본문으로 건너뛰기

[논문리뷰] Anisotropic Representations Improve Planning in JEPA World Models

링크: 논문 PDF로 바로 열기

저자: Mingu Kang, Yoori Oh, Sookyung Kim, Joonseok Lee et al.

1. Key Terms & Definitions (핵심 용어 및 정의)

  • JEPA World Models: 학습된 Representation Space에서 Action-Conditioned Dynamics를 예측하여 Planning에 활용하는 모델.
  • Anisotropic Representations: Representation Space의 각 Dimension이 Task-Alignment를 위해 서로 다른 Importance 또는 Scale을 가지도록 학습된 표현 방식.
  • SIGReg (Sketched Isotropic Gaussian Regularization): Representation Collapse를 방지하고 학습된 Feature들이 Isotropic Gaussian Distribution을 따르도록 장려하는 Regularization 기법.
  • Planning Regret: Planner가 선택한 Action Sequence의 실제 Task Cost와 최적의 Feasible Action Sequence Task Cost 간의 차이.
  • Latent Geometry: 학습된 Representation이 형성하는 기하학적 구조로, Latent Space에서의 Distance 및 Cost 계산 방식에 영향을 미침.

2. Motivation & Problem Statement (연구 배경 및 문제 정의)

JEPA-based World Models (예: LeWorldModel, LeWM)는 학습된 Representation Space에서 Action-Conditioned Dynamics를 예측하고, Euclidean Distance를 사용하여 Goal Representation에 대한 Predicted Outcome의 Cost를 Scoring한다. 이러한 접근 방식에서 Encoder는 단순한 Feature 제공을 넘어 Planning Cost의 Geometry를 결정한다. 기존 LeWM은 Representation Collapse를 방지하기 위해 Prediction Loss와 SIGReg를 결합하여 Encoder를 Jointly Training하지만, SIGReg가 장려하는 Isotropic Gaussian Distribution은 Downstream Task Alignment와 직접적으로 관련되지 않아, Planning에 적합하지 않은 Latent Geometry를 유도할 수 있다. 결과적으로, 정확한 Prediction과 Non-Collapsed Representation에도 불구하고, Isotropic Gaussian Regularization은 Task Cost와 다른 방식으로 Feasible Outcome의 순위를 매기는 Geometry를 초래하여 Planning Regret을 발생시킬 수 있다 [Figure 1, cite: 1]. 본 연구는 이러한 Prediction–Planning Gap을 해결하기 위해 새로운 Representation Regularization 방식을 제안한다.

3. Method & Key Results (제안 방법론 및 핵심 결과)

본 논문은 Prediction–Planning Gap을 해소하기 위해 Learnable Diagonal Covariance를 사용하는 ΛReg를 적용한 AnisoWM을 제안한다. AnisoWM은 기존 SIGReg의 Fixed Isotropic Gaussian Target을 Fixed-Trace 및 Anisotropy Constraint (Condition-Number κ)를 가진 Learnable Diagonal Covariance로 대체한다. 이 방법론은 Prediction Objective, Predictor Architecture, Euclidean Planner는 그대로 유지하면서, Latent Direction에 걸쳐 Variance가 할당되는 방식을 학습하여 Latent Geometry를 Task Cost에 더 잘 Alignment 시킨다. 분석 결과, Learnable Target은 Isotropic Regularization으로 인해 발생하는 Inverse-Covariance Weighting을 상쇄할 수 있으며, 과도한 Anisotropy는 Planning Regret을 증가시킬 수 있음을 보여준다.

실험 결과, AnisoWM은 네 가지 Visual Control Environment (TwoRoom, Reacher, PushT, Cube) 모두에서 LeWM 대비 Planning Success를 향상시켰다 [Figure 2, cite: 2]. 구체적으로 TwoRoom에서 93% (LeWM: 87%), Reacher에서 89% (LeWM: 86%), PushT에서 97% (LeWM: 96%), Cube에서 79% (LeWM: 74%)의 Success Rate를 달성했다. Latent Planning Cost (Jpred)의 Task Outcome Ordering Agreement 또한 네 환경 모두에서 향상되었으며, Representation-only Cost (Jenc) 기준으로는 TwoRoom, Reacher, Cube에서 향상을 보였다. Latent Cost Neighborhoods 시각화 결과, AnisoWM은 목표 주변에서 Task Cost와 더 밀접하게 일치하는 Low-Cost Neighborhoods를 생성했다 [Figure 3]. Cube에서 Latent Cost와 Task Cost 간의 Spearman’s Rank Correlation (ρ)은 0.15에서 0.91로, PushT에서는 0.65에서 0.92로 크게 증가했다. Learned Target Spectra는 Shared Anisotropy Bound κ=2에도 불구하고 환경마다 다르게 나타났으며, 허용된 Anisotropy의 증가가 Planning Performance를 단조적으로 향상시키지 않음을 확인했다.

Figure 3: 목표 주변 잠재 공간 비용

Figure 3 — 목표 주변 잠재 공간 비용

4. Conclusion & Impact (결론 및 시사점)

본 연구는 Isotropic Gaussian Regularization이 정확한 Prediction과 Non-Collapsed Representation에도 불구하고 Task-Aligned Planning Cost를 보장하지 못함을 이론적으로 밝혀냈다. 이에 대한 해결책으로 Learnable Anisotropic Gaussian Target을 사용하는 AnisoWM과 ΛReg를 제안하였으며, 이는 Prediction Objective와 Euclidean Planner를 변경하지 않고 Representation Geometry를 효과적으로 재구성한다. 이 연구는 Prediction-Driven Representation Geometry 선택을 특성화하고 Planning Regret을 줄이는 조건을 확립하였다. 실험적으로 AnisoWM은 네 가지 Visual Control Environment에서 LeWM보다 높은 Planning Success를 달성했으며, Latent Cost가 기록된 Task Outcome과 더 높은 Agreement를 보였다. 이는 World Models의 Planning 능력을 향상시키는 데 있어 Anisotropic Representations의 중요성을 강조하며, 향후 Representation Learning 및 Reinforcement Learning 분야에서 보다 효율적이고 Task-Aligned된 Planning 전략 개발에 기여할 수 있는 중요한 시사점을 제공한다.

Figure 1: 예측-계획 분리 현상

Figure 1 — 예측-계획 분리 현상

⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.

댓글

관련 포스트

Review 의 다른글