본문으로 건너뛰기

[논문리뷰] WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

링크: 논문 PDF로 바로 열기

저자: Bohai Gu, Yueyang Yuan, Taiyi Wu, Dazhao Du, Jian Liu, Xiaoyi Pang, Jie Zhang, Xiaocheng Lu, Haobin Zhong, Xiaotong Zhao, Alan Zhao, Song Guo

1. Key Terms & Definitions

  • Interactive Video World Models (IWMs): 초기 관측 및 사용자/에이전트 액션 스트림에 조건부로 미래 시각적 관측을 생성하는 학습된 시스템입니다.
  • Compounding Errors: 오토회귀(autoregressive) 생성 패러다임에서 작은 단일 스텝 오류가 누적되어 물리적, 기하학적 일관성을 저해하는 현상입니다.
  • Reversible Action Cycles: 역방향 액션 시퀀스와 결합될 때 분석적으로 초기 상태로 돌아가는 액션 시퀀스로, annotation-free supervision을 제공합니다.
  • Spatial Closure Reward: 단일 cycle 내에서 미러링된(mirrored) 순방향-역방향 프레임 쌍을 비교하여 trajectory drift에 대한 dense supervision을 제공하는 reward입니다.
  • Temporal State Consistency Reward: 반복된 cycle 실행 전반에 걸쳐 phase-aligned 상태를 정렬하여 동일한 액션이 시간 경과에 따라 drift하는 것을 패널티하는 reward입니다.
  • CycleBench: 복잡한 action structure 하에서 state-returning 능력을 진단하기 위해 고안된 benchmark suite입니다.

2. Motivation & Problem Statement

본 논문은 Interactive Video World Models (IWMs)long-horizon planningexploration에서 겪는 compounding errors 문제를 해결하고자 합니다. 기존 Reinforcement Learning (RL) 기반 post-training 방법론들은 주로 short-horizon visual quality 또는 per-step action alignment를 최적화하며, 임의의 action sequence에 대한 ground-truth future state의 부재로 인해 long-horizon drift 문제를 해결하는 데 한계가 있었습니다 (verification bottleneck). 이로 인해 WorldPlayWorldCompass와 같은 최신 모델들조차 forward-then-backward와 같은 역방향 액션 시퀀스가 시작 상태를 복구하지 못하는 spatial closure failure와, 동일한 액션이 다른 rollout step에서 일관성 없는 displacement를 생성하는 temporal consistency failure를 겪습니다 [cite: 1, Figure 1]. 특히 composite actions의 경우 pre-training dataground-truth video demonstrations가 부족하여 기존 모델들은 simple actions 대비 최대 5배accuracy collapse를 보였습니다 [cite: 1, Table 1].

3. Method & Key Results

저자들은 long-horizon video world modelsstate-returning failure를 해결하기 위해 WorldCycle이라는 Self-Verifiable Reinforcement Learning framework를 제안합니다 [cite: 1, Figure 2]. 이 프레임워크는 일반적인 action sequence로부터 closed reversible cycles 및 이들의 반복 실행을 구성하며, trajectory-level에서 두 가지 complementary rewards를 최적화합니다. 첫째, spatial closure reward는 단일 cycle 내에서 미러링된(mirrored) 순방향-역방향 프레임 쌍을 비교하여 trajectorydrift하기 시작하는 지점을 dense하게 supervision하며, sparse endpoint signal의 문제를 해결합니다 [cite: 1, Figure 2]. 둘째, temporal state consistency reward는 반복된 cycle 실행 전반에 걸쳐 phase-aligned 프레임들을 비교하여 동일한 액션이 시간 경과에 따라 drift하는 것을 패널티합니다 [cite: 1, Figure 2]. 이 두 reward는 모델이 액션을 memorized temporal patterns이 아닌 consistent state operators로 학습하도록 유도합니다.

WorldCycle은 Warm-up-and-Combine schedule을 통해 spatial rewardwarm-up한 후 두 reward를 함께 활성화하여 catastrophic forgettinggradient conflict를 방지합니다. 또한 Multi-Scale Cycle Sampling을 적용하여 temporal-position shortcut 학습을 방지합니다. 실험 결과, WorldCycle은 state returning drift를 크게 줄였습니다. short-term simple actions에서 WorldCompass 대비 Endpoint State Closure (ESC)32%, Reverse-Path Symmetry (RPS)44% 감소시켰으며, Action Accuracy0.833으로 유사한 수준을 유지했습니다 [cite: 1, Table 1]. 특히 pre-training distribution에서 out-of-domaincomposite actions에서 Action Accuracy 0.553을 달성하여 기본 모델 대비 4배, WorldCompass 대비 11% 향상된 성능을 보였습니다 [cite: 1, Table 1]. long-term 설정(381 프레임)에서는 Repeated-Cycle Stability (RCS)34% 감소하여 temporal consistency reward의 장거리 안정성 기여를 입증했습니다 [cite: 1, Table 1].

4. Conclusion & Impact

본 논문은 reversible action cycleslong-horizon video world models을 위한 dense, annotation-free supervision으로 전환하는 self-verifiable RL frameworkWorldCycle을 제시했습니다. Spatial closuretemporal state consistency rewards는 모델이 액션을 consistent state operators로 학습하도록 유도하여 reward sparsitytemporal drift 문제를 극복하고, base modelpre-training distribution 외부에 있는 composite action cycles로까지 generalization 능력을 보였습니다. 또한, video world modelsstate-transition simulators로서 평가하는 최초의 benchmarkCycleBench를 도입했습니다. 이 연구는 physically grounded world models의 중요한 기반을 제공하며, conservation lawsequivalent terminal states와 같은 다른 analytically verifiable constraints를 통해 irreversible processesself-verifiable supervision을 확장하는 미래 연구 방향을 제시합니다.

Figure 2: WorldCycle 개요

Figure 2 — WorldCycle 개요

Figure 3: CycleBench 개요

Figure 3 — CycleBench 개요

Figure 4: CycleBench 정성적 비교

Figure 4 — CycleBench 정성적 비교

⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.

댓글

관련 포스트

Review 의 다른글