[논문리뷰] Training Object Permanence in World Models
링크: 논문 PDF로 바로 열기
The paper "Training Object Permanence in World Models" by Haotian Zhang et al. introduces WROP, a benchmark and training resource for evaluating and improving object permanence and solidity in video generation models.
Now I need to extract the information as per the requested format.
Part 1: Summary
- Authors: Haotian Zhang, Fengyuan Yu, Dezhi Luo, et al.
- Keywords: Object Permanence, Video Models, World Models, Core Knowledge
1. Key Terms & Definitions I need to identify 3-5 key terms and define them.
- Object Permanence (OP): The understanding that objects continue to exist even when they cannot be seen, heard, or touched.
- Object Solidity (OS): The understanding that two solid objects cannot occupy the same space at the same time and cannot pass through each other.
- World Model: A model that can simulate a structured physical environment, typically referring to video generation models that can predict future frames.
- WROP (World Reasoning with Object Permanence): A dedicated 3D synthetic benchmark and training dataset for evaluating and training object permanence and solidity in video generation models.
- PWM-WROP: A 16B world model fine-tuned on the WROP dataset, achieving state-of-the-art performance in its category.
2. Motivation & Problem Statement
- The core problem video generation models face is their failure to accurately represent object permanence (OP) and object solidity (OS), which are fundamental cognitive priors in human physical intelligence.
- Existing video generation models, despite producing photorealistic and temporally coherent footage, often show objects vanishing behind occluders and reappearing at impossible positions, or passing through solid barriers without deflection.
- These failures are "structurally upstream," meaning they can propagate errors to higher-level scene construction or causal reasoning, hindering the development of human-like physical intelligence in world models.
- Current benchmarks for video models have limitations such as being primarily 2D, lacking V2V evaluation, and relying on VLM-based scoring, which itself exhibits core knowledge deficits.
3. Method & Key Results
- The authors propose WROP (World Reasoning with Object Permanence), a 3D synthetic benchmark constructed in Blender, featuring 150 hand-designed cognitive science-inspired tasks across six families (three for OP, three for OS). Each task generator randomizes nuisance parameters while preserving cognitive structure, yielding 10,000+ samples per task and a total 1.5M-sample training corpus and a 300-question exam.
- The evaluation protocol involves providing models with an input video (pre-event context) and a natural-language prompt, then requiring them to generate the target video (containing the key physical event and its consequence). Human evaluators judge the generated continuations against physically consistent ground truth, bypassing the limitations of VLM judges.
- They fine-tuned PWM-WROP, a 16B world model based on Cosmos3-Nano, on the WROP training corpus. This model was trained for one epoch on 1.5 million samples at a native resolution of 320x192.
- In a blind pairwise Elo study involving 14 video models (3 reference-to-video, 7 edit, and 4 continuation models), PWM-WROP ranked first among true-continuation models with an Elo score of 1679.5, and third overall. This performance was competitive with leading commercial reference-to-video systems, even at a lower native resolution.
- Performance breakdown by task family revealed that PWM-WROP excelled in Object Static Occlusion (OP-2, rank 1) and performed strongly in Baillargeonian Occlusion (OP-1, rank 3) and Container Permanence (OP-3, rank 3). However, its performance was more variable in Object Solidity tasks, with a rank of 8 in Object Collision (OS-3).
- Qualitatively, PWM-WROP demonstrated superior identity and count preservation during occlusion (e.g., G43, OP-1) and correct reference-frame updating under container displacement (e.g., G66, OP-3), unlike baselines which showed "representation dropout" or "causal decoupling" failures.
- On automatic metrics (computed at 320x192), PWM-WROP achieved the best LPIPS (0.081) and MS-SSIM (0.921), and ranked second on CLIP similarity, indicating strong perceptual and structural fidelity to the target video despite its lower resolution.
4. Conclusion & Impact
- The study concludes that WROP provides a principled foundation for evaluating and training object permanence and solidity in video generation models.
- The fine-tuning of PWM-WROP on this cognitively principled synthetic dataset significantly improved its physical reasoning capabilities, especially in object permanence tasks, making it the top-ranked true-continuation model.
- This research offers preliminary evidence that targeted training with concept-specific synthetic data is a viable approach to instill human-like physical reasoning in video generation models.
- The release of the WROP corpus, exam, model answers, scores, weights, and the PWM native-PyTorch training stack on AWS Trainium2 is expected to serve as a crucial resource for the community to advance the understanding and improvement of physical reasoning in world models.
Part 2: Important Figure Information I need to select up to 3 figures.
- Figure 1: WROP: World Reasoning with Object Permanence. -> Overall benchmark. Good candidate.
- Figure 2: Taxonomy of the six WROP task families. -> Explains the task categories. Good candidate.
- Figure 3: Generator-level composition of WROP’s 150 generators. -> Shows task distribution, less critical than Figure 2.
- Figure 4: Generator-design example from Marked_Boxes_Swap. -> Specific example, less critical than overall architecture.
- Figure 5: Human-preference Elo of the 14 models. -> Key result, good candidate.
- Figure 6: Within-family human-preference ranks across the six WROP task families. -> Detailed results, but Figure 5 is overall.
- Figure 7: Qualitative comparison on three object-permanence task families. -> Qualitative analysis, good to show model behavior.
- Figure 8: Qualitative comparison on three object-solidity task families. -> Qualitative analysis, good to show model behavior.
- Figure 9: Per-family human-preference Elo leaderboards for all six WROP task families. -> Detailed results, similar to Figure 6.
I will choose Figure 1 (overall framework), Figure 2 (task taxonomy), and Figure 5 (main results leaderboard).
Now, extract the src attributes for these figures.
- Figure 1:
URL:https://arxiv.org/html/2609.28654v1/figures/B_rounded-soft-shadow.png - Figure 2:
URL:https://arxiv.org/html/2609.28654v1/sec/2026-09-09_figure2-taxonomy-refresh-v10.png - Figure 5: This figure is described as "Figure 5: Human-preference Elo of the 14 models." in the text, but the image is not directly shown or linked in the HTML snippet. I need to check the full HTML content again to locate Figure 5's image. Looking through the provided HTML snippet, I see
Figure 5: Human-preference Elo of the 14 models.but no<img>tag directly associated with it in the provided text. Let me try to find it again by searching forfigure5or similar. I can only see a reference to it in the text. I will assume it's2609.28654v1/family_rank_heatmap__1_.png(Figure 6) or I am missing something. Let me re-check the HTML carefully. Okay, I seeFigure 5: Human-preference Elo of the 14 models. Points are Bradley–Terry strengths on an Elo scale (mean 1500, dashed line), bars are 95% rater-clustered bootstrap intervals, colour is interface class, and PWM-WROP (filled) is the top true-continuation model.But no<img>tag follows it. Thefamily_rank_heatmap__1_.pngis explicitly for Figure 6.
I should be careful not to pick a wrong image.
The figures explicitly associated with <img> tags in the browsed content are:
- Figure 1:
2609.28654v1/figures/B_rounded-soft-shadow.png - Figure 2:
2609.28654v1/sec/2026-09-09_figure2-taxonomy-refresh-v10.png - Figure 4:
2609.28654v1/sec/2026-09-09_figure3-marked-boxes-swap-v3.png - Figure 6:
2609.28654v1/family_rank_heatmap__1_.png - Figure 7:
2609.28654v1/2026-09-21_qualitative-op-g43-step3-four-balls.png - Figure 8: There is
Figure 8: Qualitative comparison on three object-solidity task families (Figure 2, bottom row).but nosrclink for it in the browsed content. It just references Figure 7 for layout. It should be a separate image but the HTML doesn't include its source.
Given the constraint "Table is often HTML table. ONLY pick figure with image <img> tag", I need to pick from the explicit <img> tags.
Figure 5 is not associated with an <img> tag in the provided HTML. The table Table 2 presents the Elo scores, which is the quantitative data corresponding to Figure 5, but Figure 5 itself is a plot. If the image for Figure 5 is not directly linked via an <img> tag, I cannot select it.
Figure 8 also does not have an <img> tag. It references the layout of Figure 7.
So my options for <img> tags are Figure 1, Figure 2, Figure 4, Figure 6, Figure 7.
I'll pick Figure 1 (overall concept), Figure 2 (taxonomy), and Figure 7 (qualitative comparison of OP).
Figure 5 is presented in Table 2, but the visual of Figure 5 itself is not there with an <img> tag. I must follow the constraint.
Let's re-evaluate:
- Figure 1: Overall WROP framework.
https://arxiv.org/html/2609.28654v1/figures/B_rounded-soft-shadow.png- Good. - Figure 2: Taxonomy of WROP tasks.
https://arxiv.org/html/2609.28654v1/sec/2026-09-09_figure2-taxonomy-refresh-v10.png- Good. - Figure 5: Not directly available as
<img>in the provided HTML. The table 2 gives the data for it. - Figure 7: Qualitative comparison of OP tasks, which is very important for understanding model failures.
https://arxiv.org/html/2609.28654v1/2026-09-21_qualitative-op-g43-step3-four-balls.png- Good.
I will use Figure 1, Figure 2, and Figure 7. I must be careful with citations. Each sentence referring to browsed results needs a citation.
Double check the constraints:
- Technical Terminology in English: Yes.
- Natural Korean Phrasing: Yes.
- Mandatory Metadata: Yes.
- Part 1 & Part 2 structure: Yes.
- 4 본문 섹션: Yes.
- First sentence direct answer: Yes.
- Figure/Table references:
[Figure N]or[Table N]. Must refer to selected figures. - Part 2 JSON format: Yes, including absolute URLs.
Looks good. Proceed with generating the output.저자: Haotian Zhang, Fengyuan Yu, Dezhi Luo, et al.
1. Key Terms & Definitions (핵심 용어 및 정의)
본 논문에서 다루는 핵심 용어들은 World Model의 Physical Reasoning 능력, 특히 Object Permanence 및 Object Solidity와 밀접하게 관련되어 있다.
- Object Permanence (OP): 시야에서 사라진 객체가 계속 존재한다는 개념으로, 인간의 핵심 인지 능력 중 하나이다.
World Model이 가려진 객체의Identity,Count,Location을 유지하는 능력을 의미한다. - Object Solidity (OS): 두 개의 고체 객체가 동시에 같은 공간을 차지할 수 없으며 서로를 통과할 수 없다는 물리적 제약에 대한 이해를 나타낸다.
World Model이Collision,Obstruction,Support Removal과 같은 물리적 상호작용의 결과를 올바르게 생성하는 능력을 지칭한다. - World Model: 현실 세계의 동적 변화를 시뮬레이션하고 예측할 수 있는 인공지능 모델을 의미하며, 특히 비디오 생성 모델이 이러한 역할을 수행할 수 있는 잠재력이 있는 것으로 간주된다.
- WROP (World Reasoning with Object Permanence):
Object Permanence및Object Solidity훈련 및 평가를 위해Blender로 구축된 3D Synthetic Dataset 및 Benchmark이다. - PWM-WROP:
WROP데이터셋으로Fine-tuning된 16B 파라미터 규모의World Model로,True Continuation모델 카테고리에서State-of-the-Art성능을 달성했다.
2. Motivation & Problem Statement (연구 배경 및 문제 정의)
기존 Video Generation Model들은 Object Permanence (OP)와 Object Solidity (OS)라는 인간 인지의 핵심 Physical Prior를 정확하게 표현하지 못하는 근본적인 한계를 지니고 있다. Photorealistic하고 Temporally Coherent한 비디오를 생성함에도 불구하고, 객체가 Occluder 뒤에서 사라지거나 불가능한 위치에 재등장하고, Solid Barrier를 통과하는 등 물리적으로 비현실적인 현상이 빈번하게 발생한다. 이러한 Core Knowledge의 결핍은 Higher-level Scene Construction이나 Causal Reasoning에서 오류를 유발할 수 있으며, 인간과 유사한 Physical Intelligence를 갖춘 World Model 구축을 저해하는 Upstream Structural Failure로 작용한다. 현재의 Video Model Benchmark들은 대부분 2차원 환경에 국한되어 있고, Video-to-Video (V2V) Inference Protocol을 포괄하지 못하며, Core Knowledge Deficit을 보이는 Vision-Language Model (VLM) 기반의 Scoring Metric에 의존한다는 구조적 한계를 갖는다.
3. Method & Key Results (제안 방법론 및 핵심 결과)
저자들은 Video Generation Model의 Object Permanence 및 Solidity 능력을 평가하고 훈련하기 위해 WROP (World Reasoning with Object Permanence)이라는 3D Synthetic Benchmark 및 Training Corpus를 제안한다. 이 WROP은 Blender를 사용하여 150개의 Hand-designed Generator로 구성된 6가지 Cognitive Task Family를 포함하며 [Figure 2], 각 Generator는 Cognitive Structure를 유지하면서 Speed, Lighting, Camera Angle 등 Nuisance Parameter를 Randomize하여 task당 10,000개 이상의 샘플을 생성한다. Training Corpus는 총 1.5M개 샘플로 구성되어 있으며, 300개 문제로 이루어진 고정된 Evaluation Exam을 함께 제공한다. 모델들은 Input Video (이벤트 이전 Context)와 Natural Language Prompt를 받아 Key Physical Event가 포함된 Target Video를 생성하도록 요구된다. 생성된 비디오는 VLM 기반의 평가 한계를 피하기 위해 인간 평가자가 Physically Consistent한 Ground Truth Animation과 비교하여 Human Preference를 평가한다.

Figure 2 — WROP 6가지 태스크 패밀리 분류
저자들은 WROP Training Corpus에 16B 파라미터 규모의 World Model인 Cosmos3-Nano를 Fine-tuning하여 PWM-WROP을 개발했다. PWM-WROP은 1.5M 샘플로 1 epoch를 Fine-tuning했으며, 320x192 해상도로 Training 및 Evaluation되었다. 14개의 Video Model (3개의 Reference-to-Video, 7개의 Edit, 4개의 True Continuation 모델)을 대상으로 한 Blind Pairwise Elo Study 결과, PWM-WROP은 1679.5 Elo 점수를 기록하며 True Continuation 모델 중 1위를 차지했고, 전체 모델 중에서는 3위에 올랐다 [Figure 5]. 이는 Frontier Commercial System들과 Competitive한 성능이며, Native Resolution이 더 낮음에도 불구하고 달성된 결과이다. Task Family별 Performance 분석에서, PWM-WROP은 Object Static Occlusion (OP-2)에서 1위, Baillargeonian Occlusion (OP-1) 및 Container Permanence (OP-3)에서 3위를 기록하며 Object Permanence Task에서 뛰어난 능력을 보였다. Qualitative Analysis를 통해 PWM-WROP은 Occlusion 중 Ball의 Lane Identity와 Count를 정확히 보존하고 (G43, OP-1), Container Displacement에 따른 Spatial Reference Frame을 올바르게 Update하는 능력 (G66, OP-3)을 입증했다 [Figure 7]. Automatic Metric 평가 (320x192 해상도)에서는 PWM-WROP이 Perceptual Distance 지표인 LPIPS에서 0.081로 Best Performance를, Structural Similarity 지표인 MS-SSIM에서 0.921로 Best Performance를 달성했으며, CLIP Similarity에서도 2위를 기록하여 Target Video와의 Perceptual 및 Structural Fidelity가 매우 높음을 보여주었다.

Figure 7 — Object Permanence 태스크 정성적 비교
4. Conclusion & Impact (결론 및 시사점)
본 연구는 Video Generation Model이 인간의 Physical Intelligence에 필수적인 Object Permanence 및 Object Solidity와 같은 Core Representational Constraint를 내재화했는지 탐색하기 위한 WROP Benchmark 및 Training Resource를 성공적으로 도입했다. WROP Training Corpus로 Fine-tuning된 PWM-WROP은 True Continuation 모델 중 Human Preference 평가에서 가장 높은 순위를 기록하며, Cognitively Principled Synthetic Data 기반의 훈련이 Video Generation Model의 Physical Reasoning 능력을 향상시키는 Viable Path임을 시사한다. 이 연구는 Video Model이 단순한 시각적 충실도를 넘어선 Structured Physical Environment를 Simulate하고 Reason하는 World Model로 발전할 수 있는 가능성을 제시한다. WROP Corpus, Exam, Model Answer, Score, Weight, 그리고 AWS Trainium2 기반의 PWM Native-PyTorch Training Stack의 공개는 연구 커뮤니티가 Video Generation Model의 Physical Reasoning 능력을 측정하고, 이해하며, 개선하는 데 중요한 기반이 될 것으로 기대된다.

Figure 1 — WROP 벤치마크 및 리소스 개요
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] Principia: Relational Physics Tests for Video Models
- [논문리뷰] PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
- [논문리뷰] Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
- [논문리뷰] WorldOlympiad: Can Your World Model Survive a Triathlon?
- [논문리뷰] OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions
Review 의 다른글
- 이전글 [논문리뷰] Rufus-Air: An Open LLM Post-Training Recipe
- 현재글 : [논문리뷰] Training Object Permanence in World Models
- 다음글 [논문리뷰] ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation
댓글