[논문리뷰] RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
링크: 논문 PDF로 바로 열기
저자: Chang Guo, Yukun Xie, Bohan Tan, Zheng Chang, Zhaokai Yin, Qianli Ma, Yingqiao Wang, Chao Liang, Zhipeng Zhang
1. Key Terms & Definitions (핵심 용어 및 정의)
본 섹션에서는 논문에서 다루는 핵심 기술 용어 및 개념들을 정의한다.
- Low Scene Entropy: 시각적 장면이 오직 하나의 유효한 task만을 허용하여, language의 역할이 중복되거나 미미해지는 특성을 지칭한다.
- High Scene Entropy: 시각적 장면이 여러 kinematically distinct한 task branches를 지원하여, vision만으로는 의도된 행동을 식별하기 불충분하고 language에 의존해야 하는 특성을 의미한다.
- Vision-Language-Action (VLA) Models: 사전 훈련된 Vision-Language Models (VLM)을 확장하여 internet-scale perception을 물리적 실행(physical execution)과 연결하는 모델들을 총칭한다.
- World Action Models (WAM): 환경 dynamics의 예측 모델링과 action generation을 결합함으로써 reactive visuomotor mapping을 넘어선 모델들을 일컫는다.
- Intent Score (IS): 정책이 각 단계에서 의도된 object, relation, waypoint, action primitive, orientation 또는 logical branch와 같은 올바른 semantic target을 선택했는지를 측정하는 지표로, 물리적 실행이 불완전하더라도 semantic grounding에 중점을 둔다.
- Execution Score (ES): 해당 subgoal(예: grasping, transporting, placing)이 물리적으로 성공적으로 완료되었는지를 측정하는 지표로, kinematic proficiency를 나타낸다.
- Hierarchical Diagnostic Protocol (L0-L3): instruction following 능력을 점진적으로 테스트하기 위해 시각적 레이아웃과 semantics에 perturbation을 가하는 4단계 프로토콜을 말한다.
L0는 in-distribution 성능을,L1은 시각적 레이아웃 변경 시의 visual grounding을,L2는 familiar visual context 하의 semantic recombination을,L3는 visual 및 semantic perturbation의 결합을 평가한다.
2. Motivation & Problem Statement (연구 배경 및 문제 정의)
본 논문은 현대 embodied agents의 높은 task success rate가 실제 instruction-following 능력과는 괴리가 있을 수 있다는 문제를 제기한다. 기존 manipulation benchmarks는 대부분 low scene entropy를 특징으로 하여, 시각적 단서만으로도 task를 유추할 수 있어 language instruction이 redundant해지거나 단지 task identifier로만 활용될 가능성을 내포한다. 이러한 visual shortcuts는 정책이 언어를 무시하거나 약하게 사용하여도 높은 success rate를 달성하는 instruction following mirage를 유발하며, 이는 deployable robots의 semantically wrong한 행동으로 이어질 수 있다. 기존 평가 프로토콜들은 visual recognition, language grounding, planning, low-level control이 얽혀 있어 instruction following 능력을 진단하기 어렵고, linguistic perturbations에 대한 VLA policies의 insensitivity가 관찰되기도 했다. 따라서 저자들은 embodied agents가 언어를 통해 의도된 행동을 선택하고 실행하는지를 진정으로 평가할 수 있는 새로운 diagnostic benchmark의 필요성을 강조한다.
3. Method & Key Results (제안 방법론 및 핵심 결과)
저자들은 embodied agents가 language를 task identification에 반드시 사용하도록 강제하는 RoboFollow 벤치마크를 제안한다. 이 벤치마크는 high-ambiguity scenes를 구성하여, 동일하거나 매우 유사한 시각적 구성이 여러 semantically valid하고 kinematically feasible한 task branches를 지원하게 함으로써, vision만으로는 의도된 행동을 식별하기 어렵게 만든다. RoboFollow는 Extrinsic Spatial Relations, Intrinsic Object Properties, Fine-Grained Action Modulation, Elementary Logical Grounding의 네 가지 scene family를 통해 instruction following을 평가하며, 각 scene은 L0 (In Distribution), L1 (Visual Grounding), L2 (Semantic Compositionality), L3 (Visual-Semantic Mixture)의 네 가지 hierarchical test levels로 난이도를 조절한다 [cite: 1, Figure 1]. 또한, semantic misunderstanding과 motor failure를 분리하기 위해 Intent Score (IS)와 Execution Score (ES)로 구성된 Multi-Stage Intent-Execution Scoring 프레임워크를 도입한다.
체계적인 실험 결과, π-series 모델을 포함한 9가지 VLA 및 WAM 모델들은 L0에서 높은 in-distribution 성능을 보였지만, L1-L3와 같은 generalization levels에서는 Intent Score가 급격히 저하되는 일관된 패턴을 보였다 [cite: 1, Table 1]. 예를 들어, π0.5 모델은 Scene 1에서 L0의 IS 99.1%에서 L1 45.5%로, L3 44.9%로 크게 하락했으며, Scene 2에서는 L0 100.0%에서 L3 34.2%로 감소했다 [cite: 1, Table 1]. 이는 기존 모델들이 언어 instruction을 shallow task identifiers로 처리하며, semantic grounding과 generalization 능력이 부족함을 시사한다. 더 강력한 VLM backbones, QA co-training, LangForce, 그리고 Classifier-Free Guidance (CFG)와 같은 완화 전략들을 적용했음에도 불구하고, L0를 넘어선 generalization gap을 해소하는 데는 실패했다 [cite: 1, Table 2]. 특히, Qwen3-VL-4B와 같은 강력한 VLM을 통합한 Qwen-GR00T도 out-of-distribution instruction following에서 catastrophic collapse를 보였으며 [cite: 1, Table 2], CFG는 오히려 L0 성능조차 저하시키는 결과를 나타냈다 [cite: 1, Table 2]. 이러한 실패 사례들은 semantic reasoning과 instruction grounding의 약점에서 비롯되며, 특히 spatial relationships, object attributes의 composition, fine-grained trajectory or orientation constraints, 그리고 logical constructs 처리 능력에서 문제가 발견되었다 [cite: 1, Figure 7].
4. Conclusion & Impact (결론 및 시사점)
본 연구는 RoboFollow 벤치마크를 통해 embodied agents가 언어 지시를 진정으로 따르는 능력에 대한 illusion을 해소하고, 실제 instruction following 능력이 취약함을 명확히 밝혔다. RoboFollow는 high-entropy scene design, hierarchical L0–L3 protocol, 그리고 decoupled intent-execution scoring을 결합하여 spatial, attribute-based, procedural, logical semantics 등 다양한 차원에서 instruction following 실패를 진단한다. 실험 결과, 현재의 VLA 및 WAM 정책들은 최소한의 시각적 및 semantic perturbations 하에서 성능이 급격히 저하되며, 이는 기존의 optimization strategies로도 해결되지 않는 critical bottleneck임을 시사한다. 이 연구는 controllable하고 reliable한 embodied agents 개발을 위해 robust instruction following에 대한 심도 깊은 연구가 필요함을 강조하며, 향후 더 넓은 data distributions 및 alternative training procedures를 통해 이 문제를 해결할 필요성을 제시한다.

Figure 1 — RoboFollow scene 디자인 개요

Figure 7 — 정책의 대표적인 실패 사례
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] Latent Action as Intention Enables Efficient Future Imagination for World Action Models
- [논문리뷰] DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
- [논문리뷰] ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
- [논문리뷰] TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
- [논문리뷰] Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
Review 의 다른글
- 이전글 [논문리뷰] Recursive self-improvement of AI research agents
- 현재글 : [논문리뷰] RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
- 다음글 [논문리뷰] StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training
댓글