[논문리뷰] PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
링크: 논문 PDF로 바로 열기
저자: DeepCybo Team, Yu Bin, Haipeng Cao, et al.
1. Key Terms & Definitions (핵심 용어 및 정의)
본 논문은 Vision-Language Models (VLM)을 Physical Foundation Models로 확장하기 위한 여러 핵심 용어와 기술을 정의합니다.
- Physical Loop: 에이전트가 환경을 관찰(observation)하고, 상호작용(interaction)하며, 환경 변화를 통해 새로운 관찰을 얻는 순환 과정을 지칭합니다. PhysBrain 1.5는 이 루프 내에서 이해(understanding), 행동(acting), 예측(predicting)의 재사용 가능한 기능을 제공합니다.
- Unified Vocabulary: 기존 언어(language) vocabulary에 행동(action) 및 시각(visual) tokens를 추가하여
V_lang U V_act U V_vis형태로 확장된 vocabulary입니다. 이를 통해 언어 응답, end-effector motion, multimodal 미래 상태를 단일 autoregressive objective 하에 jointly optimize합니다. - ActionPiece: End-effector trajectories를 discrete sequences로 인코딩하는 메커니즘으로, human motion 데이터를 로봇 행동 데이터와 동일한 형태로 표준화하여 모델이 next-token prediction을 통해 행동 시퀀스를 생성할 수 있도록 합니다.
- VQ-VAE: Future visual states(RGB 이미지, depth map, robot mask)를 discrete tokens으로 변환하는 데 사용되는 VQ-VAE (Vector Quantized-Variational AutoEncoder)입니다. 이를 통해 시각적 정보를 언어 및 행동 token과 동일한 framework에서 처리할 수 있습니다.
- Embodied Understanding Benchmarks: 모델의 physical perception, spatial and multi-view understanding, embodied cognition and planning, spatial grounding and affordance, visual trace and trajectory reasoning 능력을 평가하는 28가지 벤치마크 모음입니다.
2. Motivation & Problem Statement (연구 배경 및 문제 정의)
현재 Vision-Language Models (VLM)은 시각적 이해와 Semantic 지식, Instruction Following에 강점을 보이지만, Physical Loop 내에서의 embodied understanding, action generation, future-state prediction과 같은 물리적 상호작용 능력은 부족합니다. 기존 연구들은 Specialized 모델들을 통해 spatial reasoning, grounding, affordance, planning, execution assessment 등 특정 영역을 다루거나, 여러 기능을 통합하려는 시도를 했지만, 종종 서로 다른 supervision 방식, 데이터 포맷, temporal granularity 문제에 직면했습니다. 이러한 한계점으로 인해 다양한 embodiment와 task에 걸쳐 일반화될 수 있는 통합 학습 프레임워크가 필요하며, 특히 dense visual outputs와 temporally structured motion representations를 공동으로 학습하는 것은 큰 과제였습니다. PhysBrain 1.5는 이러한 문제를 해결하기 위해 observation, interaction, environmental change의 Physical Loop에 기반한 통합 학습 프레임워크를 제안합니다.
3. Method & Key Results (제안 방법론 및 핵심 결과)
PhysBrain 1.5는 Qwen3-VL-Instruct (8B) 모델을 기반으로, 언어, end-effector motion, 시각적 상태 타겟을 discrete sequences로 표현하고 동일한 autoregressive next-token prediction objective로 joint training합니다 [Figure 2]. 이를 위해 기존 언어 vocabulary에 ActionPiece를 이용한 행동 토큰과 VQ-VAE를 이용한 시각 토큰을 추가하여 Unified Vocabulary를 구성합니다. Pre-training은 인간 상호작용 비디오에서 physical-aware supervision을 얻고, supervised fine-tuning 단계에서는 인간 데모, 실제 로봇 trajectories, 시뮬레이션 경험을 통합합니다.

Figure 2 — PhysBrain 1.5의 전체 시스템 아키텍처를 설명하는 핵심 다이어그램
주요 정량적 결과는 다음과 같습니다.
- Embodied Understanding Benchmarks: 28가지 embodied understanding 벤치마크 스위트에서 72.5의 평균 점수를 달성하며, 모든 open-source 모델 중 새로운 State-of-the-Art를 기록했습니다 [Figure 1, Table 4]. 이는 GPT-6-Astra (73.3) 및 Gemini 3.6 Flash (73.0)와 같은 선도적인 proprietary 모델들과 대등한 성능입니다. 특히, PhysBrain 1.5는 14개 벤치마크에서 1위, 10개 벤치마크에서 2위를 차지했습니다.
- General Multimodal Understanding: 일반적인 multimodal understanding 벤치마크에서는 기본 모델인 Qwen3-VL-Instruct (8B)와 유사한 성능을 유지했습니다 [Table 5, 6]. 예를 들어, VideoMME 벤치마크에서 PhysBrain 1.5는 68.22점을 기록하여 Qwen3-VL-Instruct의 69.07점에 근접했으며, MVBench에서는 66.07점으로 Qwen3-VL-Instruct의 68.53점에 준하는 결과를 보였습니다.
4. Conclusion & Impact (결론 및 시사점)
본 논문은 PhysBrain 1.5를 통해 에이전트가 환경을 관찰하고, 행동하며, 그 변화를 바탕으로 다음 행동을 결정하는 Physical Loop를 따르는 통합 모델을 성공적으로 제시했습니다. 이 모델은 Shared Autoregressive Backbone을 통해 언어 응답, end-effector trajectories, multimodal visual states를 discrete tokens으로 표현하며 embodied understanding, action generation, future-state prediction을 지원합니다. 물리적 환경에서의 interaction 경험을 학습하여 open-source 분야에서 28개 embodied understanding 벤치마크에서 72.5의 평균 점수로 SOTA를 달성하며, 선도적인 proprietary 모델들과 동등한 성능을 보여주었습니다. 이러한 결과는 상호작용 경험을 통한 학습이 Physical Loop에 필요한 역량을 개발하는 효과적인 경로임을 시사하며, 향후 훈련 데이터 규모 확장, modality coverage 확대, 풍부한 상호작용 경험 통합을 통해 Physical AI 분야에 큰 영향을 미칠 것으로 기대됩니다.

Figure 1 — 모델의 핵심 정량적 성능 (리더보드 순위 및 점수)을 시각적으로 보여주는 주요 Figure

Table 4 — 28개 Embodied Understanding 벤치마크에 대한 PhysBrain 1.5와 비교 모델들의 상세 정량적 성능 비교 테이블
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] Show-Harness: Just a VLM Agent Can Play Robots
- [논문리뷰] Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation
- [논문리뷰] LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
- [논문리뷰] SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- [논문리뷰] HumanCLAW: Can Vision-Language Models Act Through a Body?
Review 의 다른글
- 이전글 [논문리뷰] Omni-Streaming Thinking
- 현재글 : [논문리뷰] PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
- 다음글 [논문리뷰] Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
댓글