[논문리뷰] DeltaWAM: Delta World Action Models for Bimanual Manipulation본 논문은 기존 WAM이 매 관측 단계마다 고차원 비디오를 생성하거나 dense한 비디오 표현을 반복적으로 처리함으로써 발생하는 계산 비효율성과 비효율적 추론 문제를 해결하고자 합니다.#Review#World-Action Models#Bimanual Manipulation#Visual Deltas#Streaming Delta Memory#Flow Matching#Robot Control#Efficiency2026년 9월 24일댓글 수 로딩 중
[논문리뷰] EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents본 연구는 Vision-Language-Action (VLA) 모델이 long-horizon tasks를 수행할 때 발생하는 복잡한 조율 문제와 기존 방법론의 한계를 해결하고자 한다.#Review#Vision-Language-Action (VLA)#Embodied Agent#Robot Control#Skill Orchestration#Closed-Loop System#Modular Framework#Execution Verification#Reinforcement Learning2026년 9월 7일댓글 수 로딩 중
[논문리뷰] PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control본 논문은 기존 VLA 모델들이 사전 학습된 풍부한 문맥 이해 능력을 로봇 제어의 에피소드 메모리로 활용하지 못하고, 별도의 복잡한 메모리 기법을 추가하는 구조적 비효율성 문제를 해결하고자 합니다 .#Review#Multimodal Large Language Models#Robot Control#Episode Memory#System 2 Cognition#Vision-Language-Action Models#Asynchronous Interface#End-to-End Training2026년 8월 30일댓글 수 로딩 중
[논문리뷰] GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch본 논문은 기존 WAM 방식이 추론 시 명시적인 미래 비디오 생성을 요구하여 발생하는 높은 연산 오버헤드와 실시간 제어의 한계를 해결하는 것을 목표로 합니다.#Review#World Action Models#Robot Control#Mixture-of-Transformers#AutoResearch#Inference Latency#Flow Matching#Visual Dynamics2026년 7월 15일댓글 수 로딩 중
[논문리뷰] PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence본 연구는 시점 불일치 문제로 인해 로봇 일반화에 한계가 있는 기존 VLM(Vision-Language Model)의 단점을 해결하고자 합니다.#Review#Egocentric Data#Physical Intelligence#VLM#Robot Control#Embodied AI#VQA Supervision#Human-Robot Interaction#Zero-shot Transfer2025년 12월 21일댓글 수 로딩 중
[논문리뷰] EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot Control본 연구는 기존 VLA 모델들이 가진 제한된 도메인 및 유연성 문제를 해결하고, 개방형 환경에서 인간 수준의 유연한 다중 모달 추론 및 물리적 상호작용 을 가능하게 하는 일반ist 로봇 제어를 목표로 합니다.#Review#Embodied AI#Robot Control#Vision-Language-Action Models#Multimodal Pretraining#Flow Matching#Foundation Models#Generalization#Real-world Robotics2025년 9월 1일댓글 수 로딩 중
[논문리뷰] Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies본 논문은 기존 Vision-Language-Action (VLA) 모델 디코더의 한계(고정된 순서의 autoregressive 생성 또는 continuous diffusion /flow matching 헤드의 백본 분리)를 해결하고자 합니다.#Review#Vision-Language-Action (VLA)#Discrete Diffusion#Action Decoding#Transformer#Robot Control#Masked Modeling#Adaptive Decoding#Reinforcement Learning2025년 8월 28일댓글 수 로딩 중