[논문리뷰] N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation본 논문은 정밀한 접촉이 필요한 manipulation 작업에서 기존의 순수 visuomotor 회귀 방식이 갖는 한계인 '촉각 정보의 부재'와 '미래 예측 능력의 결여'를 해결하고자 합니다.#Review#Tactile-Native#World-Action Model#Mixture-of-Transformers#Contact-Rich Manipulation#Flow-Matching#NeoForce2026년 8월 2일댓글 수 로딩 중
[논문리뷰] StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation본 논문은 기존의 게임 월드 모델이 시각적 현실성은 뛰어나지만, 게임의 핵심인 규칙(Mechanics)을 위반하는 Mechanics Inconsistency 문제를 해결하는 것을 목표로 합니다.#Review#Game World Models#Mechanics-Consistent Generation#State-Aware#Mixture-of-Transformers#Flow Matching#Game States2026년 7월 29일댓글 수 로딩 중
[논문리뷰] Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers본 논문은 다중 에이전트 환경에서 기존 비디오 확산 모델(Video Diffusion Model)이 겪는 논리적 및 기하학적 일관성 부족 문제를 해결합니다. 기존 모델들은 파편화된 에이전트별 관측치만을 조건부 컨텍스트로 사용하므로, 에이전트 간 공유되는 세계 상태를 유지하는 데 한계가 있습니다 .#Review#World Models#Multi-Agent Generation#Video Diffusion#World State Registers#Mixture-of-Transformers#Autoregressive Modeling2026년 7월 23일댓글 수 로딩 중
[논문리뷰] GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch본 논문은 기존 WAM 방식이 추론 시 명시적인 미래 비디오 생성을 요구하여 발생하는 높은 연산 오버헤드와 실시간 제어의 한계를 해결하는 것을 목표로 합니다.#Review#World Action Models#Robot Control#Mixture-of-Transformers#AutoResearch#Inference Latency#Flow Matching#Visual Dynamics2026년 7월 15일댓글 수 로딩 중
[논문리뷰] ABot-M0.5: Unified Mobility-and-Manipulation World Action Model본 논문은 모바일 매니퓰레이션(mobile manipulation) 환경에서 기존의 Embodied Learning 방식들이 겪는 구조적 한계를 해결하고자 합니다.#Review#Mobile Manipulation#World Action Model#Conditional Flow Matching#Latent Actions#Mixture-of-Transformers#Dream Forcing2026년 7월 1일댓글 수 로딩 중
[논문리뷰] Cosmos 3: Omnimodal World Models for Physical AIPhysical AI 에이전트 학습을 위한 기존의 파편화된 파이프라인은 이해(Understanding)와 생성(Generation) 모듈이 분리되어 있어 데이터 효율성과 확장성이 낮습니다.#Review#World Model#Physical AI#Mixture-of-Transformers#Omnimodal#Data-Driven Specialization#Synthetic Data#Action-Conditioned Generation2026년 6월 3일댓글 수 로딩 중
[논문리뷰] EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers본 논문은 기존의 Diffusion 기반 3D 생성 모델들이 의미론적 이해(semantic understanding)와 기하학적 추론(geometric reasoning)을 분리하여 처리함으로써 발생하는 한계를 해결하고자 합니다.#Review#Multimodal Large Language Models#Mixture-of-Transformers#3D Native Generation#Context-aware Editing#Flow Matching#Sparse Voxel Representation2026년 6월 1일댓글 수 로딩 중
[논문리뷰] Representation Forcing for Bottleneck-Free Unified Multimodal Models본 논문은 기존 UMM이 frozen VAE에 의존하여 발생하는 structural bottleneck 문제를 해결하기 위해 Representation Forcing (RF)을 제안한다 .#Review#Unified Multimodal Models#Representation Forcing#Pixel-space Diffusion#Vector Quantization#End-to-End Learning#Bottleneck-Free#Mixture-of-Transformers2026년 5월 31일댓글 수 로딩 중
[논문리뷰] HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents본 논문은 모달리티 적응형 컴퓨팅을 위한 MoT 아키텍처와 비전-언어 연결을 강화하는 Visual Latent Tokens를 핵심 방법론으로 제안합니다 . 시각적 인지 능력 향상을 위해 HY-ViT 2.0 인코더를 탑재하고, 고품질 embodied 데이터를 활용한 반복적인 사후 학습 패러다임을 설계했습니다.#Review#Embodied Foundation Models#Mixture-of-Transformers#Visual Latent Tokens#On-policy Distillation#Chain-of-Thought#Real-world Agents2026년 4월 9일댓글 수 로딩 중
[논문리뷰] UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving본 논문은 VLA 모델을 자율주행에 적용할 때 발생하는 공간 인지와 의미론적 추론 간의 근본적인 충돌 문제를 해결하고자 합니다. 기존의 VLA 시스템들은 주로 사전 학습된 2D VLM을 기반으로 하는데, 이는 강력한 의미론적 이해 능력을 갖춘 반면 자율주행에 필수적인 공간 인지 능력이 부족하다는 한계를 지닙니다.#Review#Vision-Language-Action Models#Autonomous Driving#Mixture-of-Transformers#Sparse Perception#Representation Interference#End-to-End Planning2026년 4월 2일댓글 수 로딩 중
[논문리뷰] Video-As-Prompt: Unified Semantic Control for Video Generation이 논문은 비디오 생성 분야에서 통합적이고 일반화 가능한 의미론적 제어라는 중요한 과제를 해결하고자 합니다. 기존 방법론들이 부적절한 픽셀 단위 사전 정보를 강요하여 아티팩트를 생성하거나, 특정 조건에 대한 파인튜닝이나 태스크별 아키텍처에 의존하여 일반화가 어렵다는 문제를 극복하는 것을 목표로 합니다.#Review#Video Generation#Semantic Control#Diffusion Transformers#In-Context Learning#Mixture-of-Transformers#Video-As-Prompt#Controllable Generation#Large-scale Dataset2025년 10월 27일댓글 수 로딩 중