[논문리뷰] Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue본 논문은 대화형 AI에서 발화와 신체 움직임이 분리되어 생성되는 기존 Cascade 방식의 한계점을 해결하고자 합니다. 기존 방식은 발화 생성 후 별도의 모션 모델을 수행하여 추론 지연(Latency)이 발생하며, 두 모듈 간의 공동 최적화(Joint optimization)가 불가능하다는 구조적 결함이 있습니다.#Review#Spoken Dialogue Models#Co-speech Motion#End-to-End#Multimodal Co-learning#Joint Optimization#Progressive Training#Pseudo-labeling2026년 9월 6일댓글 수 로딩 중