[논문리뷰] SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning기존 Vision-Language Models (VLMs)은 multi-view spatial reasoning을 위해 pre-trained reconstruction models의 geometric priors를 활용하는 경향이 있었습니다.#Review#Multi-view Vision-Language Models#Spatial Reasoning#3D Reconstruction#Chain-of-Thought#QA-Native Pretraining#Visual Compensation#Geometric Priors2026년 9월 28일댓글 수 로딩 중