[논문리뷰] BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation기존의 3D Vision-Language-Action(VLA) 모델들은 대규모 데이터셋에 과도하게 의존하며, 작업 수행 중 발생하는 시야 가림(occlusion)이나 과거 상태에 의존해야 하는 작업(memory-dependent task)을 해결하는 데 한계를 보입니다.#Review#Vision-Language-Action Models#3D Manipulation#Spatio-Temporal Memory#Heatmap Prediction#Data-Efficient#Robot Learning2026년 8월 5일댓글 수 로딩 중