본문으로 건너뛰기

#Spatial Intelligence

24개의 포스트

[논문리뷰] VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

댓글 수 로딩 중

[논문리뷰] SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

댓글 수 로딩 중

[논문리뷰] Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

댓글 수 로딩 중

[논문리뷰] Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

댓글 수 로딩 중

[논문리뷰] LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

댓글 수 로딩 중

[논문리뷰] S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

댓글 수 로딩 중

[논문리뷰] SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

댓글 수 로딩 중

[논문리뷰] Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

댓글 수 로딩 중

[논문리뷰] Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

댓글 수 로딩 중

[논문리뷰] SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

댓글 수 로딩 중

[논문리뷰] Context Unrolling in Omni Models

댓글 수 로딩 중

[논문리뷰] OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence

댓글 수 로딩 중

[논문리뷰] Learn2Fold: Structured Origami Generation with World Model Planning

댓글 수 로딩 중

[논문리뷰] Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training

댓글 수 로딩 중

[논문리뷰] Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models

댓글 수 로딩 중

[논문리뷰] Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems

댓글 수 로딩 중

[논문리뷰] Scaling Spatial Intelligence with Multimodal Foundation Models

댓글 수 로딩 중