[논문리뷰] To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
링크: 논문 PDF로 바로 열기
저자: Xiaobin Huang, Zilong Huang, Yang Luo, Hongchao Fan, Yiping Chen, Ting Han
1. Key Terms & Definitions (핵심 용어 및 정의)
- HoloWorld:
continuously updated cross-scale world context를 기반으로 구축된unified indoor-outdoor urban world generation프레임워크입니다. - Cross-Scale World Context ($\mathcal{C}$):
semantic descriptions,visual appearances,spatial layouts,building geometries를 연결하여generation전반에 걸쳐spatial scales간의consistency를 유지하는structured hierarchical representation(world-level, block-level, building-level)입니다. - Autoregressive Exterior Generation: 이전에 생성된
neighboring blocks에 따라urban exteriors가block-by-block으로 생성되어spatial및visual continuity를 보장하는 프로세스입니다. - Building-Level Indoor-Outdoor Correspondence: 생성된
interior scene과 그에 상응하는exterior building instance간의explicit functional,visual,spatial alignment를 의미하며,building의footprint및inherited appearance준수를 포함합니다. - Absolute Quantitative Scoring (AQS):
Structural and View Consistency (SVC),Scene Richness and Complexity (SRC),Material and Texture Fidelity (MTF),Lighting and Atmosphere (LA)와 같은 다양한dimensions에서 절대 점수(1-10)를 부여하여urban exterior generation을 평가하는metric입니다. - Shape IoU: 생성된
indoor envelope와exterior building footprint간의correspondence를 정량화하는evaluator-independent geometric metric입니다.
2. Motivation & Problem Statement (연구 배경 및 문제 정의)
기존 text-driven 3D generation methods는 대규모 outdoor environments와 세부적인 indoor scenes를 생성하는 데 상당한 발전을 이루었지만, 이러한 domains는 주로 독립적으로 synthesize되어 coherent urban world에 필요한 correspondence가 부족합니다. 이로 인해 urban exteriors와 interiors가 별도로 생성될 때 동일한 world와 일치하지 않아 scales 전반에 걸쳐 semantic, visual, geometric consistency를 유지하지 못하는 문제가 발생합니다. 현재 urban generation methods는 주로 exterior environments를 구축하는 데 중점을 두며 개별 buildings의 internal spaces를 모델링하지 않는 반면, indoor generation methods는 일반적으로 기존 urban context에 grounding되지 않은 isolated scenes를 생성합니다. 따라서 city-level planning부터 개별 building instances에 이르기까지 generation 과정에서 contextual information (semantic, visual, geometric)을 preserve, propagate, localize하는 것이 핵심적인 challenge입니다.
3. Method & Key Results (제안 방법론 및 핵심 결과)
본 논문은 HoloWorld를 제안하며, 이는 cross-scale world context($\mathcal{C}$)를 기반으로 하는 통합 indoor-outdoor urban world generation 프레임워크입니다 [cite: 1, Figure 2]. 이 world context는 사용자 description에서 초기화되어 city-scale planning부터 개별 buildings에 이르기까지 diverse world information을 점진적으로 표현하고 update합니다 [cite: 1, Figure 1, Figure 2]. HoloWorld는 evolving context와 이전에 생성된 neighboring blocks에 따라 consistent spatial organization 및 visual identity를 가진 urban exteriors를 autoregressively 생성합니다 [cite: 1, Figure 2]. 생성된 exterior representations는 3D building instances 및 footprints에 grounded되어, geometry-constrained layouts와 inherited appearance characteristics를 갖는 building-specific indoor generation을 가능하게 합니다 [cite: 1, Figure 2].
Urban exterior generation 성능 평가에서, HoloWorld는 AQS score에서 SOTA baseline 대비 평균 7.68% 향상을 달성했으며, 모든 RDR score에서 가장 높은 수치를 기록했습니다 [cite: 1, Table 1]. 특히, GPT-5.5-based evaluation에서 MajutsuCity 대비 SVC, SRC, MTF, LA AQS scores에서 각각 9.38%, 4.65%, 9.00%, 8.11%의 상대적 향상을 보였습니다 [cite: 1, Table 1]. Building-level indoor-outdoor correspondence 평가에서는 TRELLIS baseline 대비 visual coherence에서 큰 향상을 보였으며, Visual AQS가 GPT-5.5에서 5.88에서 8.21로 증가하고 RDR은 1.83에서 24.29로 증가했습니다 [cite: 1, Table 2]. 또한, HoloWorld는 Shape IoU에서 0.997을 달성하여 geometric grounding의 강점을 입증했습니다 [cite: 1, Table 2]. Ablation studies에서는 dynamic world-context updates가 building-level indoor-outdoor correspondence에 필수적이며, autoregressive neighborhood conditioning이 cross-block continuity를 유지하는 데 중요함을 확인했습니다 [cite: 1, Table 3, Table 4, Figure 4].
4. Conclusion & Impact (결론 및 시사점)
HoloWorld는 indoor와 outdoor urban scene generation을 coherent 3D urban world의 corresponding realizations로 다루는 unified framework를 제시합니다. continuously updated cross-scale world context를 통해 city-level planning에서부터 개별 buildings에 이르기까지 semantic, visual, spatial information을 transfer하고 share함으로써, 각 generated interior가 explicit correspondence를 유지하도록 합니다. 또한, context-aware autoregressive generation은 neighboring urban blocks 간의 spatial 및 visual continuity를 보존합니다. 이 연구는 기존에 분리되었던 indoor 및 outdoor generation processes를 연결하여 unified 3D urban world의 coherent components를 형성하게 함으로써, virtual worlds, embodied-agent simulation 및 interactive 3D applications 분야에 중요한 impact를 미칠 것으로 기대됩니다. 향후 연구에서는 multi-floor architectural reasoning 및 richer structural constraints로 확장할 계획입니다.

Figure 1 — 개념적 비교

Figure 2 — HoloWorld 개요

Figure 4 — 정성적 결과 및 Ablation
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] MiniWorld: Democratizing the Training of Video World Models from Scratch
- [논문리뷰] Vidu S1: A Real-Time Interactive Video Generation Model
- [논문리뷰] Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model
- [논문리뷰] LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
- [논문리뷰] Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation
Review 의 다른글
- 이전글 [논문리뷰] The Attention Triangle in Audio-Video Models
- 현재글 : [논문리뷰] To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
- 다음글 [논문리뷰] Training-Free Speech-Centric Omni Understanding with Frozen VLMs
댓글