[논문리뷰] Kaininja: Extending Native 3D Generators to the Part Level
링크: 논문 PDF로 바로 열기
저자: Ruihan Yu, Lian Fu, et al.
1. Key Terms & Definitions (핵심 용어 및 정의)
본 논문에서 다루는 핵심 용어 및 개념은 다음과 같다:
- Native 3D Generators: 단일 이미지를 입력으로 받아 3D Latent Representation을 직접 Denoise하여 3D Object를 생성하는 모델. 일반적으로 단일의 Fused Mesh를 출력한다.
- O-Voxel Representation: TRELLIS.2에서 사용하는 Sparse Voxel 구조로, 각 Active Voxel이 하나의 Dual Vertex와 Surface Patch를 저장하여 Open, Non-manifold, Enclosed Surface와 PBR Attributes를 직접 표현할 수 있게 한다.
- Dual-Volume Packing: 접촉하는 Parts가 동일한 Volume에 속하지 않도록 모든 Parts를 두 개의 Interleaved O-Voxel Volume(Stream A 및 B)으로 Packing하는 Representation으로, 기존 O-Voxel의 Part Interface Representation 문제를 해결한다.
- Rectified Flow: Diffusion Process의 Latent Space를 Denoising하기 위해 Generative Model에서 활용되는 기술로, TRELLIS.2와 KaiNinja의 핵심 Denoising Mechanism이다.
- Part-level Image-to-3D Generation: 단일 이미지 입력으로부터 각 Semantic Part가 개별적인 Self-contained Mesh로 구성된 3D Asset을 생성하는 Task.
2. Motivation & Problem Statement (연구 배경 및 문제 정의)
기존 Native 3D Generators는 단일 이미지로부터 고품질 3D Geometry를 생성하지만, 출력이 Part-level asset이 아닌 하나의 Fused Mesh 형태여서 편집, 리깅, 시뮬레이션 등 후처리(Downstream) 작업에 적합하지 않다. 기존 Part-aware 3D Generation 방법들은 대부분 3D Segmentation Network에 의존하는데, 이는 느린 처리 속도와 Segmentation Accuracy에 의해 성능이 제한되는 한계를 가진다. 또한, Part Masks나 Bounding Boxes를 사용하는 방식은 Part Overlap 또는 Fusion을 유발할 수 있다. 특히, TRELLIS.2의 O-Voxel Representation은 하나의 Voxel이 최대 하나의 Surface Sheet만 저장할 수 있으므로, 두 Part가 접촉하는 Part Interface를 효과적으로 표현하지 못하고 Collapse되는 Critical Problem이 발생한다. 더불어, 가변적인 Part 개수를 처리하기 위한 기존 방법들은 Autoregressive Generation 또는 Per-part Latent Volume allocation으로 인해 계산 비용이 크게 증가하여 확장성(Scalability) 문제가 존재한다. 마지막으로, Whole Object Pretrained Prior를 Part-level Task에 상속시키는 과정에서 Part Interface, Packed Half-objects, 두 Volume 간의 Joint Interaction과 같은 Unseen Distribution Gap을 극복해야 하는 Challenge가 있다.
3. Method & Key Results (제안 방법론 및 핵심 결과)
본 논문은 Native 3D Generator인 TRELLIS.2를 Part-level Generation으로 확장하는 KaiNinja를 제안한다. KaiNinja는 O-Voxel Representation의 Single-sheet-per-voxel 한계를 해결하기 위해 Dual-Volume O-Voxel Representation을 도입한다. 이 Representation은 Parts를 두 개의 Interleaved O-Voxel Volumes (Stream A, B)에 Packing하여 접촉하는 Parts가 동일 Volume에 속하지 않도록 한다 [Figure 3]. 이로 인해 각 Volume 내의 Geometry는 Connected-component Analysis를 통해 직접 개별 Part로 분해될 수 있으며, 별도의 Segmentation Network가 필요 없다. 제안하는 방법론은 TRELLIS.2의 Two-stage Rectified Flow를 Two-stream, Two-stage Flow로 확장한다 [Figure 2]. Stage 1 (Layout Flow)은 Coarse Occupancy를 Denoise하며, 각 Stream은 Pretrained Transformer의 자체 Copy를 가지지만 Zero-initialized Cross-volume Attention 블록을 통해 Joint Reasoning을 수행한다. Stage 2 (Refinement Flow)는 Pretrained TRELLIS.2 Backbone을 변경하지 않고 Attention Scoping을 통해 두 Stream이 Fine Geometry와 Appearance Latent를 Denoise하도록 한다.

Figure 2 — KaiNinja 파이프라인 개요

Figure 3 — Dual-volume packing 예시
Training은 Pretrained Prior를 효과적으로 상속하기 위해 Fit Each Volume, Merge Without Loss, Warm Up, Fine-tune Jointly의 다단계 Schedule을 따른다. 이 과정에서 Disjointness Penalty (ℒov)를 Layout Flow에 적용하여 Volume 간 Overlap을 줄이도록 유도한다. Inference 시에는 Merging Duplicated Occupancy와 Parts From Volumes라는 두 가지 경량(Light) Post-processing 단계를 거쳐 두 Volume을 최종 Part-separated Mesh로 조립한다. 특히, Agent-authored 3D Assets인 Articraft-10K를 포함한 다양한 Corpus를 활용하여 모델을 훈련한다.
핵심 결과로, KaiNinja는 512x512 해상도에서 단일 H100 GPU로 약 24초 만에 Part-separated Asset을 생성하며, TRELLIS.2의 Generation Speed and Quality를 유지한다 [Figure 4]. State-of-the-art Baseline인 Hunyuan3D-2.1 + X-Part와 비교했을 때, KaiNinja는 Whole-object Chamfer distance를 40% (0.0314에서 0.0186로) 감소시키고, Strict Part F-score (F10.05P)를 16% (0.597에서 0.692로) 향상시킨다 [Table 1]. 더욱이, KaiNinja는 동일한 Backbone을 동일한 Dataset으로 Fine-tuning한 결과 (F10.05W: 0.830)보다 Whole-object fidelity가 더 향상된 (F10.05W: 0.919) 성능을 보여, Dual-Volume Packing이 단순히 Parts를 담는 컨테이너를 넘어 Whole Object의 더 나은 Representation임을 시사한다 [Table 2]. Ablation Study 결과, Disjointness Penalty (ℒov)는 Stage-1 Occupancy에서 Median Cross-volume Overlap을 0.229에서 0.184로 감소시키며 [Figure 9], Post-processing의 Relabel 단계는 생성된 Part의 Component Count를 14.06에서 5.43로 효과적으로 줄여 mIoU_P와 F10.05P를 크게 개선한다 [Table 7].
4. Conclusion & Impact (결론 및 시사점)
본 논문은 Native 3D Generator인 TRELLIS.2를 Part-level Generation으로 성공적으로 확장하는 KaiNinja를 선보였다. 이 모델은 Dual-Volume O-Voxel Representation을 통해 O-Voxel의 Single-sheet-per-voxel 한계를 극복하고, Two-stream, Two-stage Adaptation으로 Pretrained Prior를 효과적으로 활용하여 Pipeline 내에 별도의 Mask나 Segmenter 없이 단일 이미지로부터 Part-separated Mesh Asset을 생성한다. KaiNinja는 Whole-object Chamfer distance 40% 감소 및 Strict Part F-score 16% 향상과 같은 우수한 State-of-the-art 성능을 달성했으며 [Table 1], 특히 Dual-Volume Packing이 Whole-object fidelity까지 향상시킨다는 예상치 못한 결과를 통해 해당 Representation의 우수성을 입증했다 [Table 2]. 이는 Part-level 3D Asset Generation 분야의 실용적 활용성을 크게 높이며, Agent-authored 3D assets를 Generative Model 훈련에 최초로 사용함으로써 새로운 데이터 소스 탐색의 중요성을 강조한다. 궁극적으로 KaiNinja는 3D Asset Creation Workflow를 효율화하고 다양한 Downstream Application에서 Part-level Control을 가능하게 하는 중요한 발판을 마련했다.
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation
- [논문리뷰] V-RAE: Rethinking Video Latent Spaces for Generation
- [논문리뷰] JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising
- [논문리뷰] ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing
- [논문리뷰] Map2World: Segment Map Conditioned Text to 3D World Generation
Review 의 다른글
- 이전글 [논문리뷰] How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
- 현재글 : [논문리뷰] Kaininja: Extending Native 3D Generators to the Part Level
- 다음글 [논문리뷰] LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
댓글