[논문리뷰] Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
링크: 논문 PDF로 바로 열기
Total of 1.
The paper is about Document Retrieval-Aware Chunking (D-RAC). I have the content. Now I need to structure the summary.
Part 1: Markdown Summary
- Authors: Uday Allu, Abhivanth Sivaprakash, Pratik Singh, Aman Manocha
- Keywords: Need to select 5-8 keywords based on the paper content. Good candidates:
RAG,Document Ingestion,PDF Normalization,Multimodal LLM,Retrieval-Aware Chunking,Markdown Conversion,Cost Efficiency,Enterprise Documents. - Key Terms & Definitions:
- Retrieval-Augmented Generation (RAG): The dominant paradigm for grounding large language models in enterprise knowledge.
- Document Retrieval-Aware Chunking (D-RAC): An extension of W-RAC for arbitrary document formats, leveraging PDF normalization and multimodal Markdown conversion for retrieval-optimized ingestion.
- PDF Normalization: The process of converting any input document format (DOCX, PPTX, XLSX, scanned images) into a standardized PDF representation.
- Multimodal Markdown Conversion: A single pass by a multimodal LLM to convert rendered PDF pages into retrieval-optimized Markdown, preserving hierarchy and converting tables into prose.
- Agentic Chunking: Applying an LLM to raw extracted text to produce semantically coherent chunks, often involving text regeneration.
- Motivation & Problem Statement:
- RAG systems struggle with heterogeneous enterprise document formats (PDFs, Word, presentations, scans) due to complex visual layouts, multi-column pages, and dense tables.
- Traditional ingestion pipelines (rule-based text extraction, OCR) often destroy reading order, flatten tables, and lose heading hierarchy, degrading retrieval quality. [Table 1]
- Existing agentic chunking methods over raw text incur high token costs and hallucination risks due to text regeneration.
- W-RAC (Web Retrieval-Aware Chunking) works well for structured web content (HTML), but PDF's presentation-centric nature lacks the reliable structure W-RAC depends on.
- Method & Key Results:
- D-RAC extends W-RAC by adding two format-agnostic stages: (i) deterministic PDF Normalization of any input document (DOCX, PPTX, XLSX, HTML, scanned images) into PDF, and (ii) Multimodal Markdown Conversion of rendered PDF pages using a multimodal LLM. [Figure 1]
- The conversion is "retrieval-aware": tables are rewritten as one self-contained prose sentence per row using column headers as context [Figure 2], decorative imagery is suppressed, and heading hierarchy is explicitly reconstructed.
- Subsequent stages use W-RAC's approach: deterministic parsing into ID-addressable units, followed by lightweight LLM-based chunk planning over identifiers, not text.
- Quantitative Results:
- The full 236-document, 795-page PDF corpus from the RAG-Multi-Corpus benchmark was converted and chunked in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks.
- D-RAC reduced chunking-stage output tokens by 95.7% compared to agentic chunking. [Table 9]
- This translated to a 77.8% reduction in chunking cost under GPT-4.1 pricing and 85.6% under Gemini 2.5 Pro pricing. [Table 10]
- Chunking time was reduced by 75%.
- In retrieval performance, D-RAC matched or exceeded agentic chunking across all seven overall metrics (e.g., Recall@6 improved from 0.717 (fixed-size) to 0.798 (D-RAC), and MRR from 0.602 to 0.690), with significant gains on temporal, comparative, and analytical queries. [Table 6, Table 8]
- D-RAC exhibits linear scalability to documents of 500+ pages, with conversion costing 3.9–5.5 s/page. [Table 4]
- Conclusion & Impact:
- D-RAC successfully extends retrieval-aware chunking to arbitrary enterprise documents by leveraging PDF normalization and a single retrieval-aware multimodal conversion pass.
- It addresses the critical limitations of traditional document ingestion methods by recovering crucial structural and semantic information from complex layouts.
- The methodology significantly reduces LLM inference costs and latency by decoupling expensive multimodal understanding from cheap, ID-based chunk planning.
- This research provides a unified, production-ready ingestion foundation for RAG systems, enabling efficient and accurate retrieval over diverse enterprise knowledge bases, even with complex PDF inputs, and facilitates rapid iteration on retrieval strategies at a lower cost.

Figure 1 — D-RAC 파이프라인

Figure 2 — 테이블 정규화 예시
Part 2: Importance Figure Information
Looking for the <img> tags and captions.
Figure 1:  -> https://arxiv.org/html/2609.24220v1/figures/pipeline.png (Overall architecture)
Figure 2:  -> https://arxiv.org/html/2609.24220v1/figures/tablenorm.png (Table normalization example)
Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, Table 8, Table 9, Table 10 are HTML tables, not images. So I will only select Figure 1 and Figure 2.
Double-check constraints:
- Technical Terminology in English: Yes, maintained.
- Natural Korean Phrasing with English terms: Yes.
- Mandatory Metadata: Authors and keywords.
- Output format: Part 1 markdown, Part 2 JSON after `
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] LLMs4All: A Review on Large Language Models for Research and Applications in Academic Disciplines
- [논문리뷰] ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
- [논문리뷰] VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
- [논문리뷰] TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
- [논문리뷰] Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors
Review 의 다른글
- 이전글 [논문리뷰] Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
- 현재글 : [논문리뷰] Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
- 다음글 [논문리뷰] EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation
댓글