본문으로 건너뛰기

[논문리뷰] Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion

링크: 논문 PDF로 바로 열기

Total of 1.

The paper is about Document Retrieval-Aware Chunking (D-RAC). I have the content. Now I need to structure the summary.

Part 1: Markdown Summary

  • Authors: Uday Allu, Abhivanth Sivaprakash, Pratik Singh, Aman Manocha
  • Keywords: Need to select 5-8 keywords based on the paper content. Good candidates: RAG, Document Ingestion, PDF Normalization, Multimodal LLM, Retrieval-Aware Chunking, Markdown Conversion, Cost Efficiency, Enterprise Documents.
  • Key Terms & Definitions:
    • Retrieval-Augmented Generation (RAG): The dominant paradigm for grounding large language models in enterprise knowledge.
    • Document Retrieval-Aware Chunking (D-RAC): An extension of W-RAC for arbitrary document formats, leveraging PDF normalization and multimodal Markdown conversion for retrieval-optimized ingestion.
    • PDF Normalization: The process of converting any input document format (DOCX, PPTX, XLSX, scanned images) into a standardized PDF representation.
    • Multimodal Markdown Conversion: A single pass by a multimodal LLM to convert rendered PDF pages into retrieval-optimized Markdown, preserving hierarchy and converting tables into prose.
    • Agentic Chunking: Applying an LLM to raw extracted text to produce semantically coherent chunks, often involving text regeneration.
  • Motivation & Problem Statement:
    • RAG systems struggle with heterogeneous enterprise document formats (PDFs, Word, presentations, scans) due to complex visual layouts, multi-column pages, and dense tables.
    • Traditional ingestion pipelines (rule-based text extraction, OCR) often destroy reading order, flatten tables, and lose heading hierarchy, degrading retrieval quality. [Table 1]
    • Existing agentic chunking methods over raw text incur high token costs and hallucination risks due to text regeneration.
    • W-RAC (Web Retrieval-Aware Chunking) works well for structured web content (HTML), but PDF's presentation-centric nature lacks the reliable structure W-RAC depends on.
  • Method & Key Results:
    • D-RAC extends W-RAC by adding two format-agnostic stages: (i) deterministic PDF Normalization of any input document (DOCX, PPTX, XLSX, HTML, scanned images) into PDF, and (ii) Multimodal Markdown Conversion of rendered PDF pages using a multimodal LLM. [Figure 1]
    • The conversion is "retrieval-aware": tables are rewritten as one self-contained prose sentence per row using column headers as context [Figure 2], decorative imagery is suppressed, and heading hierarchy is explicitly reconstructed.
    • Subsequent stages use W-RAC's approach: deterministic parsing into ID-addressable units, followed by lightweight LLM-based chunk planning over identifiers, not text.
    • Quantitative Results:
      • The full 236-document, 795-page PDF corpus from the RAG-Multi-Corpus benchmark was converted and chunked in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks.
      • D-RAC reduced chunking-stage output tokens by 95.7% compared to agentic chunking. [Table 9]
      • This translated to a 77.8% reduction in chunking cost under GPT-4.1 pricing and 85.6% under Gemini 2.5 Pro pricing. [Table 10]
      • Chunking time was reduced by 75%.
      • In retrieval performance, D-RAC matched or exceeded agentic chunking across all seven overall metrics (e.g., Recall@6 improved from 0.717 (fixed-size) to 0.798 (D-RAC), and MRR from 0.602 to 0.690), with significant gains on temporal, comparative, and analytical queries. [Table 6, Table 8]
      • D-RAC exhibits linear scalability to documents of 500+ pages, with conversion costing 3.9–5.5 s/page. [Table 4]
  • Conclusion & Impact:
    • D-RAC successfully extends retrieval-aware chunking to arbitrary enterprise documents by leveraging PDF normalization and a single retrieval-aware multimodal conversion pass.
    • It addresses the critical limitations of traditional document ingestion methods by recovering crucial structural and semantic information from complex layouts.
    • The methodology significantly reduces LLM inference costs and latency by decoupling expensive multimodal understanding from cheap, ID-based chunk planning.
    • This research provides a unified, production-ready ingestion foundation for RAG systems, enabling efficient and accurate retrieval over diverse enterprise knowledge bases, even with complex PDF inputs, and facilitates rapid iteration on retrieval strategies at a lower cost.

Figure 1: D-RAC 파이프라인

Figure 1 — D-RAC 파이프라인

Figure 2: 테이블 정규화 예시

Figure 2 — 테이블 정규화 예시

Part 2: Importance Figure Information

Looking for the <img> tags and captions. Figure 1: ![Refer to caption](2609.24220v1/figures/pipeline.png) -> https://arxiv.org/html/2609.24220v1/figures/pipeline.png (Overall architecture) Figure 2: ![Refer to caption](2609.24220v1/figures/tablenorm.png) -> https://arxiv.org/html/2609.24220v1/figures/tablenorm.png (Table normalization example) Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, Table 8, Table 9, Table 10 are HTML tables, not images. So I will only select Figure 1 and Figure 2.

Double-check constraints:

  • Technical Terminology in English: Yes, maintained.
  • Natural Korean Phrasing with English terms: Yes.
  • Mandatory Metadata: Authors and keywords.
  • Output format: Part 1 markdown, Part 2 JSON after `

⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.

댓글

관련 포스트

Review 의 다른글