IRIS v2 Chunk Structure
IRIS v2 always returns achunks array in the parse result. chunk_mode: disabled is a Reducto-only concept (soon to be deprecated) that does not exist in IRIS v2.
Chunk count and size are driven by how IRIS segments each page, not by a separate chunking enum.
These options live on
IrisParseEngineOptions inside IrisParseJobParams. They are not the same as Dex’s four rechunk strategies below—they control IRIS’s native output only.
Parse with IRIS v2
Dex Chunking Strategies
Dex offers four chunking strategies that work on any parse result—including IRIS v2 via rechunking. Use these when you need consistent, configurable chunk boundaries across documents or for embeddings/RAGParse Once, Post-chunking as needed
With IRIS 2, you parse without Dex rechunking (omit chunking_options), then apply Dex strategies on the parse result. This is the recommended pattern: one parse, many chunking experiments.Async Rechunking
For long-running documents, start the rechunk job and poll for completion:Chunking Decision Tree
IRIS v2 native (IrisParseEngineOptions):
- Many small, layout-aware chunks →
layout="rt_detr_bce"(default) - One chunk per page →
layout="whole_page" - Full-page e2e model →
e2e_ocr+e2e_response_parser - Fewer regions (fewer chunks) → raise confidence thresholds or tighten containment filters
token_size when:
- Working with LLM APIs that have token limits
- Embedding models with specific token limits
- Cost optimization
recursive when:
- General document chunking for RAG
- Preserving paragraphs and sentences
- Articles, blog posts, documentation
by_page when:
- Legal documents, forms, reports
- Page references matter
- Page structure should be preserved
by_section when:
- Documents have clear section headers
- Technical manuals, wikis, academic papers
- Semantic coherence within topics
Reducto Chunking Parse-Time - (Legacy)
When using the Reducto parse engine (soon to be deprecated), you can chunk during the initial parse instead of rechunking. Reducto’s methods are layout-aware and use document structure.
Recommendation: Use
VARIABLE for most cases, especially with embeddings.
Pattern: Retry with Different Chunking
Reducto-only decision tree (Legacy)
Use ReductoVARIABLE when:
- Still on Reducto and want layout-aware chunking at parse time
- General document processing with parser-optimal chunk sizes
BLOCK when:
- Need precise bounding box information
- Building UI overlays on documents
- Not using for embeddings
DISABLED when:
- You plan to rechunk with Dex strategies (same pattern as IRIS v2)
Next Steps
- Vector Stores: Add chunks to vector stores for semantic search
- Extract: Extract structured data from parse results
- Parse: Parse engine options and configuration (IRIS v2)

