LlamaIndex Extract v2.5 boosts accuracy with structural reasoning
LlamaIndex released Extract v2.5, improving document extraction accuracy across all tiers and adding advanced citations for better grounding.
LlamaIndex has launched Extract v2.5, a significant update to its schema-based document extraction agents that delivers higher accuracy and improved data grounding. Released on October 1, 2026, this version introduces architectural changes inspired by coding agents to handle complex document layouts more effectively without increasing per-page costs.
What happened
The core of this release is a measurable jump in performance across all three service tiers: Cost Effective, Agentic, and Agentic Plus. According to internal benchmarks on ExtractBench, the Cost Effective tier now achieves an overall value F1 score of 93.9, up from 87.1, surpassing the previous performance of the Agentic tier. The Agentic tier improved from 89.8 to 95.8, while the top-tier Agentic Plus rose from 95.1 to 96.4. These gains come with no change to pricing, offering better performance per dollar for existing users.
Beyond raw accuracy scores, the update addresses specific failure modes common in real-world document processing. Long lists, which often cause vision-language models to truncate output or lose track of repeated records, are now handled with intermediate representations that validate against the user’s schema. In one test involving a 17-page fund filing, the Cost Effective tier previously extracted only 87 of 238 holdings; with v2.5, it captured all 238, raising the document extraction score from 52.8 to 98.7.
The release also improves how the system handles records that span multiple pages and scanned forms containing mixed printed text, handwriting, and annotations. Previously, agents might merge reviewer notes with field values or stop reading at page breaks. The new version better preserves relationships between fields across pages and distinguishes original values from manual corrections, reducing the need for post-extraction cleanup.
How it works
Under the hood, LlamaIndex introduced a new agent harness purpose-built for document extraction. This architecture draws inspiration from recent advances in coding agents, hyper-tuning the model interactions to manage vision, reasoning, and verification steps more efficiently. A key feature is Structural Reasoning, which allows the agent to adapt its effort based on document complexity. For simple documents, it processes quickly with lower latency, while complex layouts trigger deeper analysis to ensure maximum accuracy.
Another major addition is Advanced Citations, now available for both Agentic and Agentic Plus tiers. This feature locates bounding boxes of supporting evidence for extracted values, even when the exact text does not appear word-for-word in the source. Grounding scores on ExtractBench jumped significantly, from 46.8 to 80.6 for Agentic and from 46.4 to 82.2 for Agentic Plus. These citations allow human reviewers to verify extracted data against the original document, facilitating reliable human-in-the-loop workflows.
The system also gained native spreadsheet extraction capabilities. Instead of flattening workbooks into unstructured text, agents now work directly with workbook cells, preserving the structured nature of the data. This approach helps maintain the integrity of tables and grids, which are often challenging for standard optical character recognition pipelines.
Key details
- Performance gains: Cost Effective F1 score rose to 93.9, Agentic to 95.8, and Agentic Plus to 96.4 on ExtractBench.
- No price increase: Per-page pricing remains unchanged despite the accuracy improvements across all tiers.
- Advanced Citations: Grounding scores improved to over 80 for top tiers, providing bounding boxes for evidence verification.
- Structural Reasoning: The agent adapts its processing effort based on document layout and information density.
- Native spreadsheet support: Agents now interact directly with workbook cells rather than flattened text representations.
- Improved long-list handling: Intermediate representations prevent truncation in documents with hundreds of repeated records.
Why it matters
For engineers building document automation pipelines, reliability is often the biggest bottleneck. Previous generations of extraction tools frequently required extensive manual review because they would miss items in long lists or confuse annotations with data. By improving grounding and structural reasoning, Extract v2.5 reduces the volume of exceptions that human operators must handle. This shift makes it feasible to automate high-stakes workflows, such as financial compliance or legal discovery, where missing a single line item can have serious consequences.
The introduction of confidence scores paired with citations also changes how teams design their review interfaces. Instead of blindly trusting the output or manually checking every field, developers can route only low-confidence extractions to human reviewers. The citations provide immediate context, allowing reviewers to jump directly to the relevant part of the source document. This targeted approach saves time and reduces cognitive load, making hybrid human-AI workflows more sustainable at scale.
What you can do
- Test v2.5 on your most difficult documents, particularly those with long lists or multi-page tables, using your existing schemas.
- Enable Advanced Citations in your configuration to generate bounding boxes for verifying extracted values.
- Use confidence scores to filter automatic processing and route only uncertain fields to human reviewers.
- Try native spreadsheet extraction for Excel-heavy workflows to preserve cell structure and relationships.
- Update your Python SDK to the latest version and use the
cite_sourcesandconfidence_scoresflags in your extraction jobs. - Experiment with the CLI tool
llpfor quick terminal-based extraction tests before integrating into larger applications.
