Introducing Turbo, our fastest extraction tier
LlamaIndex introduces the Turbo extraction tier, cutting document processing latency to 3.7s-0.5s per page by removing separate parsing and using parallel processing, specifically designed for real-time workflows.
- Turbo bypasses separate parsing, extracting directly from document pages in parallel for maximum speed
- On ExtractBench, Turbo achieves a median processing time of 3.7s per page (down to 0.5s for longer docs) with a Value F1 of 0.84
- Designed for real-time workflows like auto-filling forms, order entry, and agent-driven document processing
- Currently in Beta with fewer input types and configurations, ideal for latency-sensitive but less cost-sensitive use cases
Background: Why "Speed" Suddenly Became the Core Metric for Extraction?
Over the past two years, accuracy has been the top priority in document intelligence. After all, extracting a single wrong number from an invoice, contract, or purchase order can break downstream business logic. However, with the rise of agent workflows, extraction is no longer just a background batch job. It now sits directly in the user's waiting path or acts as a prerequisite for an agent's next tool call. In this context, speed is no longer a nice-to-have; it is the deciding factor for whether the entire system runs smoothly.
LlamaIndex's newly released Turbo extraction tier addresses this exact pain point. It is not merely a hardware upgrade but an architectural trade-off: it removes the separate parsing step used by other tiers and extracts directly from document pages in parallel. On the official ExtractBench, Turbo achieves a median processing time of just 3.7 seconds per page, dropping to 0.5 seconds for medium and long documents due to amortized fixed costs, while maintaining a Value F1 score of 0.84. This means that for many real-time interactive scenarios, extraction finally stops being the bottleneck.
Breakdown: What Makes Turbo So Fast?
Traditional document extraction pipelines typically follow a "parse first, extract later" approach. The parsing stage converts unstructured documents into machine-readable formats like text blocks or table structures. While necessary, this step is highly time-consuming. Turbo's core idea is to skip this step: it bypasses the independent parsing channel, runs extraction logic directly on the raw pages, and parallelizes page processing.
Think of it like ordering food at a restaurant. The traditional model is "waiter takes order, kitchen prepares, then serves," with each step waiting for the previous one to finish. Turbo is more like an "open kitchen" where chefs work directly on the ingredients, with multiple chefs handling different dishes simultaneously. The longer the document, the more pronounced this parallel advantage becomes, as the fixed overhead of initiating the task is spread across many pages. In contrast, other systems see latency spike sharply when processing documents over 16 pages, while Turbo's latency curve remains nearly flat.
Trend Insight: Real-Time Extraction Is Reshaping Agent Architecture
Behind this release lies a deeper trend: document processing is shifting from "background batch jobs" to "real-time interaction." Previously, we would send documents to the backend and get results hours or even days later. Now, when a user uploads an invoice, they expect the form to auto-fill within seconds; when an agent reads a contract, it needs structured data in seconds to decide its next action. Turbo's launch signals that extraction services are now specifically optimized for the real-time demands of "human-in-the-loop" and "agent-in-the-loop" workflows.
This also means future extraction services will likely feature more granular tiers, much like today's API gateways: some will prioritize extreme accuracy (for complex contracts), others will prioritize extreme speed (for real-time forms), and some will focus on cost-efficiency (for bulk archiving). Developers will need to dynamically select the optimal extraction strategy based on their business context, rather than relying on a one-size-fits-all approach.
Practical Value: How Should You Use It?
If you are building form auto-filling, purchase order entry, or agents that need to read documents in real time, Turbo is worth trying first. Its response speed allows users to complete an "upload-review-confirm" loop in a single session, or enables agents to extract data while a customer service rep drafts a reply. However, note that Turbo currently supports fewer input formats and configuration options, making it unsuitable for complex scenarios requiring highly customized parsing rules. If your task is latency-insensitive but cost-sensitive, the traditional Cost Effective tier remains the better choice.
Counterintuitive Insight: Speed Doesn't Necessarily Mean "Cutting Corners"
Many might worry that skipping the parsing step sacrifices accuracy. Yet, ExtractBench data shows Turbo's Value F1 reaches 0.84, with a relatively small gap compared to accuracy-focused tiers. This reminds us that in document intelligence, speed and accuracy are not always a zero-sum game. Through architectural optimization and parallel computing, we can compress latency below human-perceptible thresholds while maintaining reasonable accuracy. For most business scenarios, 0.84 accuracy with second-level response is far more valuable than 0.95 accuracy with a multi-minute wait.
Analysis by BitByAI · Read original