← Back to Home

Tag: 文档解析 (9 articles)

Building a Better LiteParse Skill with Evals

Through trace analysis and iterative evaluations, LlamaIndex optimized an agent's PDF parsing strategy, revealing a shift toward disciplined, data-driven agent engineering.

LlamaIndex Blog ·

Document OCR is Not Getting Commoditized

Benchmarks show specialized document OCR keeps beating top GPT models on accuracy and cost; document parsing won't be swallowed by frontier models.

LlamaIndex Blog ·

LiteParse v2.0 Runs Everywhere

LlamaIndex rewrote its lightweight PDF parser LiteParse in Rust, enabling cross-language and cross-platform (including browser) operation with up to 100x performance gains, providing critical infrastructure for real-time AI applications.

LlamaIndex Blog ·

LlamaIndex Newsletter 2026-04-14

LlamaIndex launches ParseBench, the first OCR benchmark for AI agents, and demonstrates breakthroughs in structured document understanding and multimodal reasoning, signaling a shift from text extraction to deep semantic comprehension.

LlamaIndex Blog ·

LlamaIndex Newsletter 5-19-26

LlamaIndex introduces ParseBench, the first OCR benchmark designed specifically for AI agents, alongside open-sourcing a local document parsing server and a secure sandboxed CLI agent, signaling a shift in document processing towards agent-native infrastructure.

LlamaIndex Blog ·

LlamaIndex Newsletter 7-8-26

LlamaIndex introduced Retrieval Harness and MCP restructure, enabling agents to actively traverse corpora with filesystem tools like list and grep, turning retrieval from guesswork into verification.

LlamaIndex Blog ·