Doc-V* agentic framework for multi-page Document VQA
August 21, 2026
Doc-V* is an OCR-free agentic framework that performs multi-page document reasoning through sequential evidence aggregation. The system uses a thumbnail overview followed by semantic retrieval and targeted page fetching, optimized via Group Relative Policy Optimization to improve accuracy and efficiency.
HOW THIS AFFECTS YOU
●
builderYou can implement agentic retrieval patterns to handle long, dense documents without relying on expensive OCR pipelines.
●
researcherThis offers a new benchmark for evaluating agentic reasoning in multimodal document tasks.