Cohere has unveiled Parse (parse-v5.0), a document parsing model engineered for high-volume enterprise ingestion. This 2.3-billion-parameter vision language model features an 8,192-token context window and a compact ~4.6GB footprint, built on Cohere Labs’ North-Micro-Vision-Instruct architecture. Parse accepts a PDF, PPT, or JPEG page as a base64-encoded data URI and outputs Markdown with text in reading order, tables rendered as HTML, lists, form key-value pairs, image descriptions, and bounding box coordinates—all without a separate OCR front-end.
Priced at $1.50 per 1,000 pages through the Parse API, Cohere positions the model on price-performance rather than peak accuracy. The company cites a self-reported ParseBench score of 79.2, though as noted below, this figure covers only three of the benchmark’s five dimensions.
Deployment Readiness
Yes—Parse is production-ready. It is generally available via the Cohere Parse API, Microsoft Foundry, AWS SageMaker, and single-tenant Model Vault. There is no waitlist or research license requirement, making it a practical option for immediate integration.
As of 2026, Parse aligns with industry trends toward lightweight, cost-efficient document AI, challenging larger models by offering a specialized solution that prioritizes throughput and affordability for enterprise workflows.
via MarkTechPost
