Sarvam Releases Document AI Model for English and 22 Indian Languages

Vision 2.1 adds ways to extract fields and read handwriting, but Sarvam’s own results show a wide accuracy gap between languages.

By 4 min read
Sarvam Releases Document AI Model for English and 22 Indian Languages
Sarvam Releases Document AI Model for English and 22 Indian Languages

Listen to this story

The audio brief

About 1:31
0:001:31
Read transcript
Sarvam’s new document model can pull fields from forms, organize tables that run across pages, and read handwriting in English and twenty-two Indian languages. The release, Vision 2.1, aims to turn documents into usable records—not just copy their text. Sarvam says it first identifies page regions and reading order, then has a vision-language model read them. The company reports that this improves accuracy, but hasn’t published a direct comparison of the two approaches. The results are promising in some tests, and uneven in others. Sarvam reports a score of 87.3 on olmOCR-Bench and 94.97 on OmniDocBench. On the latter, it trails PaddleOCR-VL 1.6, which scored 96.01. Those are company-reported benchmark results, not independent tests of the release. Sarvam’s new Indic OCR Bench covers 6,909 samples and measures character and word accuracy. Vision 2.1 scored 87.39 overall, but language results ranged from 97.41 for Konkani to 53.91 for Santhali. That spread is a reminder that coverage across twenty-two languages doesn’t mean equal performance. And the benchmark doesn’t test whether forms are extracted correctly, tables stay structured, or people have to fix the output. Sarvam also says serving costs are lower, but gives no new price. The key thing to watch is how the model performs on customers’ real documents—especially in lower-scoring languages and on extraction tasks its headline benchmark doesn’t measure.

Story brief

3 key points

Sarvam’s Vision 2.1 adds structured extraction from forms and multi-page tables, plus Indian-language handwriting recognition, extending its document model beyond text reading. Company-reported results are mixed: it scores 87.3 on olmOCR-Bench and 94.97 on OmniDocBench v1.6, behind PaddleOCR-VL 1.6 there. Its new Indic OCR Bench reports 87.39 overall across English and 22 Indian languages, but ranges from 97.41 for...

  1. 01

    Vision 2.1 segments page regions and determines reading order before a vision-language model reads; Sarvam reports an improvement but publishes no head-to-head test.

  2. 02

    Indic OCR Bench’s 6,909 samples measure character and word accuracy, not field extraction, table structure, or human correction burden.

  3. 03

    Extract API returns key-value pairs, tables, and form fields; Digitize structures text, while Akshar Document Agents support workflows with human review.

Sarvam has released Vision 2.1, a model designed to turn documents in English and 22 Indian languages into usable digital information. It can extract fields from forms, parse tables and recognize Indian-language handwriting, according to the company. Those additions target a harder job than simply copying text from a page: keeping the right information attached to the right field, even when a document spans pages or mixes formats.

A form needs more than a transcript

The original Sarvam Vision arrived in February with a focus on reading text from images and documents, understanding tables and producing structured output. Sarvam says users later flagged incorrect information and inconsistent results. For Vision 2.1, it says it worked on those weaknesses while adding extraction from forms and tables that stretch across multiple pages. The company has not published a before-and-after error rate for those fixes in its announcement.

The difference matters when a page is meant to become a record rather than a searchable image. Sarvam’s Extract API is designed to return key-value pairs, tables and form fields; its Digitize API converts documents into structured digital text. It also offers Akshar Document Agents for workflows with human review. These are access routes for the same release, not separate model launches.

How the model finds its place on a page

Vision 2.1 does not rely on one model reading an entire document unaided. Sarvam describes a layout parser that identifies regions of a page and a reading-order system that decides their sequence. A vision-language model then reads the content. Sarvam says this setup improves accuracy over asking the model to interpret a full page or document directly, though that improvement is a company claim rather than a published head-to-head result for the two approaches.

To teach the new extraction and handwriting tasks, the company says it combined real material with synthetic examples, including filled-in forms in several languages. It then used supervised fine-tuning, which trains on examples of desired outputs, followed by reinforcement learning that rewards checkable answers. Those choices address the kinds of content Vision 2.1 is meant to handle; they do not, by themselves, establish how reliably it will process a customer’s documents.

Strong totals, different language outcomes

In Sarvam’s evaluation, Vision 2.1 scored 87.3 on olmOCR-Bench, ahead of Infinity-Parser2 Pro at 86.1. It scored 94.97 on OmniDocBench v1.6, below PaddleOCR-VL 1.6 at 96.01. The tests measure different aspects of document reading: olmOCR-Bench uses checks across varied pages, while OmniDocBench measures how closely text, tables and formulas match the source. The rankings are Sarvam’s reported results, not an independent test of this release.

Sarvam also released Indic OCR Bench, a set of 6,909 samples: 6,609 across 22 Indian languages and 300 in English. It tests character and word accuracy using material that includes newspapers, textbooks and historical writing. Vision 2.1 scored 87.39 overall in Sarvam’s evaluation. That total is useful for comparison, but the language-level results are far from uniform.

Inside Sarvam’s Indic benchmark result
87.39Overall accuracy

Sarvam reports an overall Vision 2.1 score of 87.39 on its Indic OCR Bench.

97.41Konkani

Konkani received a score of 97.41 in Sarvam’s language-level results.

53.91Santhali

Santhali received a score of 53.91 in the same company evaluation.

In Sarvam’s results, Konkani scored 97.41 while Santhali scored 53.91. The gap makes “22 languages” a description of coverage, not a promise of equal accuracy. And this new benchmark is focused on recognizing words and characters: its overall score cannot stand in for a test of form-field extraction, table structure or whether a person must correct the finished record.

A cheaper service, without a listed new rate

Sarvam says changes to the systems that run Vision 2.1 let it serve the model more cheaply than the price announced for the original release. It also says the API has been optimized for production workloads. Its announcement does not give an exact new price, so buyers cannot calculate the savings from that statement alone. Cost and speed will need to be judged alongside errors on the particular forms, scripts and scans a team actually handles.

Sarvam itself cautions that established document benchmarks may be nearing saturation and questions how well high scores predict real-world usefulness. It says Vision 2.1 has undergone internal workflow testing. The next useful evidence would be results on representative documents outside the company’s evaluation, especially for languages where its own scores are lower and for tasks its Indic benchmark does not measure.

Sources

  1. sarvam.aiSarvam Vision 2.1 | Sarvam AI

Loading discussion...

YOUR READING SPACE

Notifications