Article Bite

LightOnOCR-3 Adds Page Grounding to OCR—But the Pipeline Moved Upstream

LightOnOCR-3 puts transcription, layout boxes, image descriptions, and chart tables into one model, though benchmark formatting and upstream annotation still matter.

Primary source LightOn AI on Hugging Face ↗ Published 2026.10.08

A document parser often needs separate OCR, layout detection, and chart interpretation stages. LightOn’s October 8 release puts those outputs behind one model call. LightOnOCR-3 comes in three named variants—0.8B, 1B, and 4B—and adds grounding to ordinary page transcription. That could simplify an inference pipeline. It does not mean the system was built without a complicated pipeline of its own.

One page, two output modes

Send a page image with an empty text prompt and the model produces the familiar transcription. Send the exact prompt grounding and it also marks content blocks with labels and boxes: ![title](x1,y1,x2,y2), for example, uses page coordinates normalized to 0–1000. The same stream can carry image descriptions and chart values as HTML tables. The 4B model card says other instructions are outside the training distribution; this is not a general-purpose document question-answering interface.

That compact inline format matters when the output feeds search or chunking. It gives a downstream system text and a place on the original page, without wrapping every block in verbose JSON. But a chart table is generated content: an unreadable point or an estimated value is not a verified measurement. A consumer still has to decide which regions to trust and how to preserve their source coordinates.

A fictional company report page has colored boxes around its title, text, image, chart, and table; matching output blocks show labels, normalized coordinates, an image description, and HTML tables.
LightOn's publisher-hosted grounding illustration pairs a fictional page with the intended structured format. It is not an actual model prediction or a measured accuracy example. Open the original full-resolution diagram ↗

The benchmark depends on its wrapper

LightOn reports these overall scores across three separate document benchmarks. This is a selection from its release tables, not a complete leaderboard or an independent LiteBites evaluation. The model-name links identify publisher-hosted checkpoints; the numbers come from LightOn’s release, not those model cards.

LightOn-reported overall scores (higher is better within each column; different benchmarks have different tasks).
Model ↗olmOCR-BenchParseBench (5 cats)fr-bench-pdf2md
LightOnOCR-3-4B ↗86.375.174.1
LightOnOCR-3-0.8B ↗85.574.670.5
LightOnOCR-3-1B ↗84.571.469.6
Infinity Parser Pro ↗87.674.363.2
Chandra 2 ↗85.870.169.0

Selected rows in the October 8 release tables. The blog calls the linked Infinity-Parser2-Pro checkpoint “Infinity Parser Pro.” Scores and processing setups should not be treated as controlled model-only comparisons; scroll the table horizontally on small screens. The older LightOnOCR-2 row is excluded because its olmOCR overall omits one category.

The release itself warns that edit-distance scoring is sensitive to formatting and applies normalization before evaluation. There is also a visible snapshot difference: the author’s October 5 benchmark repository lists 86.1 for the 4B model under a named postprocessing pipeline, rather than the blog’s 86.3. Its comparison rows differ too. The repository pins scripts and revisions, which helps reproduction, but these two sets of numbers should not be combined into one leaderboard without reconciling their pipelines.

Speed has a similar boundary. LightOn reports tests on the same 512 pages, with one H100 per model and vLLM 0.30.0. Its 0.8B and 4B models achieve their best reported olmOCR scores at a 400-DPI, 5-megapixel cap; a smaller 1,540-pixel rendering raises peak throughput but changes the input. Faster serving at lower resolution is a quality–throughput choice, not a free acceleration at identical page detail.

The pipeline moved into the training data

LightOn describes using its older OCR model plus PaddleOCR, Docling, document-layout detectors, and text-to-box alignment to construct grounding annotations. It audited ambiguous or duplicated regions rather than accepting every detected box. A much larger Qwen vision-language model supplied candidate image descriptions and chart tables; format and plausibility checks rejected some of those candidates. Unprinted chart values could be estimated from axes during annotation, not simply transcribed from the page.

That is the useful distinction: one-model inference is not one-model supervision. The release and Apache-2.0 model cards make the resulting checkpoints available, but they do not establish error rates for every document type or guarantee that chart cells match their source pixels. The 4B product name is also not an exact parameter audit: Hugging Face’s card currently displays about 5B parameters for that checkpoint.

Before replacing a parser

  • Test plain transcription and grounding separately on your own scans, forms, tables, and multi-column pages.
  • Inspect boxes and chart cells against the page image; do not treat a plausible HTML table as verified data.
  • Compare benchmark rows only with the same categories, formatting rules, render resolution, and postprocessing.
  • Budget for image resolution, output tokens, and the whole document workflow—not only model decoding.

Sources