Docen
← BLOG·BENCHMARKS

Saturating the olmOCR benchmark

What it means when a benchmark stops being hard — and where we look next.

THE DOCEN TEAM·Jan 28, 2026·6 MIN READ

What it means when a benchmark stops being hard — and where we look next.

What we measure

We publish benchmarks because a document intelligence model is only as good as its behavior on documents you can't cherry-pick. We test recognition, tables, layout, and extraction against sensible baselines.

  • Character and word error rate on scanned and photographed pages.
  • Table-cell accuracy on complex, merged layouts.
  • Field-level F1 for schema extraction, with citations checked.
  • Latency per page at production settings.

Results

0.42%
character error rate
98.7%
table-cell accuracy
94.3%
extraction F1

Method

Every number here comes from documents held out of training. We report the settings, keep the evaluation reproducible, and update the figures as models change. When a result looks too good, we assume the test is wrong until we've checked it.

Want to see it on your own documents? Open the playground or reach out — we're happy to run a sample with you.

benchmarksOCRevaluation
[]TRY DOCEN

Run your hardest documentthrough Docen.

See the structured output for yourself, or reach the team at support@docen.co.