Saturating the olmOCR benchmark
What it means when a benchmark stops being hard — and where we look next.
What it means when a benchmark stops being hard — and where we look next.
What we measure
We publish benchmarks because a document intelligence model is only as good as its behavior on documents you can't cherry-pick. We test recognition, tables, layout, and extraction against sensible baselines.
- Character and word error rate on scanned and photographed pages.
- Table-cell accuracy on complex, merged layouts.
- Field-level F1 for schema extraction, with citations checked.
- Latency per page at production settings.
Results
Method
Every number here comes from documents held out of training. We report the settings, keep the evaluation reproducible, and update the figures as models change. When a result looks too good, we assume the test is wrong until we've checked it.
Want to see it on your own documents? Open the playground or reach out — we're happy to run a sample with you.