Notes from the lab.
Releases, engineering write-ups, benchmarks, and stories from teams putting Docen to work.
Docen Parse 2.1
Faster pages, better equations, and improved reading order on dense layouts.
ReadManaged batch processing is now live
Point Docen at a queue of documents and let throughput scale itself.
Docen OCR 2
Our recognition model gets better at faint scans, handwriting, and dense multilingual pages.
Turbo extraction
A faster extraction path for high-volume, latency-sensitive workloads.
EU data residency for Docen
Options for keeping document processing within your chosen region.
Saturating the olmOCR benchmark
What it means when a benchmark stops being hard — and where we look next.
Build document processing pipelines with workflows
Chain parsing, layout, and extraction into a repeatable pipeline.
Introducing balanced extraction mode
A new mode that trades a little speed for higher recall on hard fields.
Automatically fill PDF forms with AI
Read a form's fields, map your data, and produce a completed PDF.
View your API requests in the playground
Inspect, replay, and share the exact requests you send to Docen.
Extracting tracked changes and metadata
Pull revision history and hidden metadata out of documents, not just the visible text.
How we benchmark and evaluate Docen
Our approach to measuring parsing and extraction quality — what we test, how we score it, and why we publish it.
Docen Eval for extraction
Bring evaluation to your extraction schemas, field by field.
Launch week: Docen Parse is faster again
Another round of speedups for the parsing model.
Launch week: spreadsheet parsing
Parse spreadsheets with empty cells and irregular headers, reliably.
Launch week: playground examples
A gallery of ready-made examples to explore what Docen can do.
Launch week: a new playground
Try any processor on your own documents, right in the browser.
Launch week: high-accuracy mode
A mode tuned for the documents where every character counts.
Launch week: faster tracked-changes outputs
Tracked-changes extraction, now noticeably faster.
Launch week: introducing Docen Layout for multi-page section hierarchy
Keeping section hierarchy intact across long, multi-page documents.
Launch week: document segmentation
Split long documents into clean, addressable sections automatically.
Launch week: the layout model
Region and table detection that holds together end to end.
Introducing Docen Eval
Score parsing and extraction against your own labeled documents before you ship.
The Docen SDKs
Official client libraries that drop Docen into the stack you already run.
How a presentation platform turned decks into structured data
A fast-growing presentation platform needed to read user-uploaded decks and documents reliably.
Structured extraction with citations
Every extracted value comes with a pointer to the exact span it came from.
Word bounding boxes and confidence
Every recognized word comes with a location and a confidence score. Here's how to use them.
Introducing Docen Extract
Schema-driven structured extraction with citations back to the source span.
Reducing hallucinations in document extraction
How citations and confidence keep extracted values grounded in the source.
Speeding up Docen Parse
How we cut median page latency while improving accuracy.
Structured extraction with the Docen API and long documents
Strategies for extracting from documents that don't fit in a single pass.
Turning exam papers into structured practice at scale
An exam-prep company used Docen to parse past papers, mark schemes, and diagrams into clean, structured content.
Docen Parse 2
A new generation of the parsing model, with sharper structure and stronger handwriting.
High-fidelity OCR drives accurate structured extraction
Extraction is only as good as the text underneath it. Recognition quality compounds.
Cracking math OCR
Recognizing equations and notation is its own hard problem. Here's how we approach it.
Reading purchase orders and invoices without manual entry
A procurement team replaced manual data entry with Docen and cut invoice turnaround from days to minutes.
Extracting hyperlinks from PDFs
Links are data too. Here's how Docen recovers them from PDFs.
Structured content for an AI learning company
An AI learning company needed clean, structured text from messy source material to power its tutoring models.
Parse PDFs just the way you want
Shape the output format and structure to fit what your systems expect.
Segmenting benefits and policy documents for a healthcare team
A healthcare team used Docen to split dense benefits documents into addressable, searchable sections.
Digitizing a county historical archive, page by page
A county records office turned a century of handwritten archives into searchable, structured text.
Docen Parse 1.5
A big step up in table accuracy and multilingual recognition.
Extracting part data from electronics datasheets
An electronics-sourcing platform used Docen to pull structured specs from thousands of component datasheets.
Free to start, pay as you go
A simpler way to price document intelligence: start free, then pay for what you process.
Launch week: Docen Parse 1.1
Better tables, faster pages, and steadier reading order across long documents.
Introducing Docen Parse
Our core parsing model turns any document into clean, ordered Markdown, HTML, and JSON.
Building better document intelligence for an AI-first world
Why accurate parsing and extraction are becoming core infrastructure — and how we're building models for the documents that matter.