Raleigh, N.C. – September 15, 2026 -- Nutrient has made its Data Extraction API generally available, targeting enterprises that need traceable, audit-ready structured data from PDFs, scans, images, and Office files to power AI agents and RAG systems.
Nutrient ties every extracted field to a bounding box and page reference
The extract endpoint maps documents to a customer-defined or auto-generated JSON Schema and returns each field with a page reference, source blocks, a confidence signal, and a grounding label. Fields marked "fuzzy_match" or "not_found" route automatically to human review, while high-confidence matches proceed directly into downstream workflows.
CEO says unverified AI output cannot satisfy auditors
"Most agents can pull data out of a document. Almost none of them can prove where the data came from," said Jonathan Rhyne, co-founder and CEO of Nutrient. "Every field we return points back to the exact spot in the source, and when we can't find something, we say so instead of guessing. In production, 'Trust us' isn't an answer enterprises can take to their auditors."
Understand mode scored 0.932 accuracy on a public 200-document benchmark
Nutrient's understand mode achieved 0.932 overall accuracy on the publicly available opendataloader-bench corpus, covering reading order, table structure, and heading hierarchy. Text and structure modes are scored on the public leaderboard; understand and agentic modes were evaluated internally against the same corpus, with the original comparison run on July 6, 2026. Nutrient separately published two open grounding artifacts on Hugging Face: a grounding-en evidence-scoring model under Apache-2.0, and its evaluation dataset under CC-BY-SA-4.0.
Platform processes over 100 OCR languages across four modes
The API parses documents into spatial JSON or Markdown and detects tables, forms, formulas, charts, handwriting, and checkboxes across more than 100 OCR languages. Four processing modes -- text, structure, understand, and agentic -- let teams scale compute and cost only where deeper document understanding is required.
New accounts get 5,000 free credits monthly, no card required
Access is available via REST API, a free tier, and Nutrient Data Extraction Studio for visual testing without writing code. All API communication uses HTTPS/TLS encryption, and the platform is SOC 2 Type 2 audited, with reports available under NDA; processing-run retention varies by plan, and documents are not retained for model training on plans without internal data retention.
Interactive demos target healthcare, mortgage, and legal document types
Published demos show table-row extraction from CMS-1500 healthcare claim forms with dropout-red ink grids, signature detection on mixed handwritten birth records, value isolation in dense three-column mortgage financial tables, and multi-sentence narrative extraction from legal filings spanning three pages.