Factored research introducing MEASURE, an award-winning framework for evaluating structured understanding in resume extraction.

Beyond OCR Accuracy

Award-winning CVPR research benchmarks 14 AI systems to expose where resume extraction fails beyond OCR accuracy.

Key Takeaways:

System inspection icon
OCR accuracy doesn't prove understanding.
Model evaluation icon
14 AI systems expose different failure modes.
Production reliability icon
Production reliability requires system-level evaluation.

MEASURE: Multi-stage Evaluation for Assessing Structured Understanding in Resume Extraction

Resume parsing looks straightforward: read a document and convert its contents into structured data.

Production systems face a much harder problem. Resumes combine complex layouts, columns, sections, dates, job titles, skills, education, and relationships between information. A model can correctly recognize every word on a page and still assign that information to the wrong field, section, or sequence.

Yet document AI benchmarks frequently evaluate individual tasks rather than whether the complete system correctly understands the document. Factored R&D team Elizabeth Granda Rodriguez and Rafael Mosquera-Gómez developed MEASURE to close that gap.
‍

One Framework Measures The Entire Document Pipeline.

MEASURE introduces a multi-stage evaluation framework designed around how resume extraction systems actually operate. Instead of reducing performance to one accuracy metric, the framework evaluates four distinct dimensions:

Layout Analysis

Can the model identify and segment the document's structure?

Reading Order

Can it reconstruct information in the correct sequence?

Text Extraction

Can it accurately recognize the content?

Semantic Understanding

Can it correctly interpret and classify the information it extracts?

Together, those stages create a much more demanding test of document intelligence. The question isn't simply whether the model can read a resume. It's whether it understands its structure well enough to turn it into reliable data.
‍

112 Resumes. 218 Images. 12,000+ Annotations.

Testing that question required a dataset built around real document complexity.

Factored’s R&D team assembled 112 real-world resumes converted into 218 images, spanning Machine Learning, Data Science, Software Engineering, and Data Engineering roles. The dataset includes documents in English and Spanish and more than 12,000 annotated bounding boxes and transcriptions.

Annotations capture both item-level entities, including names, dates, and roles, and section-level entities such as education, skills, and experience. This allowed the R&D team to evaluate document understanding at multiple levels rather than relying on synthetic or isolated benchmarks.
‍

14 AI Systems Prove There Is No Single Winner.

MEASURE was used to benchmark 14 AI systems across the evaluation pipeline.

The result was not one universally superior model. Different systems excelled at different stages, revealing a fundamental problem with evaluating document AI through a single metric.
‍

  • Strong OCR did not guarantee structural correctness.
  • High semantic similarity did not guarantee information landed in the correct field.
  • Strong performance at one stage did not guarantee end-to-end reliability.
    ‍

That means choosing a document AI system based on one benchmark can hide failure modes that only become visible once the system encounters real documents.
‍

The Best Model Depends On What You Need It To Get Right.

MEASURE pairs each stage with metrics designed for that specific function.

Layout analysis

Uses Mean Average Precision at different intersection-over-union thresholds.

Text extraction

Measures Word Error Rate and Character Error Rate.

Reading order

Uses BLEU and Normalized Levenshtein Distance.

Semantic understanding

Is evaluated using BERTScore.

Evaluating these dimensions independently makes it possible to locate where a system fails rather than hiding that failure inside an aggregate score. That's particularly important for production pipelines, where different errors have very different consequences.
‍

Production AI Needs More Than A Good Benchmark Score.

MEASURE points toward a broader lesson for enterprise AI.A model can look strong on a conventional benchmark while still failing the workflow it was deployed to support.

For resume extraction, misplaced employment dates, incorrectly associated job titles, broken reading order, or misclassified sections can propagate directly into downstream recruiting systems and decisions. The same principle extends beyond resumes.

Production AI requires evaluation of the complete system behavior—not just isolated model accuracy. MEASURE provides a framework for making those failures measurable before they become production problems.


Best Short Paper — CVPR 2026
‍


Research Rigor Recognized At CVPR 2026.

MEASURE received the Best Short Paper Award at AI4RWC: The 2nd International Workshop on Vision Intelligence for Real-world Challenges, held in conjunction with CVPR 2026.

The award recognizes Factored's research into a practical problem at the intersection of computer vision, document intelligence, and production AI: determining whether systems actually understand complex documents—not simply whether they can extract their text.
‍

Best Short Paper Award
AI4RWC @ CVPR 2026
Denver, Colorado · June 3, 2026

Forward Deployed Engineering.
In Your Environment.
In Your Time Zone.

Inside your standups, architecture decisions, workflows, and production environments
We build on your stack, for your business
1,000’s of AI & Data engagements across complex production environments
Discuss Your Challenge
Start Building

Continue Reading

Structure Beats Random Scale
Language families improve transfer
The Cost of Speech Adaptation
Content gains can erase speaker signal
Speech in the Wild
800K+ hours. 73 languages. 3 tasks.