
AI Codes Alongside ED Experts: Results From
1,000 Emergency Department Charts
Emergency department coding puts autonomous medical coding under one of its toughest tests.
To evaluate whether Amy AI could perform at expert-human standards, CombineHealth and Edelberg & Associates' expert coders independently coded 1,000 ED charts, with Edelberg's audit team evaluating their performance.
CombineHealth Matched or Outperformed Expert Human Coders
A Deliberately Difficult Test: High-Complexity ED Environment
Emergency department coding compresses significant complexity into a high-volume workflow. Coders must interpret complete encounters, determine E/M levels, assign procedures and modifiers, select defensible ICD-10 diagnoses, evaluate medical necessity, and work through documentation that can support more than one compliant coding interpretation.
That complexity made the ED a rigorous environment for evaluating whether CombineHealth's autonomous medical coding platform, Amy AI, could perform at expert-human standards.
For 15 days, approximately 100 charts per day were assigned separately to CombineHealth and Edelberg's expert coding team, covering 1,000 emergency department charts in total.
The outputs were not judged simply by comparing CombineHealth's codes with the human coders' codes. Instead, Edelberg's audit team independently reviewed coding decisions against defined criteria, including documentation support, coding guidelines, specificity, sequencing, medical necessity, and applicable payer requirements.
This distinction was particularly important for ICD-10, where two different coding decisions can both be defensible when supported by the clinical documentation and official coding guidelines. The evaluation therefore measured whether each coding decision was supportable—not whether AI and human coders always produced identical outputs.
CombineHealth Matched Expert-Level Coding Performance
Across the evaluation, CombineHealth achieved 95.9% E/M accuracy, 99% CPT accuracy, 100% modifier accuracy, and 93.2% ICD-10 accuracy. Edelberg's expert human coding team recorded 94.7%, 99%, 100%, and 92.2%, respectively.
At the same time, approximately 85% of charts were coded autonomously, while approximately 15% were escalated when the platform detected ambiguity, conflicting documentation, or insufficient clinical support.
~5× More CDI-Triggering Documentation Gaps Identified
CombineHealth identified approximately 5× more CDI-triggering documentation issues than the traditional workflow.
These included missing critical care time, incomplete procedure details, missing independent interpretation of tests, missing documentation of external physician discussions, and insufficient medical necessity support.
This reflects how CombineHealth's automation platform approaches coding: analyzing the complete encounter, applying the medical coding methodology, evaluating documentation sufficiency and medical necessity, and incorporating payer-aware intelligence into billing-ready coding decisions.
Operational Impact From the Evaluation
CombineHealth Matched or Outperformed Expert Human Coders
Across every measured coding dimension in the 1,000-chart evaluation
Charts coded autonomously by CombineHealth's medical coding platform
Faster coding turnaround
~12 hours with the platform vs. ~24 hours with human coders
More CDI-triggering documentation gaps identified by CombineHealth
The evaluation was designed to determine whether autonomous medical coding could operate at expert-human standards in a complex, high-volume environment—not to prove that AI was better than human coders.
The results showed that CombineHealth could deliver expert-level coding performance while increasing autonomous execution, reducing turnaround time, and identifying substantially more documentation gaps.
Go Behind the Results
The full case study shows what happened inside the 1,000-chart evaluation: where CombineHealth and expert coders reached different decisions, how Edelberg's audit team determined what was defensible, and where the platform uncovered documentation gaps that traditional coding workflows missed.
See the detailed methodology, coding-level findings, documentation patterns, and operational takeaways behind the headline results.


