ICD-10 coding with AI: how it works
How AI-assisted ICD-10-CM coding works step by step, why a general chatbot is not enough, benefits, compliance risks and how to evaluate an AI medical coding tool.
ICD-10 coding with AI uses language models to read clinical documentation, identify diagnoses and suggest candidate codes alongside the exact text that supports them. A coder validates every suggestion against the record, the Official Guidelines and payer rules. It speeds up work and improves consistency, but it does not replace the coder.
Editorial update: October 10, 2026. Published by Arkangel AI, a company that builds AI software for healthcare. Educational content; always confirm the code set version and guidelines that apply to each encounter.
What is ICD-10 (and ICD-10-CM)?
The International Classification of Diseases, 10th revision (ICD-10), is the World Health Organization standard for recording diagnoses and causes of death. In the United States, diagnoses are reported with ICD-10-CM, a clinical modification maintained by the CDC's National Center for Health Statistics, with more than 70,000 codes of up to seven characters. Inpatient procedures use ICD-10-PCS.
Two clarifications prevent common errors:
- ICD-10-CM is not the WHO's ICD-10. A tool trained on one version can suggest codes that do not exist in the other, which matters for organizations working across countries.
- ICD-11 has been in effect at the WHO since 2022, but the U.S. still uses ICD-10-CM; other countries are transitioning at different speeds.
How does AI medical coding work?
A well-designed workflow has six steps:
- Ingestion. The tool receives progress notes, discharge summaries, operative reports or results, as PDFs, text or structured data.
- Clinical extraction. A language model identifies diagnoses, procedures and findings, and distinguishes negations ("pneumonia ruled out"), history ("history of MI") and suspected conditions.
- Lookup in the official code set. Instead of "remembering" codes, the system retrieves candidates from the official tabular list for the configured version. This prevents non-existent codes.
- Rules. Specificity, principal diagnosis, combination codes and Excludes notes are applied according to the Official Guidelines and, where relevant, payer policy.
- Evidence per code. Every suggestion comes with the exact passage from the record that supports it.
- Human validation. The coder accepts, corrects or rejects; those decisions help improve the system.
Example
Note: "Patient with T2DM for 12 years, with diabetic nephropathy documented by persistent albuminuria. Denies tobacco use."
- A rushed review might code E11.9 (type 2 diabetes without complications).
- The AI highlights "diabetic nephropathy documented by persistent albuminuria" and suggests E11.21 (type 2 diabetes with diabetic nephropathy), citing the passage.
- "Denies tobacco use" generates no tobacco-use code.
- The coder confirms the code against the full record and guidelines.
Why isn't a general chatbot enough?
General models were not built to code. In a study in NEJM AI, the large language models tested matched the exact code from a code description in a minority of cases; the best, GPT-4, matched roughly a third of ICD-10-CM codes (Soroush et al., 2024). Errors included codes that do not exist.
Specialized systems combine the model with code-set retrieval, rules and memory of validated cases. In an Arkangel AI preprint under review at JMIR, a recursive learning architecture raised zero-shot clinical coding F1 from 0.318 to 0.605 after 20 iterations (JMIR Preprints). It is a preprint, not yet peer-reviewed, and was run by Arkangel's own team.
Benefits of AI-assisted coding
- Coverage: review every chart, not a sample.
- Specificity: catch documented complications, laterality or severity that are easy to miss, which also matters for risk adjustment and quality measures.
- Consistency: less variation between coders.
- Traceability: each code is tied to its evidence, which helps with audits and denials.
- Documentation gaps: flag conditions the data suggests but the note does not document, so a compliant query can go to the provider.
Specificity gaps AI commonly catches
ICD-10-CM rewards detail, and the detail is often in the note but not in the code. Typical examples:
- Heart failure type and acuity, such as acute on chronic diastolic heart failure instead of an unspecified heart failure code.
- Laterality for injuries, joints and eyes.
- Diabetes complications documented elsewhere in the chart, as in the example above.
- Pressure injury stage and site recorded in nursing notes.
- Acute versus chronic conditions, such as kidney disease stage.
Each of these should still be confirmed against the full record before it reaches a claim.
Risks and limits
- Non-existent or wrong-version codes if the system does not query the official code set.
- Negation and timing errors, such as coding a ruled-out diagnosis.
- Upcoding: suggesting codes the documentation does not support is a serious compliance risk. If it is not documented, it is not coded.
- Documentation quality: AI cannot extract what the clinician did not write.
- Privacy: only use tools contracted to process PHI under HIPAA.
How to evaluate an AI medical coding tool
- Test on a local sample of your own charts (with the necessary approvals), not just vendor examples.
- Measure per code and per encounter: precision, recall and F1, plus the share of suggestions coders accept unchanged.
- Check the version: ICD-10-CM fiscal year, ICD-10-PCS and any payer-specific edits.
- Require evidence per code and an audit trail of corrections.
- Time it: minutes per chart before and after, with the same team.
Where Arkangel AI fits
Arkangel AI surfaces candidate ICD-10 diagnoses and potential undocumented opportunities, each linked to its exact source text, so coders can validate them against the record and payer rules while final decisions stay human. On the AI for coders page you can try a demo that returns an annotated PDF; use synthetic or de-identified documents. For organizations, the platform is ISO 27001:2022 certified, with HIPAA-compliant Enterprise deployments (Trust Center). See also our guide to AI in healthcare.
Frequently asked questions
Can AI help with ICD-10 coding?
Yes. AI reads clinical documentation and suggests candidate ICD-10 codes alongside the text that supports them. It works best when it retrieves from the official code set and applies coding rules. A coder validates each suggestion against the record, the Official Guidelines and payer rules before the claim goes out.
Will AI replace medical coders?
No. AI speeds coding by suggesting ICD-10 and procedure codes with supporting text, but coders resolve ambiguity, apply guidelines and payer rules, query providers when documentation is missing and remain accountable. The role shifts from manual code lookup toward validation, quality control and documentation improvement.
What is the difference between ICD-10 and ICD-10-CM?
ICD-10 is the WHO's international classification. ICD-10-CM is the U.S. clinical modification, with more detailed codes of up to seven characters, used for diagnosis reporting and billing in the United States. Codes are not interchangeable, so tools must be configured for the version your organization reports.
Is AI medical coding HIPAA compliant?
It depends on the vendor and contract, not on AI itself. Use tools contracted to process PHI, with certified security controls, access logging and clear data-use terms. Arkangel AI offers HIPAA compliance and a signed BAA on Enterprise plans; do not paste identifiable records into consumer chatbots or into free or Pro accounts.
Sources
- WHO. International Classification of Diseases and ICD-10 browser (2019).
- CDC, National Center for Health Statistics. ICD-10-CM.
- Soroush A et al. Large Language Models Are Poor Medical Coders. NEJM AI 2024.
- Castaño-Villegas N et al. A Recursive Learning Architecture for Zero-Shot Automated Clinical Coding. JMIR Preprints 2026 (preprint; authors are Arkangel staff).