The CLAIRE Blog
Is ChatGPT Accurate for Medical Coding? Findings
What the peer-reviewed benchmarks actually measured: 33.9% exact-match on ICD-10-CM and 49.8% on CPT in the Mount Sinai study, 14-22% in otology, fabricated codes across every model tested, and why laterality, modifiers and rare codes are where general-purpose LLMs break.
Courtney HoatsonRHIA, CDIP, CCS, CRC, LSSGBTable of Contents
- Empirical Findings and Accuracy Benchmarks: What the Studies Show
- Key Limitations and Causes of Inaccuracy: Why ChatGPT Falls Short
- Compliance, Privacy, and Regulatory Considerations: The Legal Risks
- Recommended Role and Best Practices: How to Use ChatGPT Safely in Medical Coding
- Generalist AI vs. Specialized Medical Coding Solutions: Why Domain-Specific Tools Win
- Future Outlook: Advancements in LLMs for Medical Coding
- Frequently Asked Questions
- Key Takeaways
- Conclusion: The Bottom Line on ChatGPT for Medical Coding
Is ChatGPT accurate for medical coding? The short answer is no, not for autonomous code assignment. Multiple peer-reviewed studies published in NEJM AI, JAAOS, Annals of Otology, and MDPI found ChatGPT accuracy insufficient for autonomous coding, with several explicitly recommending mandatory human oversight while others note a need for further refinement.
GPT-4's best exact-match accuracy in the Mount Sinai benchmark was 33.9% for ICD-10-CM and 49.8% for CPT, far below what compliant billing demands. ChatGPT can assist with research and drafting but cannot serve as a final coder.
Specialized, code-set-grounded tools like CLAIRE's AI-powered medical coding assistant offer higher accuracy and clinical-reasoning explanations. This article reviews the empirical data, error patterns, HIPAA risks, best practices, and when to choose specialized tools over general-purpose LLMs.
Empirical Findings and Accuracy Benchmarks: What the Studies Show
The Mount Sinai benchmark, published in NEJM AI, evaluated 7,697 ICD-9-CM, 15,950 ICD-10-CM, and 3,673 CPT codes from 12 months of EHR data. GPT-4 achieved exact-match accuracy of 45.9% for ICD-9-CM, 33.9% for ICD-10-CM, and 49.8% for CPT. All tested models, including GPT-3.5, GPT-4, Gemini Pro, and Llama2-70b, "often generated codes conveying imprecise or fabricated information." The study introduced CodeSTS, a manual similarity grading system that classifies non-exact codes as "equivalent," "generalized but correct," or incorrect. GPT-4 generated most "equivalent" codes, while GPT-3.5 generated most "generalized but correct" codes.
An MDPI study of 150 randomized cases found overall accuracy of approximately 31% in both English and Portuguese. Procedure-code error rates were roughly four times higher than diagnostic codes, with quantified error rates reaching 62.8 to 65.6% for procedure codes; excess or unrelated codes were also among the most frequent error types. A JAAOS hand-surgery study of 90 patients across three surgeons showed CPT accuracy of 91.5% but ICD-10 of only 23.9%. Laterality errors appeared as a persistent failure mode, though they were not ranked as the single most common error across these studies.
An otology billing study of 191 operative notes found ChatGPT-3.5 achieved 22% exact match and 32% partial match, while ChatGPT-4 reached 14% exact and 33% partial. Cochlear implantation sensitivity was 94 to 96%, but cartilage grafting was 0 to 4.2%. Both models "failed to apply modifiers," causing underbilling by assigning fewer wRVUs (work Relative Value Units) than human coders.
A MIMIC-IV/MS-DRG study, available through PubMed Central (PMC), achieved top-1 accuracy of 68.1% and top-5 of 90.0% on the 30 most frequent DRG codes, but only with a specialized meta-prompt, many-shot learning, and dynamic in-context learning framework. Cross-study comparisons should be cautious: benchmarks focus on frequent codes, and rare-code performance is likely worse.
Key Limitations and Causes of Inaccuracy: Why ChatGPT Falls Short
Code hallucination is the most dangerous failure mode. LLMs generate fabricated or non-existent codes that sound plausible but are not valid in the current code set. The NEJM AI study found all tested models, including GPT-3.5, GPT-4, Gemini Pro, and Llama2-70b, "often generating codes conveying imprecise or fabricated information."
Laterality errors remain a persistent failure mode. Targeted prompts emphasizing laterality improved ICD-10 accuracy from 23.9% to 40% in hand surgery, still far below acceptable thresholds. Modifier omissions create direct reimbursement risk: the otology study found both models "failed to apply modifiers," assigning fewer wRVUs than human coders. The MDPI study found procedure-code error rates roughly four times higher than diagnostic codes, with excess or unrelated codes among the most frequent errors.
Rare and complex codes degrade performance. Code frequency and shorter descriptions correlated with higher exact-match rates (P<0.05). ChatGPT also lacks clinical context: it cannot interpret ambiguous documentation, query physicians, or apply AHA Coding Clinic guidelines, and it has no real-time access to current code tables. A searchable medical code reference built into a workflow solves this by grounding every suggestion in live code tables.
Compliance, Privacy, and Regulatory Considerations: The Legal Risks
Entering protected health information into consumer ChatGPT is not covered by a Business Associate Agreement (BAA). Organizations must use HIPAA-compliant API tiers or de-identify all data. As industry analysis confirms, ChatGPT is not automatically HIPAA-compliant.
Inaccurate coding, especially systematic underbilling, can trigger CMS Recovery Audit Contractor (RAC) reviews, OIG audits, and False Claims Act exposure. The HHS OIG routinely audits AI-assisted coding tools under its Work Plan.
ChatGPT does not reference AHA Coding Clinic or AMA CPT Assistant, both authoritative sources auditors use. AHIMA and ACDIS require that AI-generated codes be reviewed by credentialed professionals (CCS, CDIP, CPC) before submission. AHIMA's Code of Ethics obligates coders to ensure documentation integrity. CPT codes are copyrighted by the AMA and require proper licensing.
Recommended Role and Best Practices: How to Use ChatGPT Safely in Medical Coding
Human-in-the-loop oversight is non-negotiable. Multiple published studies, including research in NEJM AI, JAAOS, and the Annals of Otology, found ChatGPT accuracy insufficient for autonomous coding, and several explicitly recommend mandatory human oversight. Frame ChatGPT as a research and drafting assistant, not a final coder.
Structured prompting helps but has limits. Targeted prompts emphasizing laterality improved ICD-10 accuracy from 23.9% to 40% in hand surgery. Basic prompt styles, including zero-shot, one-shot, multishot, and chain-of-thought, showed no statistically significant difference (P=0.27 for ICD-10, P=0.62 for CPT). Advanced domain-specific frameworks are needed.
Retrieval-Augmented Generation (RAG) grounds LLMs in current code sets, reducing hallucination by constraining outputs to valid codes. This is a core architectural advantage of specialized tools over general-purpose ChatGPT. Every AI-suggested code must be verified against the official code set, AHA Coding Clinic guidance, and payer-specific rules before submission.
Safe use cases include drafting CDI queries, looking up clinical concepts, generating education scenarios, and summarizing documentation, always with human review. CLAIRE's AI Medical Coding Assistant applies a coder-built approach with clinical-reasoning explanations for each code, and its compliant CDI query generation tool creates physician-friendly queries that address the gaps where ChatGPT falls short.
Generalist AI vs. Specialized Medical Coding Solutions: Why Domain-Specific Tools Win
General-purpose LLMs like ChatGPT, Gemini, and Llama are trained on broad internet data, not on ICD-10-CM, CPT, or HCPCS code tables, AHA Coding Clinic updates, or CMS National or Local Coverage Determinations. This is why hallucination and outdated-code rates remain high.
Specialized coding AI tools offer a different architecture: trained or fine-tuned on actual code sets, grounded via RAG to current code tables, designed to explain clinical reasoning behind each suggestion, and built with compliance guardrails for HIPAA and CMS audit alignment.
CLAIRE was built by coders for coders, embedding domain expertise that general-purpose LLMs lack. Its integrated database for ICD-10-CM, CPT, HCPCS, and ICD-10-PCS codes reduces the need to switch between separate code-reference tools, and every suggested code includes a clinical-reasoning explanation so coders can verify the logic, not just accept a black-box output.
No published study has evaluated domain-specific coding vendors head-to-head with ChatGPT, a gap in the literature. However, the DRG study showed specialized prompt frameworks lifted GPT to 68.1% top-1 and 90.0% top-5 accuracy, suggesting both fine-tuning and RAG are viable paths.
Future Outlook: Advancements in LLMs for Medical Coding
GPT-4 generally outperformed GPT-3.5 on exact-match accuracy in the Mount Sinai and hand-surgery studies, though GPT-3.5 occasionally led on "generalized but correct" codes and one otology metric. Future models may close the gap, but general-purpose training alone will not solve code-set grounding and hallucination.
Fine-tuning on coding scenarios, AHA Coding Clinic advice, and CMS guidelines could push accuracy beyond zero-shot prompting. RAG is the most promising near-term approach, grounding LLMs in real-time access to current ICD-10-CM, CPT, HCPCS, and ICD-10-PCS code tables.
GPT models serve as helpful assistants to human coding specialists but are not yet equipped to fully replace expert judgment.
No published study recommends fully autonomous AI coding today. A realistic timeline for narrow, high-frequency code categories is 3 to 5 years with RAG and fine-tuning; complex, multi-code encounters will take longer. Academic medical centers and NIH-funded studies continue to evaluate LLM coding accuracy.
Frequently Asked Questions
Is ChatGPT accurate for medical coding?
No. Multiple peer-reviewed studies report exact-match accuracy ranging from 14% to 49.8% across code sets and specialties, far below the near-perfect rates required for compliant billing. The Mount Sinai benchmark published in NEJM AI remains the largest evaluation to date. Researchers across studies consistently recommend that credentialed coders validate all AI-suggested codes before submission, and several explicitly call for mandatory human oversight.
Can ChatGPT assign ICD-10 codes accurately?
Not reliably. ICD-10-CM exact-match accuracy ranges from 23.9% in hand surgery to 33.9% in large benchmarks. Common errors include omitted laterality, incorrect specificity, and fabricated codes. Targeted prompts improved accuracy from 23.9% to 40% in one study, still far below thresholds for compliant billing.
Can ChatGPT assign CPT codes accurately?
CPT accuracy ranges from 49.8% exact match in the Mount Sinai benchmark to 91.5% in hand surgery. Otology studies found only 14% to 22% exact matches. ChatGPT also frequently failed to apply modifiers, tending to underbill by assigning fewer work RVUs than human coders.
Is ChatGPT HIPAA-compliant for medical coding?
Standard consumer ChatGPT is not HIPAA-compliant. Entering protected health information into consumer AI tools risks HIPAA violations without a Business Associate Agreement. Organizations must use HIPAA-compliant API tiers or de-identify all data. CMS and OIG also scrutinize AI-assisted coding for audit risks under the OIG Work Plan.
Can ChatGPT replace human medical coders?
No. Multiple published studies, including research in NEJM AI, JAAOS, and the Annals of Otology, found ChatGPT accuracy insufficient for independent medical coding. Several explicitly recommend human-in-the-loop oversight, and credentialed professionals (CCS, CDIP, CPC) must validate all code assignments before submission.
Key Takeaways
- ChatGPT is not accurate enough for autonomous coding; even the largest benchmark found GPT-4's best exact-match rates fell well below 50% for both ICD-10-CM and CPT.
- Multiple peer-reviewed studies found ChatGPT accuracy insufficient for autonomous coding, with several explicitly recommending mandatory human oversight before any AI-suggested code is finalized.
- Code hallucination, laterality errors, modifier omissions, and poor rare-code performance are common failure modes.
- HIPAA, CMS audit exposure, and AHIMA ethical obligations make unvalidated AI coding a liability.
- Specialized, code-set-grounded AI tools built for coders offer higher accuracy and compliance guardrails.
Conclusion: The Bottom Line on ChatGPT for Medical Coding
Is ChatGPT accurate for medical coding? The accumulated evidence says no. The largest benchmark found GPT-4's best exact-match rates at 33.9% for ICD-10-CM and 49.8% for CPT, and the gap between those results and the near-perfect accuracy that compliant billing demands remains wide. Multiple peer-reviewed studies recommend human expert review before any code is finalized.
ChatGPT can assist with research and drafting, but HIPAA, CMS, and AHIMA obligations make unvalidated AI coding a liability. Specialized, code-set-grounded tools built by coders offer a safer alternative.
Ready to move beyond general-purpose AI? Explore CLAIRE's AI coding assistant to see how coder-built tools deliver the accuracy and compliance your workflow demands.
Reviewed by a CCS, CDIP-credentialed coding specialist with experience in CDI and AI-assisted coding workflows.
Related Posts
The Future of Medical Coding: Trends and Predictions for 2026-2030
The future of medical coding centers on deeper human-AI collaboration, with the AI medical coding market projected to grow from $2.99 billion in 2025 to over $10 billion by 2035. Key trends include autonomous coding for routine cases, predictive analytics for documentation improvement, point-of-care coding guidance, and specialized AI for complex specialties. Human coders will focus on complex cases, quality assurance, and clinical judgment while AI handles routine processing with increasing sophistication.
Read moreBest CDI Software Tools for Hospitals: 2026 Comparison
A fair comparison of Iodine, Solventum, Nuance, SmarterDx, Ambience, Optum, Dolbey and CLAIRE across AI capability, EHR integration, pricing and KLAS scores, grounded in the ACDIS 2026 survey and the new AHIMA/ACDIS query practice guidelines.
Read moreMedical Coding Accuracy: Proven Strategies to Achieve 95%+ in Your Organization
Achieving 95%+ medical coding accuracy requires a comprehensive approach combining guideline knowledge, quality assurance, continuous education, and technology support — including AI tools that reduce error rates by 30–50%.
Read more
Experience Clinical Clarity Today
Join medical coding professionals who trust CLAIRE for accurate, explained guidance. Start your free trial - no credit card required. No EMR integration needed.
The AI Medical Coding Assistant,
Built for Real-World Clinical Workflows
© 2026 CLAIRE IT AI. All rights reserved.