The CLAIRE Blog
AI HCC Gaps in Clinical Documentation: 2026 Guide
How NLP and large language models surface conditions the record supports but the claim never carried: FHIR ingestion, extraction split from HCC mapping, V28 trumping logic, and why MEAT validation and the compliant query stay with the coder.
Valentina GallegosBA, CPC, CRCTable of Contents
- Executive Summary: AI's Role in Identifying HCC Gaps
- Understanding HCCs and Documentation Gaps
- How AI Identifies HCC Gaps: The Core Mechanism
- Key AI Capabilities for Comprehensive HCC Gap Closure
- Operational and Financial Benefits of AI in HCC Documentation
- Challenges and Limitations of AI in HCC Coding
- Implementing AI for HCC Gap Identification: Best Practices
- Step-by-Step AI Workflow for HCC Gap Closure
- Frequently Asked Questions About AI and HCC Gaps
- Key Takeaways
"Can AI help identify HCC gaps in clinical documentation?" The short answer is yes, and the mechanics behind that capability are already reshaping how risk adjustment teams operate. Healthcare organizations face mounting pressure to capture every clinically supported diagnosis while staying defensible under CMS audit scrutiny. Artificial intelligence, specifically natural language processing and large language models, offers a way to scan unstructured clinical notes at scale and surface conditions that documented care supports but coded claims miss.
Executive Summary: AI's Role in Identifying HCC Gaps
Yes, AI can help identify HCC gaps in clinical documentation, and the technology has moved well beyond pilot testing into production risk adjustment workflows. Natural Language Processing (NLP) and Large Language Models (LLMs) analyze unstructured clinical notes, discharge summaries, and lab results to find documented conditions that never made it into coded claims. The result is improved Risk Adjustment Factor (RAF) score accuracy, faster chart review, and stronger RADV audit readiness.
But the technology has clear boundaries. AI surfaces possibilities, flags potential gaps, and triages charts so human coders focus where it matters most. The actual coding decision, the MEAT criteria validation, and the compliant query generation still belong to trained professionals. AI augments the workflow; it does not replace the human judgment that keeps submissions defensible.
- AI uses NLP and LLMs to scan unstructured EHR data for undocumented or under-documented conditions
- The technology improves RAF accuracy, chart review speed, and RADV audit readiness
- Human coders remain essential for validating findings, applying MEAT criteria, and generating compliant provider queries
Understanding HCCs and Documentation Gaps
Hierarchical Condition Categories (HCCs) are risk adjustment categories CMS uses to calculate RAF scores for Medicare Advantage beneficiaries. Each HCC carries a weight that reflects the expected cost of caring for a patient with that condition. Accurate documentation and coding directly determine whether an organization receives appropriate reimbursement for the complexity of its patient population.
Documentation gaps take several forms. ICD-10-CM has required heart failure type specificity (systolic, diastolic, combined) since its implementation in 2015; the V28 (FY2024) update modified the heart failure code set but did not newly impose this requirement. A chronic condition coded last year might not be re-documented in the current year, dropping it from the risk score entirely. Or a patient's medication list showing insulin and metformin might suggest diabetes management that no claim captures.
The financial and compliance stakes are significant. Under-documented conditions mean lost revenue, while unsupported diagnoses create RADV audit exposure. CMS requires every submitted diagnosis to trace back to a face-to-face encounter with supporting documentation in the medical record, coded per the ICD-10-CM Official Guidelines. Organizations must also certify, under 42 CFR 422.504(l), that submitted data is accurate, complete, and truthful.
- Missing specificity: conditions documented without the clinical detail V28 requires, such as heart failure without systolic or diastolic classification or CKD without a documented stage
- Missing MEAT evidence: conditions listed in the problem list but lacking documentation that the provider monitored, evaluated, assessed, or treated them during the encounter
- Uncaptured chronic conditions: diagnoses coded in prior years but not re-documented in the current year, which drops them from the RAF calculation
How AI Identifies HCC Gaps: The Core Mechanism
The process starts with data ingestion. AI tools connect to Electronic Health Records (EHRs) via FHIR APIs, pulling clinical notes, structured data, and claims history using resources like DocumentReference, Condition, Observation, and Encounter. This FHIR-native approach creates a canonical patient record rather than relying on flat clinical data exports, which is an architectural best practice for reliable NLP pipelines.
Once the data is ingested, NLP and LLMs go to work on the unstructured text. Extraction models such as AWS Comprehend Medical, Azure Health Insights, or fine-tuned LLMs identify condition mentions in narrative notes. A separate HCC mapping layer then translates those extracted concepts into ICD-10-CM codes and assigns them to V28 HCC categories. Splitting extraction from mapping lets teams update one without rebuilding the other.
The filtering step is where gaps surface. AI compares extracted conditions against current-year coded diagnoses and removes already-captured conditions. What remains is a ranked list of suspect conditions, ordered by HCC weight and confidence score. Pattern recognition adds another layer: if a patient's medication list includes insulin, the AI flags potential diabetes. If a discharge summary mentions cognitive decline requiring 24-hour supervision, that suggests dementia even if the word itself never appears.
V28 trumping logic adds complexity. The model has 115 HCC categories arranged in hierarchical groups. When a patient has both diabetes with chronic complications (HCC 37) and diabetes without complication (HCC 38), the more severe HCC 37 trumps HCC 38. But the clinical documentation must support the higher-severity code specifically. AI can identify the hierarchy, but confirming that documentation supports the more specific code requires human review.
- Ingest structured and unstructured EHR data via FHIR APIs using DocumentReference, Condition, Observation, and Encounter resources
- Extract condition mentions from unstructured text using clinical NLP tools or fine-tuned LLMs
- Map extracted concepts to ICD-10-CM codes and V28 HCC categories in a separate mapping layer
- Filter out already-coded conditions and rank remaining suspects by HCC weight and confidence score
Key AI Capabilities for Comprehensive HCC Gap Closure
AI brings several distinct capabilities to HCC gap closure, each addressing a different point in the documentation lifecycle. AI tools that identify documentation gaps process unstructured data from clinical notes, discharge summaries, and consultation reports to find condition mentions that claims data misses. Real-time point-of-care prompts surface suspected HCCs pre-visit or during the encounter, often embedded directly in EHRs like Epic and Cerner via SMART on FHIR integrations. Instead of a monthly PDF gap report, clinicians see relevant conditions at the moment of care.
- Unstructured data analysis: scans narrative notes, discharge summaries, and consultation reports for documented-but-uncoded conditions
- Point-of-care surfacing: delivers suspect condition lists within EHR workflows using SMART on FHIR, InBasket, or side-panel embedding
- Automated MEAT validation: checks whether documentation supports that the provider monitored, evaluated, assessed, or treated each condition
- V28 trumping logic application: applies hierarchy rules to identify or recommend the likely HCC category for human review when multiple related conditions exist
- Proactive annual recapture: identifies chronic conditions from prior years that require re-documentation each year for prospective risk adjustment
Operational and Financial Benefits of AI in HCC Documentation
The operational gains are measurable. AI can reduce risk adjustment coordination from approximately 30 minutes to 3 minutes per member, according to Pelica Health case study data. A 2026 study on HCC-Coder, an NLP system mapping narrative documentation to 115 CMS HCC codes, achieved a macro-F1 of 0.779 and a micro-F1 of 0.756. A 4 percentage point F1-score advantage over GPT-4o is confirmed in at least one study (GPT-5 low reasoning effort F1=0.84 vs. GPT-4o F1=0.80 for localized disease classification); actual F1 gaps range from 2 to 14 percentage points, and GPT-4o does outperform comparison models in some studies.
The financial impact compounds at scale. For an ACO with 15,000 beneficiaries, a 0.05 RAF lift can result in approximately $8.1M in annual RAF-adjusted revenue, based on Vizier Health data. Beyond the numbers, AI helps reviewers identify potential unsupported diagnoses and perform pre-submission checks, while retaining human validation as the control. Pre-submission validation catches hierarchy conflicts, unsupported combination codes, and codes lacking a qualifying encounter, shrinking audit exposure at the source.
AI also reduces administrative burden by triaging charts. Instead of reading every line of a 40-page chart, human coders review only flagged records. This workflow optimization means coding teams spend their time on validation and query generation rather than initial screening, which is where their expertise adds the most value.
- Time savings: reduces risk adjustment coordination from approximately 30 minutes to 3 minutes per member (Pelica Health)
- Accuracy: HCC-Coder NLP system achieved macro-F1 of 0.779 and micro-F1 of 0.756 (2026 study)
- Financial impact: a 0.05 RAF lift yields approximately $8.1M annually for a 15,000-beneficiary ACO (Vizier Health)
- Audit readiness: pre-submission validation catches unsupported codes before they reach CMS
Challenges and Limitations of AI in HCC Coding
Human oversight is non-negotiable. AI can find a diagnosis mentioned in a chart, but it cannot reliably determine whether the provider actually monitored, evaluated, assessed, or treated that condition during a specific visit. That distinction matters because MEAT criteria, while useful as a documentation-review mnemonic, do not replace the official ICD-10-CM, CMS, encounter, or payer rules that determine whether a diagnosis is reportable. The collaboration between AI and human coders is what makes the workflow effective.
AI also produces errors. Tools may infer specificity not explicitly documented, defaulting to the most general code when documentation is ambiguous. Under V28's stricter requirements, this tendency creates real audit exposure. If AI suggests E11.65 (Type 2 diabetes with hyperglycemia) but the note only says "diabetes, sugars running high" without a documented glucose reading or A1C, that code gets rejected in a RADV audit and the RAF score gets clawed back.
Data quality matters as much as model quality. AI performance depends on the completeness and clarity of EHR data. Poor documentation limits what any tool can surface. AI can also miss conditions documented narratively rather than in structured fields. A provider writing "patient's cognitive decline has worsened, now requiring 24-hour supervision" has documented dementia progression, but AI might not flag it if the word "dementia" never appears.
The regulatory landscape continues to evolve. The January 2026 Kaiser settlement of $556M highlights the compliance risks associated with automated diagnosis submission. As one analysis puts it, "The RADV audit does not review the software output. It reviews the physician's note." "The AI suggested it" is not a defensible position. AI is a tool to augment human expertise, and that boundary must be clear in any implementation.
- Human-in-the-loop requirement: AI surfaces possibilities but cannot validate MEAT criteria or confirm clinical specificity
- Hallucination risk: AI may infer undocumented specificity or default to general codes, creating V28 audit exposure
- Data quality dependency: incomplete or ambiguous EHR documentation limits AI effectiveness
- Regulatory risk: the January 2026 Kaiser settlement ($556M) underscores the compliance stakes of automated diagnosis submission
Implementing AI for HCC Gap Identification: Best Practices
EHR interoperability comes first. Use FHIR standards and SMART on FHIR protocols to integrate AI tools into existing EHR workflows, including Epic, Cerner, and Meditech. The AI should ingest clinical data, surface pre-visit suspect condition lists, and deliver point-of-care prompts within existing clinician workflows without requiring separate logins or manual data entry.
Integration into clinical documentation integrity programs matters as much as the technology itself. AI should surface suspected HCCs pre-visit and route opportunities to the appropriate human reviewer, whether that is a coder or clinician. The AI handles initial screening and triage; human coders validate findings, evaluate documentation against MEAT criteria, and generate compliant provider queries. Governance should include monitoring false positive rates, tracking RADV defensibility of AI-suggested diagnoses, and maintaining documentation of clinical reasoning for every suggestion.
CLAIRE fits naturally into this workflow. Built by coders for coders, CLAIRE interprets clinical documentation and suggests ICD-10-CM, CPT, and HCPCS codes with clear explanations of clinical reasoning and coding pathways. It generates compliant, clinically supported CDI queries that improve physician response rates. The searchable medical code reference integrates ICD-10-CM, CPT, HCPCS, and ICD-10-PCS codes into a single platform, reducing the need to switch between tools. CLAIRE provides instant expert-level clarification, removing the delays and bottlenecks that traditional peer consultations create.
- Use FHIR and SMART on FHIR standards for standards-based EHR integration with Epic, Cerner, and Meditech
- Route AI-surfaced suspect conditions to human coders for MEAT validation and compliant query generation
- Establish governance: monitor false positive rates, track RADV defensibility, and log every suggestion for compliance
- Choose AI tools built by coders for coders to provide stronger evidence of domain alignment, subject to validation
Step-by-Step AI Workflow for HCC Gap Closure
The AI workflow for HCC gap closure follows five sequential steps, from data ingestion through claim tracking. Each step builds on the previous one, with human validation as the critical checkpoint before any diagnosis reaches CMS. Pre-submission validation at the end of the workflow catches errors before they create audit exposure.
- Data Ingestion: AI ingests clinical notes, structured EHR data, and claims history via FHIR APIs to create a canonical patient record using DocumentReference, Condition, Observation, and Encounter resources.
- Condition Extraction: NLP and LLMs extract clinical concepts from unstructured text. A separate HCC mapping layer then translates those concepts into ICD-10-CM codes and V28 HCC categories.
- Gap Identification: AI filters out conditions already coded in the current year, ranks remaining suspect conditions by HCC weight and confidence score, and applies V28 trumping logic for hierarchy conflicts.
- Human Validation: A coder or CDI specialist reviews flagged charts, confirms documentation support against MEAT criteria, and generates compliant provider queries when documentation needs clarification.
- Claim Submission and Tracking: Validated diagnoses are submitted, and AI tracks the claim through the 277CA (payer acknowledgment), MAO-004 (Medicare Advantage Organization submission), MMR (Monthly Membership Report), and MOR (Model Output Report) lifecycle to confirm closure.
Frequently Asked Questions About AI and HCC Gaps
This section addresses the questions medical coders and CDI specialists ask most often about AI's role in HCC gap identification, RADV defensibility, V28 transition impacts, and EHR integration.
Is AI-captured HCC documentation RADV-defensible?
Yes, if the underlying medical record supports the diagnosis. CMS does not distinguish between human- and AI-identified codes during RADV audits; the question is whether the clinical documentation substantiates the submitted diagnosis. AI can help surface potential gaps, but a human coder must validate that documentation meets MEAT criteria and supports the code before submission.
Does using AI for HCC coding trigger CMS audit risk?
Using AI itself does not trigger audit risk. However, submitting diagnoses based solely on AI suggestions without proper clinical documentation support does. Robust governance and human-in-the-loop validation are essential to help reduce the risk that AI-driven queries lead to unsupported code additions rather than genuine documentation.
How does AI integrate with existing EMR systems?
AI tools typically integrate with EHRs like Epic and Cerner via FHIR APIs and SMART on FHIR protocols. This allows AI to ingest clinical data, surface pre-visit suspect condition lists, and deliver point-of-care prompts within existing clinician workflows without requiring separate logins or manual data entry.
Is AI for HCC gap identification a clinical decision support tool or a billing tool?
It functions as both. AI provides clinical decision support by surfacing undocumented conditions for provider review at the point of care. It also supports billing accuracy by helping confirm that all documented conditions are captured for risk adjustment. However, it cannot replace human coders who apply ICD-10-CM guidelines, payer rules, and MEAT criteria validation.
Key Takeaways
- AI can identify HCC gaps by using NLP and LLMs to scan unstructured clinical notes for documented-but-uncoded conditions
- The technology improves RAF accuracy, chart review speed, and RADV audit readiness, with measurable time and financial benefits
- Human-in-the-loop validation remains essential: AI surfaces possibilities, but coders confirm MEAT compliance and clinical specificity
- V28's stricter specificity requirements and hierarchical trumping logic demand careful AI implementation with human oversight
- CLAIRE supports the HCC gap closure workflow by interpreting documentation, suggesting codes with clinical reasoning, and generating compliant CDI queries
Ready to strengthen your HCC documentation workflow? CLAIRE's AI-powered coding and CDI assistant helps coders interpret clinical documentation, suggest accurate ICD-10-CM codes with explanations of clinical reasoning, and generate compliant provider queries that get responses. See how CLAIRE can fit into your risk adjustment process and close documentation gaps with confidence.
The AI makes the coder faster. The coder makes the AI accurate. Remove either one and you have a problem.
Related Posts
AI Solutions for Outpatient Medical Coding: Challenges and Opportunities
Outpatient coding is the highest-volume environment in healthcare. AI hits 94-97% accuracy on E&M selection, procedure coding, and modifiers, cutting turnaround from days to hours and lifting accuracy 20-35%.
Read moreBest AI Tools for Inpatient Medical Coding 2026
Inpatient coding resists automation in ways outpatient does not. A look at CodaMetrix, Nym Health, 3M 360 Encompass and CLAIRE across MS-DRG accuracy, CDI integration and compliance, and why 55-70% of encounters still reach a human coder.
Read moreAutonomous Medical Coding: How Close Are We in 2026?
Autonomous coding is already live for radiology, pathology, and routine visits at 97-99% accuracy. Complex cases still need human review. Industry consensus: 40-50% of encounters fully autonomous by 2027, 70-80% by 2029.
Read more
Experience Clinical Clarity Today
Join medical coding professionals who trust CLAIRE for accurate, explained guidance. Start your free trial - no credit card required. No EMR integration needed.
The AI Medical Coding Assistant,
Built for Real-World Clinical Workflows
© 2026 CLAIRE IT AI. All rights reserved.