The Office of Inspector General doesn’t usually hand medical coders a clean, quantified example of what goes wrong when a diagnosis code outruns its documentation. Its May 28, 2026 report on acute stroke coding in Medicare Advantage did exactly that. Auditors reviewed a sample of 97 enrollees whose plans had submitted an acute stroke diagnosis for risk adjustment. None of the 97 had a matching inpatient or outpatient hospital record confirming a stroke in the same service year. Every single sampled code was unsupported, and OIG estimated the resulting overpayment at $462 million.
That is not a story about fraud. It is a story about a coding workflow that let a high-value HCC through without checking whether the rest of the chart agreed with it. It is also a clear illustration of the kind of failure agentic AI, applied correctly, is built to catch.
What OIG Actually Found
The report, numbered A-02-23-01020, focused narrowly on one diagnosis category: acute stroke codes submitted through physician or other health care professional data records, without a corresponding acute stroke diagnosis appearing on any inpatient or outpatient hospital claim for that enrollee during the same payment year. Acute stroke carries a substantial HCC risk-adjustment weight, which makes it an attractive target for chart review vendors working on contingency, and an equally attractive target for auditors looking for a payment integrity story.
OIG’s recommendation to CMS was procedural, not punitive: build a system edit that flags any acute stroke diagnosis lacking a linked hospital record before it factors into a plan’s risk score. That recommendation matters more than the dollar figure. It says, in plain terms, that the fix is not more retrospective auditing. It’s validation at the point of coding.
Why This Particular Code Slipped Through
Acute stroke is a documentation-dependent diagnosis in a way that many chronic HCC conditions are not. A patient with a history of stroke can appear in a chart repeatedly, in problem lists, medication rationales, and review-of-systems notes, without a current acute event ever having occurred. A coder or an automated system working from an isolated note, rather than the full encounter picture, can reasonably read “CVA” in a progress note and assign the active code. CMS’s own MEAT standard (Monitor, Evaluate, Assess, Treat) exists precisely to prevent that kind of read, but MEAT only works if something is actually cross-checking the encounter record, not just scanning the note in front of it.
That cross-check is where most coding workflows, human and automated alike, are thinnest. Chart review at scale tends to optimize for finding codes, not for confirming the absence of contradicting evidence elsewhere in the record.
Where Agentic AI Changes the Workflow
A single-pass classification model has the same blind spot as the workflow OIG audited: it reads a note, predicts a code, and moves on. An agentic system is built differently. Instead of one inference step, it runs a sequence of checks, each with a defined job, before a code is allowed to reach a claim or a risk-adjustment submission.
Evidence matching before submission
For an HCC like acute stroke, an evidence-matching agent would query the patient’s full encounter history for the payment year, not just the note being coded, and look for a corroborating inpatient or outpatient hospital record. No match, no automatic submission. The code gets routed for review instead of going straight into the risk score, which is the exact control OIG asked CMS to build at the system level.
Human-in-the-loop escalation
Agentic doesn’t mean unsupervised. The workflow that actually holds up under audit routes low-confidence or unsupported codes to a human coder with the specific gap identified, for example “no linked hospital encounter for acute stroke in CY2026,” rather than either auto-approving the code or auto-rejecting it. That keeps a person in the loop on the judgment call while removing the manual burden of finding the gap in the first place.
A validation-first agentic workflow for high-risk HCCs generally includes:
- Cross-referencing each HCC-eligible diagnosis against linked inpatient and outpatient encounters before it’s submitted for risk adjustment
- Flagging codes that lack MEAT-criteria documentation for coder review instead of auto-submitting them
- Maintaining a timestamped evidence trail that maps each submitted code back to its supporting documentation
- Surfacing confidence scores so compliance teams can prioritize audit-prone diagnoses like acute stroke, sepsis, and major depressive disorder
- Re-running validation whenever new encounter data arrives, rather than treating a code as settled at first submission
What This Means for MA Plans and Coding Teams
RADV audits now cover all eligible Medicare Advantage contracts annually rather than a small annual sample, which means the exposure OIG quantified for one diagnosis category in one payment year is no longer a rounding error a plan can absorb quietly. A coding workflow that can produce, on demand, the specific encounter evidence behind every high-weight HCC is no longer a nice-to-have; it’s the difference between a clean audit response and a multi-week scramble through scanned charts.
That evidence trail is also the practical test of whether an AI coding tool is actually agentic or just a labeling model with better marketing. A system that can only tell you what code it assigned isn’t the same as one that can tell you why, and show its work against the encounter record, when an auditor asks.
Coding teams evaluating AI tools this year should ask vendors a direct question: when your system assigns a high-risk HCC like acute stroke, what evidence does it check before submission, and can it show that evidence six months later during a RADV request? If the answer is “trust the model’s confidence score,” that’s not a validation workflow. It’s the same gap OIG just measured at $462 million.
Medikode’s automated medical coding platform builds evidence-linked validation into the coding workflow itself, so high-risk HCCs are checked against the full encounter record before submission, not caught after the fact in an audit letter. Learn more about Medikode’s automated medical coding platform.