AI Upcoding and the New False Claims Act Risk for Coders

·

Medical coding automation was sold to hospitals and payers as a productivity fix: faster charts closed, fewer backlogged claims, less coder burnout. A growing body of claims data suggests it has also become something else — a measurable source of False Claims Act exposure. For coding and compliance teams, that changes the question from “does the AI save time” to “can we defend every code it touched.”

The exposure is now measured, not anecdotal

An analysis from Blue Health Intelligence (BHI), reviewing anonymized commercial inpatient claims from Blue Cross Blue Shield plans covering roughly 62 million members over a three-year period ending March 2025, put estimated excess spending tied to AI-driven coding at $2.3 billion nationally — $663 million in inpatient claims and $1.67 billion in outpatient exposure. The study found a 9% rise in inpatient costs between 2023 and 2024 tied to more aggressive coding patterns, and flagged specific diagnosis categories where the shift was sharpest: acute posthemorrhagic anemia coding in maternity admissions jumped from 4% to 12.3% at hospitals with the fastest AI adoption.

BHI’s David Wennberg summarized the tension plainly: hospitals gain real productivity from automating documentation and coding, but “careful attention must be paid to ensure that newly billed diagnoses reflect the true acuity and treatment level.” That is a compliance statement as much as a cost one — a diagnosis code that doesn’t reflect true acuity is the exact fact pattern the False Claims Act targets.

Where the legal exposure actually starts

The pattern isn’t hypothetical. In January 2026, Kaiser Permanente agreed to pay $556 million to resolve False Claims Act allegations that it submitted inflated Medicare Advantage diagnosis codes to CMS for risk adjustment payments. It’s the kind of settlement that predates the current AI coding wave in its origins, but it establishes exactly the theory regulators and qui tam relators are now applying to AI-assisted coding: that a diagnosis code submitted without adequate clinical support to justify it — regardless of what generated the suggestion — is a false claim.

What regulators actually look for

OIG and DOJ investigators reconstruct the coding decision, not just the resulting code. That means asking what evidence supported a given CPT or ICD-10-CM assignment, whether a human reviewer had a real opportunity to catch an error, and whether the organization’s own coding tool created a record of that reasoning. A tool that outputs a code with no retrievable rationale leaves the health system defending a black box in litigation.

Why black-box coding suggestions raise the risk profile

Many AI coding tools on the market today are pattern-matching engines: they’re trained to predict the code a human coder would likely assign given similar documentation, without necessarily tying that prediction back to a specific guideline, NCCI edit, or piece of clinical evidence in the note. That’s fine for speed. It’s a liability problem when a payer audit or OIG sample asks why a particular E/M level or HCC category was selected, and the honest answer is “the model’s training data suggested it was likely.”

Agentic coding architectures — where the system identifies supporting evidence in the note, checks it against coding guidelines, and reconciles ambiguity before proposing a code — produce a fundamentally different artifact: a reasoning trail a compliance officer can actually audit after the fact, not just a code.

What an audit-ready coding workflow looks like

Coding and compliance teams evaluating AI tools should be asking for specific, inspectable outputs, not just accuracy claims. At minimum, that means:

  • A documented evidence citation for every code, pointing to the specific note text that supports it
  • Guideline or NCCI-edit validation attached to the suggestion, not just a confidence score
  • A clear flag when documentation is ambiguous or insufficient, routed to human review rather than auto-submitted
  • A versioned, retrievable audit trail showing what was suggested, what a human changed, and why
  • Coding-pattern monitoring that would catch the kind of category-level drift BHI identified before an external auditor does

What compliance leaders should do now

None of this requires abandoning AI-assisted coding — the productivity case is real, and the alternative of unassisted manual coding at current claim volumes isn’t realistic for most organizations. It does mean treating “can we show our work” as a procurement requirement, not an afterthought. Coding leaders should ask vendors for a sample audit trail before signing, run periodic internal reviews of AI-assisted codes against the same criteria an OIG auditor would use, and make sure every AI-suggested code that reaches a claim has a human sign-off that’s actually meaningful, not a rubber stamp built into the workflow to hit throughput targets.

The organizations that get audited in 2027 won’t be the ones that used AI to code faster. They’ll be the ones that can’t explain, code by code, why the AI was right. Medikode’s automated medical coding platform is built around exactly this distinction — every suggested code is tied to cited clinical evidence and validated against current coding guidelines, so coding and compliance teams have a defensible record, not just a faster one.