AI Surgical Coding Accuracy: What Procode’s Study Shows

·

Most claims about AI coding accuracy come from vendor benchmarks that never see peer review. On July 28, 2026, Procode Inc. published one that did. The AI-powered surgical billing company announced a $10 million Series A led by Health Velocity Capital, and tied the raise to a retrospective comparative study in PRS Global Open, the peer-reviewed, open-access journal of the American Society of Plastic Surgeons. The study measured how well Procode’s fine-tuned hybrid LLM codes real operative reports compared to general-purpose AI models and professional human auditors, and the results are a useful data point for anyone deciding how much to trust AI-generated CPT codes.

What the study actually measured

The paper, titled “Artificial Intelligence for Automated CPT Coding in Plastic and Reconstructive Surgery Using a Fine-Tuned Hybrid LLM,” compared Procode’s model against OpenAI GPT-5, Google Gemini 2.5 Pro, Anthropic Claude Sonnet 4.5, and external professional coding auditors. Researchers ran all four systems and the human auditors against 120 real operative reports pulled from plastic and reconstructive surgery cases, split evenly across three difficulty tiers: 40 easy, 40 medium, and 40 high-complexity cases.

The difficulty tiers

High-complexity cases were defined as operative reports requiring four or more CPT codes to bill correctly, the scenario where bundling rules, modifier logic, and multiple-procedure sequencing tend to break both general models and human coders. Easy cases involved straightforward, single-procedure documentation with little ambiguity in code selection.

The head-to-head numbers

Procode’s model hit 86.7% overall accuracy across all 120 cases. The best-performing general-purpose LLM, OpenAI’s model, managed 35.8%. Human professional auditors landed at 42.5%, meaning a specialty-tuned coding model more than doubled both the strongest generic AI and trained human reviewers. On the easy tier, Procode coded every case correctly. On the hardest tier, the gap widened further: Procode held 80% accuracy while the general-purpose LLMs fell to a range of 5% to 12.5%.

“Our research shows that AI trained specifically for surgical coding can reliably outperform both humans and generic models,” said Kameron Rezzadeh, MD, FACS, cofounder and Chief Medical Officer of Procode, in the Series A announcement. “That’s a step change in what billing accuracy looks like for private practice surgeons.”

Why general-purpose LLMs collapsed on hard cases

The widening gap between easy and hard cases is the real signal in this data, more than the headline accuracy number. A single-procedure note has one obvious answer, and most modern LLMs can pattern-match their way to it. A four-code operative report with staged flap reconstruction, revision work, and bundled supply codes requires a model to reason about sequencing, modifier stacking, and payer-specific bundling edits simultaneously. General-purpose LLMs, trained broadly rather than on coding-specific logic and documentation patterns, degrade fast under that load. That collapse from roughly a third correct on easy-to-medium work down to single digits on complex cases is consistent with what coding teams already experience when they’ve tried off-the-shelf AI tools on anything beyond routine E/M or single-procedure claims.

What the funding signals for surgical RCM

The $10 million round follows Procode’s stealth exit in March 2026, when the company launched with $4 million and the acquisition of The Auctus Group, a billing organization serving more than 300 plastic surgery and dermatology providers. This round funds two additional acquisitions of surgical billing companies, which Procode plans to layer its coding model into using the same playbook: buy an established RCM operation, then apply AI to the coding workflow from charge capture through claim submission. That acquire-and-automate approach is becoming a recognizable pattern among specialty-focused coding AI vendors, and it says something about where investors think the near-term value sits: not in replacing billing companies, but in buying them and making their coding faster and more defensible.

What this means for CDI and coding teams

A peer-reviewed, three-tier accuracy comparison is a rarer artifact than another vendor’s internal benchmark, and it’s worth pulling a few concrete takeaways from it rather than treating the 86.7% figure as a universal claim:

  • Specialty-tuned models outperformed general-purpose LLMs by a wide margin specifically as case complexity rose, not just on average — the accuracy gap on high-complexity cases was roughly 6x to 16x.
  • Human auditors, while more consistent than general LLMs, still landed under 50% overall accuracy on this case mix, underscoring how much operative report complexity strains manual review too.
  • Documentation quality is doing real work here: a model can only code as well as the operative note supports, which keeps CDI teams squarely in the loop even as coding automation improves.
  • The study’s scope (120 cases, one specialty) means the numbers shouldn’t be generalized to every surgical service line without similar validation.

The documentation dependency doesn’t go away

None of this removes the need for strong clinical documentation. A fine-tuned coding model still reads the operative report, not the procedure itself, so ambiguous laterality, missing device details, or vague closure descriptions will degrade any coding system’s accuracy, AI or human. If anything, a study showing an 86.7%-vs-35.8%-vs-42.5% split raises the bar for what “good enough” documentation needs to produce, since the ceiling on coding accuracy is only as high as the note underneath it.

Peer-reviewed benchmarks like this one are still uncommon in the AI coding space, and coding and CDI teams evaluating vendors should ask for this kind of independent, tiered validation rather than accepting a single blended accuracy number. Medikode’s automated medical coding platform is built around that same principle: pairing coding automation with documentation-aware validation so accuracy holds up on the complex cases, not just the routine ones. Learn more about Medikode’s automated medical coding platform.