RevCycleAIAI
Revenue Cycle Intelligence · Published Daily
RevCycleAI Research

Frontier AI Model Performance in RCM Denials

Run A — Calibration · 20 synthetic denial cases · Preliminary results

Our first calibration used a controlled set of RCM denial scenarios to validate the evaluation framework, response structure and scoring approach before moving to a larger primary study.

Calibration — not a model ranking. These scores are preserved because they showed where the evaluation framework was working and where it needed stronger controls. They should not be interpreted as production accuracy, a definitive comparison of model families, or the RevCycleAI Research publication baseline.
Run A / calibration
Cases 20
Purpose Framework validation
Status Historical calibration

What the calibration exposed

Four findings from Run A that shaped the next phase of the research.

01
Some dimensions didn't separate the models

Grounding, deadline handling and completeness were perfect across the set. Those checks needed to be more discriminating before the primary evaluation.

02
Model tiers weren't perfectly matched

Gemini Flash represented a speed/cost tier rather than the comparable flagship tier. Later evaluation uses matched flagship-class models.

03
Repeatability hadn't been measured

A single run can make small score gaps look more meaningful than they are. The next phase repeats cases to measure stability.

04
The scoring framework needed stronger controls

Grounding and deadline handling were strengthened so credit reflects actual decision quality rather than easier structural signals.

Initial calibration results

These are the results that led to the findings above. They are shown for transparency, not as a model ranking.

How to read this: the overall score combined disposition, next-action selection, grounding, deadline handling, escalation judgment and completeness. Run A was designed to tell us whether the evaluation framework behaved sensibly — not to establish a final leaderboard.
Evaluation dimensionClaude Sonnet 5GPT-5.6 SolGemini 3.7 FlashGrok 4.6
Overall score0–100 calibration composite86.894.090.890.0
Disposition accuracyDid it correctly identify how the denial should be handled?85.0%100%95.0%95.0%
Next-action accuracyDid it choose the right operational next step?75.0%80.0%75.0%70.0%
GroundingDid it stay within the facts provided?100%100%100%100%
Escalation judgmentDid it recognize when escalation was needed?75.0%90.0%85.0%90.0%
Deadline handlingDid it correctly account for filing or appeal timing?100%100%100%100%
CompletenessDid it provide all required parts of the answer?100%100%100%100%
Critical failuresDid it make a high-risk error in the recommended action?0000
Grounding violationsDid it introduce unsupported facts, rules or requirements?0000
The important result wasn't who scored highest. Run A showed that several dimensions were saturated, the provider tiers were not perfectly matched, and a single pass could not tell us whether small score differences were repeatable. Those observations informed the controls and validation steps used in the next phase.
Why preserve Run A? Calibration is part of the research process. Showing what the first run exposed makes the evolution of the methodology visible. Run A remains historical calibration data and is not directly comparable with later runs using the revised rubric and matched model tiers.