RevCycleAIAI
Revenue Cycle Intelligence · Published Daily
RevCycleAI Research

Frontier AI Model Performance in RCM Denials

Run A — Calibration · 20 synthetic denial cases · Preliminary results

Our first calibration used a controlled set of RCM denial scenarios to validate the evaluation framework, response structure and scoring approach before moving to a larger primary study.

Calibration — not a model ranking. These scores are preserved because they showed where the evaluation framework was working and where it needed stronger controls. They should not be interpreted as production accuracy, a definitive comparison of model families, or the RevCycleAI Research publication baseline.
Run A / calibration
Cases 20
Purpose Framework validation
Status Historical calibration

What the calibration exposed

Four findings from Run A that shaped the next phase of the research.

01
Some dimensions didn't separate the models

Grounding, deadline handling and completeness were perfect across the set. Those checks needed to be more discriminating before the primary evaluation.

02
Model tiers weren't perfectly matched

Gemini Flash represented a speed/cost tier rather than the comparable flagship tier. Later evaluation uses matched flagship-class models.

03
Repeatability hadn't been measured

A single run can make small score gaps look more meaningful than they are. The next phase repeats cases to measure stability.

04
The scoring framework needed stronger controls

Grounding and deadline handling were strengthened so credit reflects actual decision quality rather than easier structural signals.

Initial calibration results

Historical Run A results, shown for transparency.

Evaluation dimensionClaude Sonnet 5GPT-5.6 SolGemini 3.7 FlashGrok 4.6
Overall score86.894.090.890.0
Disposition accuracy85.0%100%95.0%95.0%
Next-action accuracy75.0%80.0%75.0%70.0%
Grounding100%100%100%100%
Escalation judgment75.0%90.0%85.0%90.0%
Deadline handling100%100%100%100%
Completeness100%100%100%100%
The important result wasn't who scored highest. Run A exposed saturated dimensions, mismatched provider tiers and the need to measure repeatability before treating small gaps as meaningful.