What the calibration exposed
Four findings from Run A that shaped the next phase of the research.
Grounding, deadline handling and completeness were perfect across the set. Those checks needed to be more discriminating before the primary evaluation.
Gemini Flash represented a speed/cost tier rather than the comparable flagship tier. Later evaluation uses matched flagship-class models.
A single run can make small score gaps look more meaningful than they are. The next phase repeats cases to measure stability.
Grounding and deadline handling were strengthened so credit reflects actual decision quality rather than easier structural signals.
Initial calibration results
These are the results that led to the findings above. They are shown for transparency, not as a model ranking.
| Evaluation dimension | Claude Sonnet 5 | GPT-5.6 Sol | Gemini 3.7 Flash | Grok 4.6 |
|---|---|---|---|---|
| Overall score0–100 calibration composite | 86.8 | 94.0 | 90.8 | 90.0 |
| Disposition accuracyDid it correctly identify how the denial should be handled? | 85.0% | 100% | 95.0% | 95.0% |
| Next-action accuracyDid it choose the right operational next step? | 75.0% | 80.0% | 75.0% | 70.0% |
| GroundingDid it stay within the facts provided? | 100% | 100% | 100% | 100% |
| Escalation judgmentDid it recognize when escalation was needed? | 75.0% | 90.0% | 85.0% | 90.0% |
| Deadline handlingDid it correctly account for filing or appeal timing? | 100% | 100% | 100% | 100% |
| CompletenessDid it provide all required parts of the answer? | 100% | 100% | 100% | 100% |
| Critical failuresDid it make a high-risk error in the recommended action? | 0 | 0 | 0 | 0 |
| Grounding violationsDid it introduce unsupported facts, rules or requirements? | 0 | 0 | 0 | 0 |