Accuracy by model
Point estimate with 90% confidence interval, sorted from highest to lowest accuracy.
Open →Accuracy alone doesn’t tell you whether a model knows the limits of its own knowledge. This page checks whether each model’s stated confidence actually predicts how often it’s right — the property that matters most before trusting AI output on a real regulatory decision.
For each self-reported confidence level, we plot the model’s actual accuracy. Points on the dashed diagonal are perfectly calibrated; points below it mean the model was overconfident.
Weighted average gap between confidence and accuracy across bins — lower is better.