Calibration explorer

Does the model know what it doesn’t know?

Accuracy alone doesn’t tell you whether a model knows the limits of its own knowledge. This page checks whether each model’s stated confidence actually predicts how often it’s right — the property that matters most before trusting AI output on a real regulatory decision.

Does the model know what it doesn’t know?

For each self-reported confidence level, we plot the model’s actual accuracy. Points on the dashed diagonal are perfectly calibrated; points below it mean the model was overconfident.

Actual accuracy

Number of questions

Expected Calibration Error (ECE):

Weighted average gap between confidence and accuracy across bins — lower is better.

Explore the other views

Leaderboard

Accuracy by model

Point estimate with 90% confidence interval, sorted from highest to lowest accuracy.

Open
Performance over time

Tracking model performance as new versions and benchmarks arrive

Each point is one model’s accuracy on the day we tested it.

Open