Tracking model performance as new versions and benchmarks arrive
Each point is one model’s accuracy on the day we tested it.
Open →The headline view: how each model scores on the same fixed set of 143 questions, with uncertainty made explicit through confidence intervals rather than hidden behind a single number.
Point estimate with 90% confidence interval, sorted from highest to lowest accuracy.
| Model↕ | Provider↕ | Accuracy↕ | 90% CI↕ | ECE (lower is better)↕ |
|---|
Each point is one model’s accuracy on the day we tested it.
Open →For each self-reported confidence level, we plot the model’s actual accuracy.
Open →