Capabilities & Benchmarking
How well do frontier LLMs know the Danish Building Regulations?
We test each model against the same fixed set of questions grounded in Bygningsreglementet.
Leaderboard
Accuracy by model
Point estimate with 90% confidence interval, sorted from highest to lowest accuracy.
Show data table
| Model↕ | Provider↕ | Accuracy↕ | 90% CI↕ | ECE (lower is better)↕ |
|---|
Calibration explorer
Does the model know what it doesn’t know?
For each self-reported confidence level, we plot the model’s actual accuracy.
Actual accuracy
Number of questions
Expected Calibration Error (ECE):
Weighted average gap between confidence and accuracy across bins — lower is better.
Methodology
How we build and score the benchmark
The question set is drawn from real provisions of the Danish Building Regulations.
Every model is given the same questions, with the same prompt structure.
Note
Figures on this page are placeholders pending the final published dataset.