Capabilities & Benchmarking

How well do frontier LLMs know the Danish Building Regulations?

We test each model against the same fixed set of questions grounded in Bygningsreglementet.

Leaderboard

Accuracy by model

Point estimate with 90% confidence interval, sorted from highest to lowest accuracy.

Show data table
Full leaderboard: accuracy, confidence interval and calibration error (ECE) per model.
Model Provider Accuracy 90% CI ECE (lower is better)
Calibration explorer

Does the model know what it doesn’t know?

For each self-reported confidence level, we plot the model’s actual accuracy.

Actual accuracy

Number of questions

Expected Calibration Error (ECE):

Weighted average gap between confidence and accuracy across bins — lower is better.

Methodology

How we build and score the benchmark

The question set is drawn from real provisions of the Danish Building Regulations.

Every model is given the same questions, with the same prompt structure.

Note

Figures on this page are placeholders pending the final published dataset.