Measuring how well AI understands the Danish Building Regulations
Byggeriets LLM independently benchmarks large language models on their ability to interpret and apply Bygningsreglementet (BR18/BR23).
Six frontier models, one Danish rulebook
Every model answers the same set of questions drawn from the Danish Building Regulations.
Top accuracy
Gemini 3.5 Flash currently leads the leaderboard.
90% CI 89.7–97.7%Models evaluated
Frontier models from OpenAI, Anthropic and Google.
GPT-5.5 · GPT-5.4 · Opus 4.7 · Sonnet 4.6 · Gemini 3.5 Flash · Gemini 2.5 ProBenchmark questions
Built from real provisions in the Building Regulations.
Bygningsreglementet (BR18/BR23)Best calibration (ECE)
Gemini 2.5 Pro is the most honest about its own uncertainty.
Expected Calibration ErrorAccuracy spread
Gap between the strongest and weakest model tested.
89.5% – 93.7%Avg. question length
Typical length of a benchmark question, in characters.
Range 200–1800 charactersWhich model gets the Building Regulations right?
A live leaderboard of model accuracy with 90% confidence intervals, plus a calibration explorer.
Look inside the benchmark dataset
Explore the composition of the question set that underlies every result on this site.
Our research, written up
Reports and papers describing our methodology, findings, and their implications.
Latest from Byggeriets LLM
An independent benchmark, built for the Danish construction industry
Byggeriets LLM is a research initiative that evaluates how reliably large language models can answer questions grounded in Bygningsreglementet.
More about the projectDeveloped in collaboration with
Get new benchmark results in your inbox
A short note whenever we add a model, publish a report, or update the dataset.
By subscribing you agree to receive occasional emails.