CladBench – an open benchmark for AI on UK building regulations

Hacker News
Read full post
CladBench is an open benchmark evaluating large language models on UK and EU building regulations with 536 questions across 12 categories. Seven models were tested, with Claude Opus 4.7 scoring highest and Gemini 2.5 Pro and GPT-5 performing comparably. The benchmark highlights varying model performance depending on task category rather than overall score.

More in LLM & Text Generation

DeepSeek V4.1 Flash now available on AI Gateway

Covered by 2 sources

Cohere Debuts Open-Weight 218B Mixture-of-Experts Machine Translation Model

Unite.AI