Business & EnterpriseAlibaba Backs Ex-Staffer’s AI Testing Lab at $2.5 Billion ValueBTCovered by 2 sources11h ago
LLM & Text Generation·8 min readBenchMIRT: What are LLM benchmarks actually measuring?HHugging Face9d ago
Machine Learning·14 min readThe Open ASR Leaderboard Adds Its First Global South LanguageHHugging Face14d ago
Agents·5 min readDeepSeek launches an experimental multimodal model to rival AnthropicTTCovered by 2 sources20d ago
Cybersecurity·6 min readReading Zhipu’s GLM-5.3 results past the headline numberAAI News (TechForge)23d ago
LLM & Text Generation·6 min readCladBench – an open benchmark for AI on UK building regulationsHHacker News24d ago
Machine Learning·11 min readStreaming benchmark and recommendation results to MLflow with Amazon SageMaker AIAAWS BlogJul 6
LLM & Text GenerationFable 5 was beating GPT 5.5 on every major benchmark. Then the US government pulled it offline.TThe Next WebJun 14
LLM & Text Generation**Introducing SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding**HHugging FaceMar 19