Machine Learning6 min reading time

Smaller, faster, safer: running Kimi and GLM at scale

Hacker News
Read full post
Cloudflare's Workers AI enhances serving of large models like Moonshot's Kimi K-series and Z.ai's GLM by quantizing KV cache to 8-bit floats, compressing model weights, and protecting shared caches, doubling context capacity and improving efficiency without accuracy loss.

More in Machine Learning

Machine Learning3 min read

OpenAI Releases GPT-6 Astra for Coding and Computer Use

InfoQ (AI, ML & Data)
Machine Learning4 min read

Arlequin AI raises €28M to build novel AI models that learn complex relationships at scale

SiliconANGLE
Machine Learning4 min read

DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

The Next Web