Machine Learning28 min reading time

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s

Hacker News
Read full post
The WASTE inference engine runs the 2.78 trillion parameter Kimi K3 model on a 64 GB consumer laptop by streaming model experts from disk, using only about 29 GB RAM at 0.5 tokens per second. This approach enables handling extremely large models without requiring server-grade hardware, though inference speed remains slow.

More in Machine Learning

Machine Learning4 min read

DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

The Next Web
Machine Learning3 min read

Nvidia and Palantir fine-tune a 30B Nemotron model for Nvidia’s supply chain. It beats a model 18 times its size.

Covered by 3 sources
Machine Learning6 min read

CoreWeave Puts Field Engineers Inside Customer Teams for Physical AI

Covered by 2 sources