Machine LearningDev9 min reading time

Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU

MarkTechPost
Read full post
Researchers from UC Berkeley and UT Austin developed FreeToken, an edge-native MoE serving engine enabling large models like 753B GLM-5.2 to run on a single workstation GPU by elastically mapping computation across available hardware. FreeToken supports interactive speeds for models ranging from 35B on laptops to 753B on workstations, targeting solo developers and small teams needing cost-effective local inference.

More on this story


More in Machine Learning

Machine Learning3 min read

OpenAI Releases GPT-6 Astra for Coding and Computer Use

InfoQ (AI, ML & Data)
Machine Learning4 min read

Arlequin AI raises €28M to build novel AI models that learn complex relationships at scale

SiliconANGLE
Machine Learning4 min read

DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

The Next Web