Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU
MarkTechPost
Read full postResearchers from UC Berkeley and UT Austin developed FreeToken, an edge-native MoE serving engine enabling large models like 753B GLM-5.2 to run on a single workstation GPU by elastically mapping computation across available hardware. FreeToken supports interactive speeds for models ranging from 35B on laptops to 753B on workstations, targeting solo developers and small teams needing cost-effective local inference.

- FreeToken Unlocks Frontier MoE Inference on Consumer Hardware via Dynamic Co-Execution· InfoQ (AI, ML & Data)


