BREAKING
744B GLM-5.2 Runs on 25GB, No GPU
0B
total params
0B
active per token
0M
context tokens
How Colibri Fits It in RAM
1Dense layers in RAM 9.9GB
2Experts on disk 370GB int4
3Stream only active experts
425GB footprint total
0tok/s
cold generation
0tok/pass
with MTP int8
0tok/s
potential ceiling
GLM-5.2 vs Claude Opus 4.8
GLM SWE-bench Pro62.1
GLM Terminal-Bench81
Claude Opus 4.885
A Signpost for Local AI
AI NEWS BLITZ
A pure-C engine just ran a 744-billion-parameter model on an ordinary PC with no graphics card.