BREAKING
744B GLM-5.2 Runs on 25GB, No GPU
0
B
total params
0
B
active per token
0
M
context tokens
How Colibri Fits It in RAM
1
Dense layers in RAM 9.9GB
↓
2
Experts on disk 370GB int4
↓
3
Stream only active experts
↓
4
25GB footprint total
0
tok/s
cold generation
0
tok/pass
with MTP int8
0
tok/s
potential ceiling
GLM-5.2 vs Claude Opus 4.8
GLM SWE-bench Pro
62.1
GLM Terminal-Bench
81
Claude Opus 4.8
85
A Signpost for Local AI
AI NEWS BLITZ
A pure-C engine just ran a 744-billion-parameter model on an ordinary PC with no graphics card.