BREAKING
Grug-12B Cuts Reasoning Tokens ~70%
0%
fewer tokens
0%
quality gap
0B
parameters
Tokens Per Response
Gemma-4228.53
Grug-12B68.94
Built Cheaply on Open Weights
1Base Gemma-4-12B
2QLoRA on 5,740 rows
335 min on one A100
4Merged safetensors
Speed Gains vs Reliability Risk
Strengths~40% faster
34 vs 24 tokens/sec
Good for agents and scraping
1,134+ downloads
Limitsexperimental
Unstable on multi-hop
~47% pass@1 coding
Format not always suppressed
A Promising Proof of Concept
AI NEWS BLITZ
An open-source fine-tune of Gemma-4 slashes thinking tokens while nearly matching base quality.