A developer has released Grug-12B, an experimental open-source fine-tune of Google's Gemma-4 that claims to use roughly 69.8% fewer thinking tokens than the base model while staying within about 2% of its performance. The model, published on Hugging Face in early July 2026, aims to bring the kind of efficient internal reasoning seen in leading closed models to a freely downloadable 12-billion-parameter system.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.