BREAKING
Gemma 4 31B Hits 1,800+ Tokens/Sec
0
t/s
throughput
0
x
vs GPU endpoint
0
x
vs Anthropic model
0
Gemma 4 index
0
Claude Haiku 4.5
0
s
time to first token
Why Speed Makes Agents Practical
1
Call tool
↓
2
Observe result
↓
3
Retry fast
↓
4
Verify
0
k
paid context
0
$/M
input tokens
0
$/M
output tokens
Open Model Chases Fast Inference
AI NEWS BLITZ
Google's Gemma 4 31B now runs on Cerebras at over eighteen hundred tokens per second.