BREAKING
Gemma 4 31B Hits 1,800+ Tokens/Sec
0t/s
throughput
0x
vs GPU endpoint
0x
vs Anthropic model
0
Gemma 4 index
0
Claude Haiku 4.5
0s
time to first token
Why Speed Makes Agents Practical
1Call tool
2Observe result
3Retry fast
4Verify
0k
paid context
0$/M
input tokens
0$/M
output tokens
Open Model Chases Fast Inference
AI NEWS BLITZ
Google's Gemma 4 31B now runs on Cerebras at over eighteen hundred tokens per second.