BREAKING
Gemma 4 E2B Hits 255 tok/s In-Browser
0
tok/s
Decode speed
0
K
Context window
0
B
Effective params
255 tok/s vs Earlier Figures
New
255
Before
84
LiteRT-LM
76
An LLM Wrote the Inference Engine
1
Generate kernels
↓
2
Verify output
↓
3
Improve code
↓
4
255 tok/s
Praise and Cautions
Upsides
●
Data stays in browser
●
Instant, no install
●
Cross-platform
Cautions
●
Tied to M4 Max
●
Reproducibility unclear
●
Long-task stability
Agentic Kernels for On-Device AI
AI NEWS BLITZ
An open model now runs at 255 tokens per second inside a plain web browser.