BREAKING
Gemma 4 E2B Hits 255 tok/s In-Browser
0tok/s
Decode speed
0K
Context window
0B
Effective params
255 tok/s vs Earlier Figures
New255
Before84
LiteRT-LM76
An LLM Wrote the Inference Engine
1Generate kernels
2Verify output
3Improve code
4255 tok/s
Praise and Cautions
Upsides
Data stays in browser
Instant, no install
Cross-platform
Cautions
Tied to M4 Max
Reproducibility unclear
Long-task stability
Agentic Kernels for On-Device AI
AI NEWS BLITZ
An open model now runs at 255 tokens per second inside a plain web browser.