ainewsblitz.com

Breaking

Google's Gemma 4 31B Multimodal Model Now Runs on Cerebras at Over 1,800 Tokens Per Second

  • Foundation Models
  • Infra & Chips
  • AI Agents

Google DeepMind's open-weight Gemma 4 31B model is now available in public preview on Cerebras Inference, where it runs at more than 1,800 tokens per second—a speed the company bills as the fastest multimodal inference available today. The launch pairs one of Google's most capable open models with Cerebras's wafer-scale hardware, promising near-instant responses for real-time visual and agentic workloads.

Continue reading

The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.

$20
Read this article
$29/month
Unlimited — all 5,716 articles, the full archive, and comprehension quizzes
Save 72%
$98/year
≈ $8.17/month
Unlimited, billed once a year