Google DeepMind's open-weight Gemma 4 31B model is now available in public preview on Cerebras Inference, where it runs at more than 1,800 tokens per second—a speed the company bills as the fastest multimodal inference available today. The launch pairs one of Google's most capable open models with Cerebras's wafer-scale hardware, promising near-instant responses for real-time visual and agentic workloads.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.