In the second half of 2026, the AI industry's spotlight may fall on companies focused less on building bigger models and more on cutting the cost of running them, a view gaining traction across the sector.
The Inference Era
OpenAI Engineers Reportedly Halved AI Inference Costs — Overnight
Using a technique applicable to existing models, they cut the cost of running AI by more than 50% — a sign the industry has shifted from building bigger models to running them cheaper.
50%+
Inference cost cut reported on existing models
85%
Of corporate AI budgets now spent on inference
74%
NVIDIA's AI inference chip share (up from 66%)
Gartner Forecast · Cost to Run a Trillion-Parameter LLM
More than 90% cheaper by 2030 vs. 2025
Up to 100× more efficient than early-2022 models.
Prices Are Falling Fast
How cheap running models has become
Gemini 3.1 Flash
$0.10 in / $0.40 out per M tokens — ~99.7% cheaper than 2023
DeepSeek V4-Pro
~$0.86 per M words out — roughly 1/28th of Claude Opus 4.7 , near-parity on SWE-bench
Llama 3.2 3B
$0.06 input per M tokens — 70–90% cheaper than closed-source
The Upside
Lower user prices & relaxed rate limits
Improved gross margins and a competitive edge
Routing simple tasks to weaker models cuts cost 85% while keeping 95% of quality
The Caution
The specific technique was not disclosed
Skepticism over whether the claim holds in production
Wariness of cost spikes — a $200/mo bill ballooning to $10,000
Where the Workload Is Going
AI compute is shifting from training toward inference — and toward open, efficient models.
Palantir and NVIDIA now deploy open Nemotron models in air-gapped sovereign environments. Some U.S. government customers have switched from proprietary models, valuing the balance of cost, performance and control.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…