A single one-shot prompt asking three leading AI models to build a black hole simulation produced strikingly different results, reviving debate over whether benchmark scores capture the qualitative differences that matter for creative and physics-oriented tasks.
June 2026 · Frontier Model Test
Same Prompt, Three Black Holes: What Benchmarks Don't Show
One instruction — "build a black hole simulation" — handed to three frontier models produced strikingly different renders, reviving debate over whether leaderboard scores capture the qualitative gaps that matter for creative and physics work.
GLM-5.2
Z.ai
Cartoony
Open-weight MoE · 1M context · low cost
Fugu Ultra
Sakana AI
Abstract
Fable 5-class claim · OpenAI-compatible API
Claude Fable 5
Anthropic
Physics-engine-like
1M context · strongest · high price
Cost per million tokens
Capability came at a steep price gap — the strongest render also carried the heaviest bill.
Fable 5 costs roughly 7× more on input and 11× more on output than GLM-5.2.
The Argument
"Benchmarks do not show this kind of difference — but the outputs do."
Top models are converging on standard leaderboards, blurring what separates them.
A physics-heavy render — accretion disk geometry — exposed gaps a numeric score can't convey.
Qualitative, task-specific evaluation is becoming the more meaningful way to choose.
Why it matters
On creative and simulation tasks, output quality — physical realism vs. stylized abstraction — reveals differences that leaderboard parity hides.
The caveat
This rests on a single prompt — anecdotal, not a ranking. A different instruction or seed could shift the outcomes entirely.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…