BREAKING
NVIDIA claims GB300 up to 25x efficiency
Judge AI infra as a Pareto curve
Throughput vs responsiveness
Throughput
per MW
●
Max total tokens
●
Whole facility
Responsiveness
per user
●
Faster replies
●
Single user
0
B300 GPUs
0
Grace CPUs
0
TB/s
NVLink
0
PFLOPS
FP4 Tensor
Per-watt gains vs Hopper by model
DeepSeek V4 Pro
25
GLM5.1
20
Kimi K2.6
10
Real gains stay workload-dependent
AI NEWS BLITZ
NVIDIA says its new GB300 NVL72 delivers up to 25 times better efficiency than Hopper.