BREAKING
LLM Memory: ~3.6 Bits per Parameter
0
bits
per parameter
0
B
max model size
0
MB
1.5B capacity
How They Measured It
1
Train 500K-1.5B models
↓
2
Use random bitstrings
↓
3
Strip generalization
↓
4
Measure capacity
Precision Barely Adds Capacity
bfloat16
3.6
float32 low
3.51
float32 high
3.83
Why the Framework Matters
The Insight
●
Splits memorization vs generalization
●
Explains grokking and double descent
Implications
●
Bigger data dilutes memorization
●
Membership inference gets harder
Meta, DeepMind, Cornell, NVIDIA
AI NEWS BLITZ
A new ICML 2026 paper pins down how much language models actually memorize.