BREAKING
LLM Memory: ~3.6 Bits per Parameter
0bits
per parameter
0B
max model size
0MB
1.5B capacity
How They Measured It
1Train 500K-1.5B models
2Use random bitstrings
3Strip generalization
4Measure capacity
Precision Barely Adds Capacity
bfloat163.6
float32 low3.51
float32 high3.83
Why the Framework Matters
The Insight
Splits memorization vs generalization
Explains grokking and double descent
Implications
Bigger data dilutes memorization
Membership inference gets harder
Meta, DeepMind, Cornell, NVIDIA
AI NEWS BLITZ
A new ICML 2026 paper pins down how much language models actually memorize.