BREAKING
LLM Memory Pinned: ~3.6 Bits Per Parameter
0
bits
per param bf16
0
bits
per param fp32
0
GB
7B model verbatim
How They Isolated Raw Storage
1
Train on random bitstrings
↓
2
No structure to learn
↓
3
Measure verbatim recall
From Memorizing to Generalizing
1
Fill capacity by memorizing
↓
2
Saturation point
↓
3
Grokking: generalize
Privacy Implications
Harder to detect
●
Membership inference scaling laws
●
Each point is a tiny share
Still at risk
●
Rare or unique strings
●
Precision tuning barely helps
A Sharper Vocabulary for LLM Behavior
AI NEWS BLITZ
A new ICML paper puts a hard number on how much a language model can memorize.