BREAKING
LLM Memory Pinned: ~3.6 Bits Per Parameter
0bits
per param bf16
0bits
per param fp32
0GB
7B model verbatim
How They Isolated Raw Storage
1Train on random bitstrings
2No structure to learn
3Measure verbatim recall
From Memorizing to Generalizing
1Fill capacity by memorizing
2Saturation point
3Grokking: generalize
Privacy Implications
Harder to detect
Membership inference scaling laws
Each point is a tiny share
Still at risk
Rare or unique strings
Precision tuning barely helps
A Sharper Vocabulary for LLM Behavior
AI NEWS BLITZ
A new ICML paper puts a hard number on how much a language model can memorize.