BREAKING
llm.c trains GPT-2 in pure C/CUDA
0
%
faster vs PyTorch
0
M
GPT-2 params
Minimalist vs full framework
llm.c
pure C/CUDA
●
No PyTorch or cPython
●
Kernel-level optimization
●
Training focused
PyTorch
3.3M lines
●
11,000+ files
●
Mature ecosystem
●
General purpose
0
min
training time
0
$
cost
0
x
A100 80GB
Features and known limits
1
MPI + NCCL multi-node
↓
2
bfloat16 + Flash Attention
↓
3
CPU lags PyTorch
↓
4
MPI/NCCL setup hurdle
MIT-licensed, built to teach internals
AI NEWS BLITZ
Andrej Karpathy's llm.c pretrains GPT-2 using nothing but raw C and CUDA.