BREAKING
llm.c trains GPT in pure C/CUDA
0
%
faster vs PyTorch
0
%
less memory
0
%
speed gain
From GPT-2 124M to GPT-3 1.6B
1
GPT-2 124M
↓
2
GPT-3 miniseries
↓
3
Up to 1.6B params
Training costs at a glance
GPT-2 124M
single node
●
~90 minutes
●
~$20
GPT-2 1.6B
8x H100
●
~24 hours
●
~$600-672
Low-level, framework-free training
MIT-licensed and open on GitHub
AI NEWS BLITZ
Andrej Karpathy's llm.c pretrains GPT models in pure C and CUDA, no PyTorch needed.