BREAKING
NVIDIA AI Agent Trains Vision Model
0%
Baseline
0%
After training
The Autoresearch Loop
1Configure environment
2Train model
3Evaluate results
4Propose next test
NeMo Gym vs NeMo RL
NeMo GymApache 2.0
Evaluate and RL-train models
Scales via Ray to thousands of envs
NeMo RLApache 2.0
Post-training RL library
Supports GRPO and Qwen3
A proof of concept, not a benchmark
Toward agentic RL research
AI NEWS BLITZ
NVIDIA shows a coding agent that trained a vision model almost on its own.