ainewsblitz.com

Breaking

New BabyVision Benchmark Shows Top AI Trails Toddlers on Visual Tasks

  • Foundation Models
  • Research & Papers

Frontier multimodal AI models score below the average three-year-old on basic, language-free visual reasoning tasks, according to a new benchmark called BabyVision. On the 388-task evaluation, the top performer, Gemini3-Pro-Preview, reached only 49.7% accuracy, far short of the 94.1% posted by adults and roughly 20 points below six-year-olds.

Continue reading

The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.

$20
Read this article
$29/month
Unlimited — all 9,674 articles, the full archive, and comprehension quizzes
Save 72%
$98/year
≈ $8.17/month
Unlimited, billed once a year