BREAKING
FutureShow benchmarks AI on forecasts
How FutureShow tests AI agents
1Real events
2Polymarket bet
3Log & compare
4Live dashboard
Round 1 accuracy vs models
DeepSeek95.4
GPT-592.5
Gemini87.3
0%
DeepSeek vs Human
0%
GPT-5 vs Human
0%
Gemini vs Human
Open experiment, real caveats
Strengths
Open-source on GitHub
Real-money PnL tracking
Live public dashboard
Caveats
Models lag human accuracy
Differing forecast counts
Tool costs vary
Round 2 now underway
AI NEWS BLITZ
An HKU-linked team just launched FutureShow to test frontier AI on real-world predictions.