BREAKING
FutureShow benchmarks AI on forecasts
How FutureShow tests AI agents
1
Real events
↓
2
Polymarket bet
↓
3
Log & compare
↓
4
Live dashboard
Round 1 accuracy vs models
DeepSeek
95.4
GPT-5
92.5
Gemini
87.3
0
%
DeepSeek vs Human
0
%
GPT-5 vs Human
0
%
Gemini vs Human
Open experiment, real caveats
Strengths
●
Open-source on GitHub
●
Real-money PnL tracking
●
Live public dashboard
Caveats
●
Models lag human accuracy
●
Differing forecast counts
●
Tool costs vary
Round 2 now underway
AI NEWS BLITZ
An HKU-linked team just launched FutureShow to test frontier AI on real-world predictions.