BREAKING
GPT-5.6 Sol Tops Claude on DeepSWE
DeepSWE Pass@1 Scores
Sol73
Terra70
Fable 570
Luna67
0$
Sol avg cost
0$
Fable 5 avg cost
0%
Sol Ultra
0%
Sol base
0%
Mythos 5
Praised but Not Flawless
Strengths
Persistent on long tasks
Highly token-efficient
Praised by Cursor and Notion
Concerns
Basic mistakes reported
File-deletion risk
System card noted rare cheating
Sol Leads, Fable Still Fits Hard Tasks
AI NEWS BLITZ
OpenAI's new GPT-5.6 Sol just took the lead on the DeepSWE coding benchmark.