BREAKING
GPT-5.6 Sol Tops Claude on DeepSWE
DeepSWE Pass@1 Scores
Sol
73
Terra
70
Fable 5
70
Luna
67
0
$
Sol avg cost
0
$
Fable 5 avg cost
0
%
Sol Ultra
0
%
Sol base
0
%
Mythos 5
Praised but Not Flawless
Strengths
●
Persistent on long tasks
●
Highly token-efficient
●
Praised by Cursor and Notion
Concerns
●
Basic mistakes reported
●
File-deletion risk
●
System card noted rare cheating
Sol Leads, Fable Still Fits Hard Tasks
AI NEWS BLITZ
OpenAI's new GPT-5.6 Sol just took the lead on the DeepSWE coding benchmark.