Devin maker Cognition on July 8, 2026 released version 1.1 of FrontierCode, its benchmark for measuring the real-world quality of AI-generated code. The update centers on clearer guidelines for fair internet use and refined grading criteria. Details are laid out on the company blog .
July 8, 2026 · Cognition
FrontierCode 1.1: grading AI code like a tech lead
The benchmark asks not "does it pass tests?" but "would maintainers actually merge this?" The 1.1 update sharpens fair-use rules and grading — while even the best models still land in the low teens.
Top model score — still far from production
Even leading models clear only a fraction of the production-quality bar.
81%
fewer misclassification errors vs SWE-Bench Pro
<1%
unfair-use rate after new prompt + verifier
75
overly strict blocker criteria relaxed to non-blockers
Subset structure, revised in 1.1
The noisiest tier was retired; reporting now covers two subsets.
Extended
150
tasks · reported
Main
100
tasks · reported
Diamond
50
hardest · retired (noisy)
How grading works
A pass requires clearing every blocker — quality is a rubric-weighted average.
Blocker criteria
Mandatory to merge. Any unmet blocker → score of zero.
Non-blockers
Quality signals, not required — weighted into the average.
Fair-use detection uses prompts + verifiers (not blocklists) — chosen for evasion resistance and scalability. Documentation is fine; upstream PR diffs and solution mirrors are not.
✓ What developers praise
Practical framing — "would you actually merge this code." Maintainers from Uppy, Mattermost and Budibase say it grades "like a tech lead" and respects subjective quality.
! What critics note
Grading is strict — top scores near 13% suggest models remain far from production quality. Some observers call for independent verification.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…