Real-SWE Benchmark - Fable 5.1 Solves 38.8% of Private Enterprise Code, and Public Coding Benchmarks May Be Leaking

Specific Labs scored eight frontier coding agents on 10 private enterprise tasks the models have never seen. Fable 5.1 leads at 38.8%; every score is below the SWE-bench Verified floor. The 32-point gap is the empirical size of training-data recall.

September 17, 2026 - Specific Labs published Real-SWE the week of September 11-13: 10 private enterprise tasks (billing, tax, data migration) with the production stack those tasks actually run on (AWS simulators, Docker, Kubernetes, PostgreSQL, MongoDB, Slack, Email). 640 scored runs total, 8 per task averaged. The code is licensed from real companies - which is the contamination-resistant bit. SWE-bench Verified is public benchmark code that lives in training data. Real-SWE is private enterprise code that does not.

Yesterday’s post covered five frontier launches in ten days. Those same models just got benchmarked on code they have not memorized - and the leader drops to 38.8%.

The ranking: 8 models, 10 tasks, 640 runs

Model Score
Fable 5.1 + Claude Code 38.8%
GPT-6 Astra + Codex CLI 33.8%
Gemini 3.8 Flash + Gemini CLI 31.2%
GLM 5.3 28.8%
Grok 4.6 / Muse Spark 1.3 (tied) 23.8%
Kimi K3 18.8%
GPT-5.6 SolCodex CLI 16.2%

Every score on this list is below the SWE-bench Verified leaderboard floor. The best frontier coding agent on private enterprise code fails 60%+ of the time, per The New Stack (September 15). The same models that headline-fronted the September rush - Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash - drop to 31-39% when they cannot memorize the answer.

For 6 of the 10 tasks, every model scored below 15%. No model cleared 15% on a majority of tasks. That is the kind of number that breaks a procurement argument built on public benchmark scores.

The 32-point gap: SWE-bench Verified vs Real-SWE

The SWE-bench Verified leaderboard sits above 70% for the top models. Real-SWE’s best score is 38.8%. The 32-percentage-point gap is the empirical size of training-data recall on public benchmark tasks. Yesterday’s post mentioned the 119x price spread across the top 15 models; today the calibration number is the gap between what these models recall and what they generate.

The same model is being judged twice. Once on code it has seen. Once on code it has not. Public benchmark scores are recall scores. Private benchmark scores are generation scores. Procurement arguments built on public scores are recall arguments - and the recall argument is the weaker argument the moment a buyer asks for a private-repo number.

The contamination argument

Ken Ashe made the structural problem explicit: “Every issue, every repo, every fix lives on GitHub, which means it lives in the training data of basically every frontier model.” That is the argument Real-SWE is built on. SWE-bench Verified scores are inflated by training-data recall on the public benchmark tasks. Real-SWE’s licensed private code is contamination-resistant by design - the closed-book exam the open-book exam never was.

The Real-SWE methodology writeup is not yet public. Treat the setup as confirmed by Specific Labs’ own disclosure, not third-party verified. The 32-point gap is the empirical size of recall on the benchmark tasks. It is not the size of all contamination - other vectors (documentation, blog posts, Stack Overflow answers) are not measured here.

Cost-per-solved-task: Fable 5.1 is not the efficiency winner

The structural argument collapses without the cost row. Fable 5.1’s per-task cost on Real-SWE exceeded Gemini 3.8 Flash’s $2.50/task reference. Pay roughly 4x more for the leader and you get 7.6 percentage points more. The cost-per-solved-task ratio is the metric that matters for procurement, and Fable is not the efficiency winner.

The procurement question flips: which number are you buying - the 95% public score or the 38.8% private score - and at what per-task cost? A model that headlines at 95% on SWE-bench Verified and drops to 31-39% on Real-SWE is being judged twice. Once for capability it can recall, once for capability it has to generate. Yesterday’s launch narrative does not survive that second judgment intact.

What changes for builders

If you’re scoring coding agents on private repos. Real-SWE is the first benchmark specifically designed to test contamination resistance. If you have private code that has not been pushed to GitHub, you have a benchmark target the frontier models have not memorized. Use it. The 38.8% leader is the ceiling, not the floor, for any agent you intend to deploy on private codebases.

If you’re building a procurement argument. Public benchmark scores are now split into two tracks. SWE-bench Verified and SWE-bench Pro measure recall-on-public-code. Real-SWE measures generate-on-private-code. The same model will score differently on each, and the procurement conversation needs to name which one the budget is buying. A vendor that quotes a 95% public score and declines to share a private-repo score is signaling which number they want you to see.

If you’re tracking the cost curve. The cost-per-solved-task ratio is the structural metric. Gemini 3.8 Flash at $2.50/task is the reference point. Fable 5.1’s higher per-task cost is justified only if the 7.6-point accuracy gain compounds across enough tasks to clear the cost premium. For most enterprise workloads - billing, tax, data migration - that arithmetic does not close.

If you’re maintaining SWE-bench-style harnesses. This is the August 10 follow-up. The SWE-bench Pro harness-variance post argued that harness design inflates scores on public benchmarks. Real-SWE adds the contamination half. A harness on public code measures recall plus harness variance. A harness on private code measures generation alone. The two failure modes are different and require different mitigation.

What this does not touch

Methodology writeup not yet public. Specific Labs has not published the Real-SWE methodology paper. The 38.8% / 33.8% / 31.2% numbers are from Specific Labs’ own disclosure and from Winzheng and Pebblous’s reporting on the disclosure. Treat them as the official benchmark result, not as third-party verified. The contamination framing is consistent across Ken Ashe, Winzheng, and Pebblous - which is the multi-source confirmation available, not the academic peer review the data would normally require.

Contamination beyond training-data recall. The 32-point gap is the empirical size of training-data recall on the benchmark tasks. It is not the size of all contamination. Other contamination vectors - documentation, blog posts, Stack Overflow answers, public discussions of the tasks - are not measured here. The 32-point figure is a floor on the contamination effect, not a ceiling.

Public benchmark quality as a category. Real-SWE is one private-repo benchmark. It does not retire SWE-bench Verified. Public benchmarks remain useful for relative ranking among models. They are not useful for absolute claims about coding-agent capability on private enterprise code. The argument is not “stop using public benchmarks” - it is “stop using public benchmarks as the sole evidence in a procurement conversation.”

The other 5% of the model story. Yesterday’s post covered cyber gating, architectural efficiency, the 119x price spread, and Anthropic’s cancelled price hike. Real-SWE is the empirical mirror - it is what those same models look like on private code. The launch narrative and the benchmark narrative are not in conflict; they are two frames on the same set of models.

The structural read

The model layer is racing to zero margin faster than the routing layer can capture it. Yesterday’s post framed that as a price-curve story. Today’s post frames it as an evaluation story. The two are the same story: every procurement argument that survives on a 95% public score is going to fail when the buyer asks for a private-repo number.

The first lab or vendor that publishes both numbers on the same page - public score, private score, cost-per-solved-task included - wins the procurement conversation for the next two quarters. Specific Labs just gave them the template. The September model rush did not change the math; it changed the visibility of the math. The buyers who have not yet asked for a private-repo benchmark score are about to.

Sources