EarlyTerms

SlopCodeBench

Rising · Emerged · 139 days old · Last reviewed
Search / mo
~365 /mo
Competition KD
Stage
Rising
measured 2026-08-10 sources · 8

SlopCodeBench is a benchmark that scores AI coding agents not on a single pass/fail attempt but on how their code quality decays as they repeatedly extend their own solutions across a sequence of evolving specifications.

Gabriel Orlanski's team at University of Wisconsin–Madison, backed by DARPA, NSF, and Snorkel AI, published the benchmark on March 25, 2026, finding no agent solved any of its 20 problems end-to-end; it resurfaced on July 27, 2026 when HumanLayer's viral Opus 5 benchmark post hit 400 points on Hacker News.

It's a fitness test that keeps re-weighing the same runner after every added mile instead of judging one sprint.

EarlyTerms Pro

See nascent terms 7 days before everyone, unlock every stage filter, and get weekly early alerts.

Why is it emerging now?

TL;DR

A March 2026 UW-Madison benchmark measuring how coding agents' code quality erodes over iterative tasks went viral on July 27, 2026 after HumanLayer benchmarked Claude Opus 5 against it and posted results to Hacker News.

5 forces driving coverage — scroll →

Search Interest

peak ~365/mo
updated 2026-08-10
~365/mo ~182/mo 0
2026-07-12 2026-07-27 2026-08-10
Term Lifecycle
  1. Nascent
    0–7 days
  2. Emergent
    8–30 days
  3. Validating
    31–90 days
  4. Rising ← now
    91–180 days
  5. Established
    180 days +

Outlook

6-month signal projection and commercial timeline.

Signal medium
Revenue moderate

Unsaturated leaderboard (best model 28%) plus DARPA/NSF backing gives it staying power as each new flagship model gets benchmarked against it.

Risk · Rival long-horizon benchmarks (AgentWorldBench, ProgramBench) could fragment attention before SlopCodeBench becomes the default citation.

Analogs · SWE-bench · HumanEval · Terminal-Bench

Monetization timeline
  1. now
    Explainer content wide open

    No dedicated comparison or tutorial content exists yet despite 400-point HN thread.

  2. 3-6mo
    Model vendors cite scores publicly

    Labs start referencing checkpoint solve rates in launch posts, driving comparison-content demand.

  3. 6-12mo
    Erosion metrics enter CI tooling

    Verbosity/erosion scoring gets adapted into commercial code-review and agent-QA products.

Competition & Opportunity for term “SlopCodeBench”

Signals derived from the tracked queries, the term's monetization cards, and its cluster neighbors. Heuristic except where marked measured (Google KD).

Content Gap
1 queries tracked
Led by General (1)
1 Suggest-only tails — long-tail opening
Revenue Potential
0% commercial-intent queries
2 monetization angles mapped
Mostly informational — pre-commercial
Build Difficulty
High (heuristic)
Stage: rising — late entry — verify the gap first
0 / 12 default TLDs taken
9 related terms already published
Heuristic · signals: tracked queries, term monetization cards, cluster neighbors

Ideas for term “SlopCodeBench”

Buildable pitches — turn this term into an article, site, product, post, newsletter, video, or course. Steal any card and run with it.

Article
SlopCodeBench vs SWE-bench: Why 'Solving a Problem' Isn't the Same as 'Not Making It Worse'

Zero SERP competition on this exact comparison; developers are actively debating single-shot vs iterative benchmarks post-HN thread.

Article
What Is SlopCodeBench? The Benchmark Measuring AI Code Rot

Definitional explainer capturing search demand spiking off the July 2026 HN post, before mainstream tech media covers it.

Article
SlopCodeBench Leaderboard Explained: How Opus 5, GPT-5.5, and Kimi K2.6 Stack Up

Model-comparison content anchored to the live leaderboard; refreshes naturally as new checkpoints get added.

Product
A SlopCodeBench score-change alert bot for new model releases

The official leaderboard has no notification layer; a Discord/Slack bot pinging on new entries fills an obvious gap for AI researchers.

Product
An erosion/verbosity linter that reuses SlopCodeBench's 41 quality metrics on private repos

The benchmark's deterministic code-quality metrics (complexity concentration, clone rate, dependency entropy) are open-source and reusable outside the benchmark itself.

Video
I Watched Opus 5 Write Itself Into a Corner on SlopCodeBench — 6-Hour Livestream Recap

HumanLayer's write-up already has the visuals (defect trajectories, cost charts); a narrated video recap targets a broader audience than the raw blog post.

Post HN / r/MachineLearning
The Benchmark That Finally Measures Vibe Coding's Real Cost

Opus 5 wrote five times more functions than its predecessor to gain four extra percentage points of accuracy — and still failed every problem end-to-end.

Post Newsletter / LinkedIn
Why an Unsaturated Benchmark Is More Valuable Than a Saturated One

Best score on SlopCodeBench: 28%. That's the whole pitch — a benchmark frontier labs can't game their way to 90% on yet.

What People Search

Long-tail queries from Google Suggest + Trends. Volume and competition are heuristics — directional, not audited. Content Type comes from query shape.

Keyword
Competition
Content Type
slopcodebench
Very Low
General
Updated 2026-08-10 · sources: Google Trends, Google Suggest · Competition is heuristic

SERP of term “SlopCodeBench”

What searchers see today — organic results on top, paid ads if anyone's bidding. Ad density is a real-time commercial signal.

FAQ

What is SlopCodeBench?

SlopCodeBench is a benchmark that scores AI coding agents not on a single pass/fail attempt but on how their code quality decays as they repeatedly extend their own solutions across a sequence of evolving specifications.

Why is SlopCodeBench emerging now?

A March 2026 UW-Madison benchmark measuring how coding agents' code quality erodes over iterative tasks went viral on July 27, 2026 after HumanLayer benchmarked Claude Opus 5 against it and posted results to Hacker News.

When did SlopCodeBench emerge?

Publicly emerged around 2026-03-25 (about 139 days ago as of 2026-08-11). EarlyTerms first recorded a pipeline signal on 2026-07-28.

Related Terms

Other terms in the same space — aliases, subtypes, competitors, and neighbors to explore next.

Explore next
Also mentioned
  • Related SWE-bench

Sources

Primary URLs this report cites — open any to verify the claim yourself.

  1. 01 SlopCodeBench paper (arXiv) arxiv.org
  2. 02 SlopCodeBench full text (HTML) arxiv.org
  3. 03 SprocketLab/slop-code-bench — official GitHub repo github.com
  4. 04 Live leaderboard — scbench.ai scbench.ai
  5. 05 Snorkel AI — Measuring code erosion as agents iterate snorkel.ai
  6. 06 Snorkel AI — SlopCode Bench leaderboard page snorkel.ai
  7. 07 HumanLayer — Benchmarking Opus 5 on SlopCodeBench github.com
  8. 08 Hacker News thread — Benchmarking Opus 5 on SlopCodeBench news.ycombinator.com
Opportunity radar
More terms breaking out right now
View →