SlopCodeBench
SlopCodeBench is a benchmark that scores AI coding agents not on a single pass/fail attempt but on how their code quality decays as they repeatedly extend their own solutions across a sequence of evolving specifications.
Gabriel Orlanski's team at University of Wisconsin–Madison, backed by DARPA, NSF, and Snorkel AI, published the benchmark on March 25, 2026, finding no agent solved any of its 20 problems end-to-end; it resurfaced on July 27, 2026 when HumanLayer's viral Opus 5 benchmark post hit 400 points on Hacker News.
It's a fitness test that keeps re-weighing the same runner after every added mile instead of judging one sprint.
See nascent terms 7 days before everyone, unlock every stage filter, and get weekly early alerts.
Why is it emerging now?
A March 2026 UW-Madison benchmark measuring how coding agents' code quality erodes over iterative tasks went viral on July 27, 2026 after HumanLayer benchmarked Claude Opus 5 against it and posted results to Hacker News.
Search Interest
-
Nascent0–7 days
-
Emergent8–30 days
-
Validating31–90 days
-
Rising ← now91–180 days
-
Established180 days +
Outlook
6-month signal projection and commercial timeline.
Unsaturated leaderboard (best model 28%) plus DARPA/NSF backing gives it staying power as each new flagship model gets benchmarked against it.
Risk · Rival long-horizon benchmarks (AgentWorldBench, ProgramBench) could fragment attention before SlopCodeBench becomes the default citation.
Analogs · SWE-bench · HumanEval · Terminal-Bench
-
nowExplainer content wide open
No dedicated comparison or tutorial content exists yet despite 400-point HN thread.
-
3-6moModel vendors cite scores publicly
Labs start referencing checkpoint solve rates in launch posts, driving comparison-content demand.
-
6-12moErosion metrics enter CI tooling
Verbosity/erosion scoring gets adapted into commercial code-review and agent-QA products.
Competition & Opportunity for term “SlopCodeBench”
Signals derived from the tracked queries, the term's monetization cards, and its cluster neighbors. Heuristic except where marked measured (Google KD).
Ideas for term “SlopCodeBench”
Buildable pitches — turn this term into an article, site, product, post, newsletter, video, or course. Steal any card and run with it.
Zero SERP competition on this exact comparison; developers are actively debating single-shot vs iterative benchmarks post-HN thread.
Definitional explainer capturing search demand spiking off the July 2026 HN post, before mainstream tech media covers it.
Model-comparison content anchored to the live leaderboard; refreshes naturally as new checkpoints get added.
The official leaderboard has no notification layer; a Discord/Slack bot pinging on new entries fills an obvious gap for AI researchers.
The benchmark's deterministic code-quality metrics (complexity concentration, clone rate, dependency entropy) are open-source and reusable outside the benchmark itself.
HumanLayer's write-up already has the visuals (defect trajectories, cost charts); a narrated video recap targets a broader audience than the raw blog post.
Opus 5 wrote five times more functions than its predecessor to gain four extra percentage points of accuracy — and still failed every problem end-to-end.
Best score on SlopCodeBench: 28%. That's the whole pitch — a benchmark frontier labs can't game their way to 90% on yet.
What People Search
Long-tail queries from Google Suggest + Trends. Volume and competition are heuristics — directional, not audited. Content Type comes from query shape.
SERP of term “SlopCodeBench”
What searchers see today — organic results on top, paid ads if anyone's bidding. Ad density is a real-time commercial signal.
FAQ
What is SlopCodeBench?
SlopCodeBench is a benchmark that scores AI coding agents not on a single pass/fail attempt but on how their code quality decays as they repeatedly extend their own solutions across a sequence of evolving specifications.
Why is SlopCodeBench emerging now?
A March 2026 UW-Madison benchmark measuring how coding agents' code quality erodes over iterative tasks went viral on July 27, 2026 after HumanLayer benchmarked Claude Opus 5 against it and posted results to Hacker News.
When did SlopCodeBench emerge?
Publicly emerged around 2026-03-25 (about 139 days ago as of 2026-08-11). EarlyTerms first recorded a pipeline signal on 2026-07-28.
Related Terms
Other terms in the same space — aliases, subtypes, competitors, and neighbors to explore next.
- Part of coding-agents Coding Agents is the category name for AI developer tools that act on code autonomously — reading a repo, planning a change, editing… →
- Part of agentic-coding Agentic coding is the software-development pattern where an autonomous AI agent plans, writes, tests, and iterates on code against a… →
- Competitor programbench ProgramBench is a software-engineering benchmark that tests whether AI agents can reconstruct a complete, working codebase from only a… →
- Competitor agentworldbench AgentWorldBench is an evaluation suite that measures how accurately a language model predicts what happens next inside an agent's… →
- Related ai-slop AI slop is a pejorative for generative-AI content — text, images, video, pull-requests — that is technically fluent but intellectually… →
- Related context-rot Context rot is the measurable degradation in large-language-model output quality as input length grows, even when the prompt stays well… →
- Related agent-harness An agent harness is the middleware between a large language model and the real world — code that runs the agent loop, calls tools,… →
- Related claude-opus-5 Claude Opus 5 is Anthropic's cost-efficient frontier LLM, priced at $5/$25 per million tokens — half of flagship Claude Fable 5 — built… →
- Related long-running-agents Long-running agents are AI agents designed to sustain work across multiple context windows, persisting state through structured… →
- Related
Sources
Primary URLs this report cites — open any to verify the claim yourself.
- 01 SlopCodeBench paper (arXiv) arxiv.org ↗
- 02 SlopCodeBench full text (HTML) arxiv.org ↗
- 03 SprocketLab/slop-code-bench — official GitHub repo github.com ↗
- 04 Live leaderboard — scbench.ai scbench.ai ↗
- 05 Snorkel AI — Measuring code erosion as agents iterate snorkel.ai ↗
- 06 Snorkel AI — SlopCode Bench leaderboard page snorkel.ai ↗
- 07 HumanLayer — Benchmarking Opus 5 on SlopCodeBench github.com ↗
- 08 Hacker News thread — Benchmarking Opus 5 on SlopCodeBench news.ycombinator.com ↗