Livenerf: Has Claude Opus 5.5 Been Nerfed Yet?
On September 22, 2026, Claude Opus 5.5 launched. Eight days later, the familiar post-release question was already back: did the model get nerfed? The honest answer, as of September 30, is that nobody has enough evidence yet. The latest Livenerf update, posted September 29, reports six completed days of a 30-day series. Days 1–10 form the baseline, while days 11–20 and 21–30 are decision windows, putting the first possible call around October 24. (anthropic.com)
That waiting period is the point. Livenerf is a public benchmark, meaning a repeatable collection of tests used to measure a system over time. Instead of treating a disappointing answer or a handful of screenshots as proof, it asks a narrower question: does the same served version of Claude Opus 5.5 perform differently after launch?
“Nerfed” means more than a smaller model
People use nerfed to describe several different things. The company might route requests to a different model, lower the amount of reasoning used, change safety filters, alter the client, or replace the model snapshot behind the same name. A user experiences all of those as a change in quality, even though only one involves changing the model itself.
The official Opus 5.5 documentation makes this distinction especially important. Adaptive thinking is always enabled, the default effort level is medium, and the model behaves differently from Opus 5 even when an application changes no code. Anthropic’s help documentation also describes automatic fallback, where some requests can move from Opus 5.5 to Opus 5 or Opus 4.8 after a safety classifier intervenes. (platform.claude.com)
Livenerf measures one particular serving path: Opus 5.5 running through headless Claude Code on a Claude Max subscription. Headless means the command-line tool runs without an interactive desktop session. The protocol fixes effort at high, pins the Claude Code version, and rejects samples that were retried, refused, or partly served by another model. That does not describe every Claude user, but it creates a clear target instead of pretending that one test can measure every deployment. (github.com)
A frozen exam beats a pile of anecdotes
The heart of the project is a frozen panel. Think of it as an exam whose questions are locked in a drawer before the experiment starts. The panel contains 78 questions selected from GPQA Diamond, MMLU-Pro, competition-math sets, and AIME problems. They came from a much larger screening pool of 2,336 questions, with the easiest and hardest items mostly removed because questions the model always answers the same way cannot reveal much change.
Each day, the same panel is run again. That matters because a fresh random test can accidentally be easier or harder than yesterday’s test. With a frozen panel, each question is compared with itself over time. The basic idea looks like this:
baseline[item] = mean(score[item] for day in 1..10)
window[item] = mean(score[item] for day in 11..20)
delta = mean(window[item] - baseline[item] for item in panel)
This is conceptual pseudocode, not a command from the repository. The important detail is the pairing: a difficult question remains difficult in both periods, so item difficulty drops out of the comparison. The control arm also runs Opus 5 on the same harness and GPQA questions. If both models move together, the explanation may be a platform or harness change rather than an Opus 5.5 regression. (github.com)
The scoring is intentionally unglamorous. Answers are checked with fixed rules, such as a known answer key or an exact mathematical result. There is no second language model deciding whether the first model sounded correct. That would add another moving part, and a drifting judge could manufacture a result as easily as a drifting model.
The plumbing is part of the experiment
Repeatability often fails in the details nobody notices. Livenerf pins Claude Code to version 2.1.280, records a harness code hash, uses a fixed system prompt, and runs in an empty working directory without tools, Model Context Protocol servers, settings files, hooks, or automatic memory. Those restrictions make the test less like a normal coding session, but they also make a later score easier to interpret.
The logs preserve more than the final answer. They record the model identity, timing, usage, and output-token count. Tokens are the small pieces of text a language model processes and generates, so output length can act as a rough clue about how much reasoning or explanation occurred. In the project’s pre-series validation, lowering effort reduced output tokens by about 62%, while the accuracy signal was noisier. That makes token count a useful warning light, but not a verdict on capability.
Statistics decide, not one strange afternoon
Before the series began, Livenerf recorded its decision rule in a pre-registration. A pre-registration is a public plan written before the result is known. Under the current rule, a change must clear a 99% interval in two consecutive 10-day windows, move by at least three percentage points, use the same harness identity, and keep the sample error rate below 5%.
Those requirements are deliberately demanding. The instrument’s estimated minimum detectable effect is about 7.5 percentage points per 10-day window, so a small decline may remain invisible. The validation also could not distinguish an Opus 5 swap from Opus 5.5 at the 99% level in one batch of samples. That limitation is valuable to publish: a benchmark earns trust by describing what it can miss, not only what it hopes to catch.
So, has Opus 5.5 been nerfed?
Not according to Livenerf, because Livenerf does not have a post-baseline result yet. As of September 30, 2026, it is still collecting the reference period against which later runs will be judged. Saying “no nerf detected” would already be too strong; the accurate statement is “not enough data for a verdict.”
That restraint is the most interesting part of the project. If a sustained regression appears, the paired score, token trend, classifier-event rate, and control arm should help narrow down what changed. If the numbers stay inside the noise, the result will be less exciting but more useful than a viral complaint: no measurable change under this exact test. Either way, Livenerf turns a vague feeling about Claude Opus 5.5 into a record that can be checked day by day.
Comments (0)
No comments yet. Be the first to respond!
Leave a Comment
Your comment will be visible after review.