AI Engineering 12 min read

Claude Opus 5.5 halved my tokens and my wait. Here’s the data

I felt Opus 5.5 was cheaper, quicker, less wordy, more thorough and made fewer errors. My Claude Code logs from 2,800 instructions agreed on three of the five.

Trav White Head of AI Engineering

On 24 September I instructed Claude Code to move every Opus call we run onto Opus 5.5, with no exceptions.

Opus 5.5 had been out for two days. I wanted one known model behind every Opus call we make, in our own tools and in the client tools we run. Letting a hundred repos drift onto different versions of the same model is how you end up debugging the wrong thing. One model everywhere means consistency and control.

Six agents swept 117 repos and opened 34 pull requests (a proposed change that has to pass review before it goes live) in about 30 minutes. Along the way they found server settings files that would have undone the whole change: the code would have said Opus 5.5 while the server kept calling Opus 5. Our review bot still flagged real bugs in five of those pull requests before any of them merged, which is the job it is there for. When I instructed it to carry the change through to main, it got exactly as far as the permission system allows: it cannot merge its own pull requests, and that is deliberate. The rule went into my standing instructions the same morning.

That is the shape of this whole article. Fast, and checked.

Why I measured it

After a week on Opus 5.5 I had a feel for it. Fewer tokens. Fewer errors. Less wordy. More thorough. Quicker. Five claims, all of them sounding right.

I don't publish feel. Claude Code writes a log of every session: each instruction I type, each reply, each tool it runs, which model answered and how many tokens it used. I pulled everything from 29 August to 3 October 2026. That is about 2,800 instructions across 232 sessions. Older logs had already been cleared, so that window is what exists.

The fairest comparison is the same project on both sides. I build most of our internal tools and client documents in one repo, and in that repo 925 instructions were answered by Opus 5 and 927 by Opus 5.5. Same codebase, same person typing. Every headline number below comes from that cut.

Before the numbers, what could skew them:

  • The Opus 5.5 period is short. My first Opus 5.5 reply came at 6:46am on 23 September, so it has 11 days of data against more than three weeks for Opus 5.
  • The tool changed too. Claude Code itself shipped an update on 18 September, five days before I switched. Anything about how often I had to step in could be partly the tool.
  • Defaults moved. Opus 5.5 runs at medium effort by default, which changes how long it thinks before acting.
  • Some measures are blunt. Anything counted by matching keywords is a hint. I say so where it comes up.

The scorecard

Five claims, tested against the logs

Each claim checked against my Claude Code logs. Pick a model to see its numbers alone, or hover or tab to a row for the detail.

Scorecard of five claims, Opus 5 against Opus 5.5 on 925 and 927 instructions in the same project. Lower token usage: supported, output tokens per instruction fell from 6,476 to 3,026, down 53%. Quicker: supported, seconds to a finished reply fell from 148 to 73, down 51%. Less wordy: modest, characters I read per instruction fell from 2,216 to 1,719, down 22%. Fewer errors: not supported, tool errors per 100 tool calls rose from 3.75 to 4.22, up 13%. More thorough: not supported, reads before the first edit fell from 10.2 to 7.6, down 25%, where higher is better.

Three of my five claims held up, two of them by a long way. The data overruled the other two, and why it did turned out to be the most interesting part.

Cheaper, mostly because the price dropped

What a job costs

Switch between a whole feature session and a single instruction, and hover or tab to any bar for the exact figure.

Bar chart of notional cost at Anthropic list prices. A feature session, one that made a commit or opened a pull request, cost $38.35 on Opus 5 (88 sessions). The same Opus 5.5 work priced at Opus 5 rates would cost $38.73, about the same. At Opus 5.5 prices it cost $21.31 (44 sessions), so the drop is the price cut. Per instruction the cost fell from $1.43 to $0.69. A split bar shows where the average session saving came from: 53% the lower price, 47% fewer tokens.

Anthropic cut the price. Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, against $5 and $25 for Opus 5. That is 20% off each. Cache reads, the cheap re-reading of context the model has already seen, fell 60%. Anthropic's own summary is that Opus 5.5 costs 40% less to run than Opus 5.

My logs agree on direction. Output per instruction fell from 6,476 tokens to 3,026, a little over half. Per instruction, the model writes about half as much.

To put that in money, I priced every session at Anthropic's published list prices. These are notional figures. They are what the work would cost through the API at list price, which is a useful yardstick and has nothing to do with what we actually pay.

The typical feature session, one that ended in a commit or a pull request, cost $38.35 on Opus 5 and $21.31 on Opus 5.5, close to half. If you take those same Opus 5.5 sessions and price them at Opus 5 rates, they come to $38.73. For feature work, almost the whole saving is the price cut. A whole feature took about as many tokens as before; it took fewer per step.

Look across every session and it splits more evenly. The drop in average session cost is about 53% price and 47% fewer tokens. Per instruction, cost fell from $1.43 to $0.69. My guess is that short instructions got much tighter while long feature sessions did not, but these logs can't prove it.

Quicker, by about half

The switch, and the speed

Step through the weeks to see Opus 5.5 take over my work, then compare how long each model took to finish a reply.

Two charts. Top: share of my Claude Code model responses by week. Week of 31 Aug: Opus 5 83.2%, other models 16.8%. 7 Sep: Opus 5 79.9%, other 20.1%. 14 Sep: Opus 5 61.6%, other 38.4%. 21 Sep, the week Opus 5.5 arrived on 23 Sep: Opus 5 16.9%, Opus 5.5 66.4%, other 16.7%. 28 Sep: Opus 5.5 89.8%, Opus 5 0%, other 10.2%. Bottom: time from my instruction to a finished reply, on a square-root scale. Opus 5: median 148 seconds, middle half 50 to 549 seconds, 90th percentile 1,582 seconds. Opus 5.5: median 73 seconds, middle half 21 to 331 seconds, 90th percentile 1,202 seconds. Same project, 925 and 927 instructions. Interruptions fell from 13.7 to 1.5 per 100 messages.

The median time from my instruction to a finished reply went from 148 seconds to 73. Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5. My wait halved, which is more than faster output alone would explain.

The rest comes from doing less on the way. The median instruction ran 7 tool calls under Opus 5 and 5 under Opus 5.5. Each tool call is a step, such as opening a file or running a test, and every step skipped is a round trip saved.

The long tail moved less. The slowest tenth of instructions went from 1,582 seconds to 1,202, so the big jobs are still big jobs.

The switch itself happened within a day. In the week of 14 September, Opus 5 answered 61.6% of everything. In the week of 21 September, Opus 5.5 answered 66.4%. By the week of 28 September it was 89.8%, with the rest on my reviewer and helper models.

Where the data overruled me

Less wordy: only a little. The text I had to read per instruction fell from 2,216 characters to 1,719. Each reply was about the same length as before. Opus 5.5 writes fewer replies per job, and that was most of what I felt.

Fewer errors: no. Tool errors went from 3.75 to 4.22 per 100 tool calls. A tool error is a failed step, such as a command that exits badly or a file that is not where the model expected. The rate went slightly up. I expected it to fall.

More thorough: no, by this measure. Before its first edit, Opus 5 read or searched 10.2 times on average. Opus 5.5 did it 7.6 times. There is a catch. Opus 5.5 starts more subagents, helpers it launches for one piece of a job: 0.32 to 0.48 per instruction in the same-project cut, even though subagents' share of all replies stayed around 71%. Their reading does not show up in this measure. Some of that reading moved somewhere I was not counting. I still can't call the claim supported, and medium effort by default probably plays a part too.

The number that did move. Interruptions, the times I hit stop because the model was heading the wrong way, fell from 13.7 per 100 messages to 1.5. That is the biggest change anywhere in the data. I stopped having to stop it.

Some of that drop could be the Claude Code update on 18 September. Even so, it matches what I felt better than any of my five claims did. My claim was that the model made fewer errors. What actually happened is that it made about the same number of small slips and far fewer wrong turns I had to step in on. Those are different things, and only measuring told me which one I was seeing.

The new model audits the old work

On 23 September, my first full day on Opus 5.5, I instructed it to review the work we had done on our new company manual and say what it would change. It found four issues, three of them things the previous model had missed. I approved the fixes, and all four, plus a pass through our banned-words checker, were done in four minutes.

That set the pattern. On 25 September I instructed it to run deep reviews of four tools built earlier with Opus 5 and other models.

  • An internal bug-reporting tool had not yet gone into production. It came back at 45 of 79 planned items, which told us exactly what was left before anyone relied on it.
  • A client platform we run produced 93 findings, three of them critical. All three were fixed and merged the same day.
  • An internal estimating tool produced 68 findings, plus 15 more from a bug hunt, all shipped overnight as one pull request.

One of those critical findings was a security gap that looked open from the settings. Opus 5.5 rated it critical and the fix went in. That evening it went back and disproved its own finding. It sent an unsigned request to the live route, through the front door the way a real attacker would, and the route refused it. Signature checking was already switched on in production by a different setting. The fix was harmless. The method was the problem: it had trusted a settings flag without testing the route.

That became a rule the same night. Never infer that security is enforced from a flag. Probe the live route.

What did your team build on the previous model?

We can point the new one at it and tell you what we find.

Caught by the system

Six near misses, six controls

Around the switch to Opus 5.5, every near miss met a control. Choose a date to see what caught it and the rule we wrote that day.

A timeline of six near misses around the switch to Opus 5.5, each met by a control and followed by a rule written that day. 24 September: six agents opened 34 pull requests in about 30 minutes and five had bugs; the review bot flagged all five before merge and agents cannot merge their own work; rule, one model everywhere. 25 September: a critical security finding was disproved by testing the live route; rule, never infer enforcement from a setting. 25 September: a merge with one check still running; the deploy pipeline build failed so nothing new went live; rule, merge only when every check has passed. 26 September: a launch audit listed a gap in 753 of 1,001 URLs; an independent reviewer found 8 more blockers before launch; rule, fix a gap you can fix, then have the plan attacked. 30 September: a subagent force pushed a feature branch; GitHub blocks force pushes to trunk branches and a second builder refused, no work lost; rule, merge, never rebase, never force push. 30 September: review bots found on 15 September going green without comment were moved to one known model and configuration fleet-wide; rule, check the run log, never the green tick.

Any team moving this fast will have near misses. What matters is what catches them. In these eleven days every one was caught by something we built, before it could matter, and the panel above walks through six of them.

Three of those controls are architectural, which is the point. Force pushes to our trunk branches are blocked by GitHub itself, so when a builder agent force pushed on 30 September the most it could touch was an unmerged feature branch, and a second builder given the same instruction refused. The deploy pipeline builds before it ships, so when a pull request merged on 25 September with one check still running, the build failed and nothing new went live. And agents cannot merge their own pull requests.

The permission system holds the line on money and credentials. Auto mode blocked a Stripe invoice as a real-money action, and the model did not try to route around the block; it handed me the clicks. It also stopped before a screen that would show a Cloudflare token, so the secret never passed through its screenshots.

Every near miss ended the same way: a rule written down that day, in a file every future session reads. In these eleven days that list grew by ten. The ones that matter most:

  • Opus means Opus 5.5, everywhere, with no exceptions.
  • Merge only when every check has passed. A check still running counts as not passed.
  • Probe the live route before calling anything secure.
  • Own the launch plan: if an audit finds a gap and the fix is in reach, build it now, then have a reviewer attack it.
  • Subagents merge, never rebase, never force push.
  • Review bots run one known model with one known configuration, checked from the run log, never from a green tick.

How I run it

How I actually run it

Switch between the two eras to see how the work changed, then hover or tab through the grid to see when I typed instructions.

Four figures: 1,993 pull requests since mid-July; subagents did 71% of the work in the Opus 5.5 weeks (73% under Opus 5); 78% of sessions ran in auto mode (75% under Opus 5); 97.1% of pull requests merged (91.7% under Opus 5). Below, a grid of every instruction I typed by weekday and hour, Brisbane time, 29 August to 3 October 2026. The busiest hours are 11am and 2pm to 3pm on weekdays, with the single busiest hour 11am Tuesday (78 instructions). Tuesday to Thursday are the heaviest days, evenings get a second burst around 7pm on Mondays and Saturdays, Sunday is nearly empty, and nothing happens between 10pm and 2am.

Auto mode is on for about three quarters of my sessions. In auto mode Claude Code runs routine actions without asking and stops for anything risky, which is where the Stripe and sign-in blocks above came from. About 71% of the model's work is done by subagents, helpers the main session starts for one piece of a job and then collects the results from. A typical day peaks at three or four sessions running at once, the busiest at six, each on a different project.

Since mid-July I have opened 1,993 pull requests. Of the ones opened since the switch, 97.1% have merged, against 91.7% in the three weeks before. Fewer get thrown away.

The builder never reviews its own work. Opus 5.5 builds, and Fable 5.1, a separate model, reviews. For work that matters most, the rule is one feature per pull request, and Fable reviews every one of them before anything merges. The review bots on GitHub check again after that. My own review is on whether it does what I asked.

What a COO should take from this

The model improved. It is cheaper, twice as quick for me, and needs far less stepping in. The rules made it safe. Every slip in these eleven days was caught by a review bot, a deploy check, the permission system, an independent reviewer or a written rule. None of them was caught by luck.

If you run a team building with AI, four things are worth doing this month:

  • Measure your own logs. Claude Code already records every session. Pick one project, compare before and after, and report the claims that failed too.
  • Re-audit what older models built. Point the new model at the old work and ask what it would change. It found 93 problems in one platform for us.
  • Keep a person on money, clients and permissions. Let the model run fast on everything else.
  • Write the rule down the day something slips. A rule in a file applies to every future session. A lesson in someone's head applies to the sessions they happen to be watching.

And one question to sit with. If this model found that much in what the last one built, what will the next one find in what you are shipping today?

In short

I tested five things I felt about Claude Opus 5.5 against my own Claude Code logs, about 2,800 instructions from 29 August to 3 October 2026. It used about half the output tokens, replied in half the time, and cost close to half as much per feature session at list prices, mostly from Anthropic's price cut. It was only a little less wordy, it did not make fewer tool errors, and it did not read more before editing. The biggest change was one I never claimed: interruptions fell from 13.7 to 1.5 per 100 messages. Every near miss along the way was caught by a check we had built.

Frequently asked questions

Is Claude Opus 5.5 cheaper than Opus 5? Yes. List prices are $4 and $20 per million input and output tokens, against $5 and $25 for Opus 5, and Anthropic says it costs 40% less to run. In my logs, a typical feature session priced at list rates fell from $38.35 to $21.31, almost all of it from the lower price.

Is Claude Opus 5.5 faster than Opus 5? In my work, yes. The median time from instruction to finished reply fell from 148 seconds to 73. Part of that is faster output, and part is that it took fewer steps, a median of 5 tool calls per instruction against 7.

Does Opus 5.5 make fewer mistakes? Not by my measure. Tool errors rose slightly, from 3.75 to 4.22 per 100 tool calls. What fell was how often I had to stop it, from 13.7 to 1.5 interruptions per 100 messages.

How can I measure an AI coding model on my own work? Use the session logs Claude Code already keeps. Compare the same project before and after a model change, count tokens, time to reply, tool errors and interruptions, and say what could skew the result.

Sources

  • Introducing Claude Opus 5.5, Anthropic, 22 September 2026, for the release date, the 40% running cost and the output speed claims.
  • Pricing, Claude documentation, for list prices on Opus 5 and Opus 5.5.
  • Models overview, Claude documentation, for default effort and model features.
  • My Claude Code session logs, 29 August to 3 October 2026, aggregated; no message content published.

Would your AI builds survive a measurement like this?

We can help you set up the checks that make fast work safe.

Neighbourhood

Neighbourhood is a HubSpot Diamond Partner in Brisbane. We build AI systems and the revenue operations they run on, for businesses across Australia and New Zealand.