July 25, 2026

Claude Opus 5 Is Here: Anthropic's Best Everyday Model, and the Reviewers Are Split

Anthropic's Claude Opus 5 is out, and the early reviews are unusually split. What changed, what testers like Claire Vo and Ethan Mollick found, and my own four-model test.
Daan van Rossum
By
Daan van Rossum
Founder & CEO, Lead with AI

Presented by

Anthropic has released Claude Opus 5, and it is available now on every plan and platform.

Opus is the tier of Claude for long analysis, multi-step projects, building things. Opus 5 replaces Opus 4.8 as the new default model on Claude Max and the strongest model available on Claude Pro.

Anthropic's framing is that Opus 5 "comes close to the frontier intelligence of Claude Fable 5 at half the price," but the initial reviews do not seem to share that excitement.

Still, from its makers perspective, Fable 5 is their most capable model, while Opus 5 is the one they expect you to use every day.

Flagship AI Newsletter
The AI Newsletter That Makes You Smarter, Not Busier
Join over 30,000 leaders and receive our insights on AI platforms, implementations, and organizational change management.
FlexOS Course - AI Content Accelerator - Testimonial Badge

What Anthropic Says Changed

Anthropic says Opus 5 is best-in-class on its coding and professional knowledge-work evaluations.

Some other standout capabilities, as per Anthropic: 

It thinks before it answers, by default. On Opus 4.8 you could switch extended reasoning on, as we showcase in our lesson on Advanced Models. On Opus 5 it is always on, and what you control instead is an Effort setting with five levels: Low, Medium, High (default), Extra, and Max.

Anthropic's technical notes say effort matters more here than on any previous Opus model, so the same question can produce a meaningfully different answer depending on where that dial sits. Higher effort also burns your usage limits faster.

It reads a million tokens at once. The context window is 1 million tokens, as both the default and the maximum, with quality claimed to hold across the whole context window. If that holds up in real use, you can drop in a full year of board papers rather than feeding them in chapters.

It checks its own work without being told to. Anthropic's docs instruct developers to delete the "include a final verification step" lines they wrote for older models, because leaving them in makes Opus 5 over-verify.

It talks more, and it writes longer. Anthropic says default responses and written deliverables run longer, and the model narrates its progress more often, although in my initial tests this wasn't the case. Other reviewers were also less pleased about it.

It is Anthropic's most aligned model so far, scoring lowest of any recent Claude model on Anthropic's own automated audit of misaligned behavior. (Although this hasn't been verified by objective others.)

The price did not move. What is good is that Opus 5 costs exactly what Opus 4.8 did, and on Claude Pro and Max it's even included, although usage limits continue to apply as always.

Why the Early Testers All Say the Same Strange Thing

A handful of people had the model for a week before general release, and their reviews rhyme. They are impressed by the output and irritated by the experience, which is not the pattern you would predict from the launch post.

Claire Vo, host of the How I AI podcast, put it bluntly: she hates working with it. She calls the model "neurotic" and "apologetic," and was especially critical about the verbosity, which she names "Claude slop." Then (hilariously) the leaderboard came in and Opus 5 won, with top scores on front-end design and prototyping.

Her summary: it is her most loathed colleague, and it does the best work.

Claire Vo’s leaderboard reveal, where Opus 5 wins her blind test

Dan Shipper of Every called it a hard model to love. His team found it argued with instructions, stopped before the work was finished, and did not play well with the skills and plugins they had built for earlier models. Then they deleted those skills and started from scratch, and it got dramatically better. His colleague Kieran Klaassen found the same from another angle: lower thinking levels produced better results than higher ones.

Zapier CEO Wade Foster reported that Opus 5 took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the owner, summarizing for retention operations. Previous models did not pass this test, but Opus 5 scored 100%.

Ethan Mollick, the famed Wharton AI professor, described it as a good model if a quirky one: it could match or beat Fable 5 on shorter tasks, but seemed less ambitious on longer ones and would not deliver as complete a set of work.

As a bottom line from these field reports I'd say to 1) temper your enthusiam for this 'groundbreaking new model' and 2) have a critical look at old prompts, custom instructions, projects, and skills that may not work as desired anymore.

The fastest path to a good result with Opus 5 might be deleting them and starting clean, which is a pain and a warning sign of how quickly things can change and how unreliable these platforms are for long-term efforts you invest your time and money into.

My Own Test: Four Models, One Real Business Problem

I gave Opus 5 and GPT 5.6 Sol (both on two levels of thinking/effort) the same live problem from our business: based on all you know about us, create a scalable transformation product that produces a real result within 30 days.

Opus 5 versus GPT 5.6-Sol

All four come up with pretty much the same idea (which I found surprising), but Opus 5 (Extra) gave the best strategic answer, while GPT-5.6 Sol (Pro) gave the answer I would actually operate from.

Sol anticipated how I work and what I would need to turn an idea into a plan, and that was more useful than the answer that followed my prompt most literally. (A fair test, I think, because I use Claude and ChatGPT roughly equally.)

Some details on the test:

Four models, one real strategy problem
Model Strongest contribution Where it fell short Verdict
Opus 5 (Extra) Made the sharpest strategic choice. Narrowed to one specific customer group, tied the first small deployment to a larger opportunity, and proposed a simple one-page artifact to prove the result Thin on how to actually execute it Best strategy
GPT-5.6 Sol (Pro) Broke the business into clear component parts, defined the stages work moves through, gave a concrete six-week build plan, and named the points where human judgment stays Less decisive about who to serve first Most usable
GPT-5.6 Sol (High) The strongest operating manual: standards for what counts as done, SOPs, guardrails, and a structured log of results to learn from Qualified almost anyone as a customer, so it made no real market choice Best process
Opus 5 (Max) The best thinking on adoption, including a requirement that other people actually use the thing, and a curated pattern library with fit criteria and security tiers An unrealistic shipping cadence and a scope that kept expanding Best on adoption

My conclusion is that as always, using multiple models will remain a must for real advantages.

In my ideal scenario I'd combines two of them: Opus 5 (Extra) for the customer segment, the result structure, and the enterprise expansion logic, then GPT-5.6 Sol (Pro) for the workflow stages, delivery ideas, human review frameworks, and the program structure.  

The Bottom Line: Claude Opus 5

Opus 5 as well as GPT 5.6-Sol are incredibly capable models, and most of us will barely leverage them to their fullest abilities (see 'the capability overhang'.) So go ahead and put it to the test:

  • Try it on one hard thing this week. Not a summary or an email, but your biggest strategic problem. Compare it against what the cheaper model produces.
  • If the output of existing skills and prompts feels off, delete your setup and start clean. The early testers are consistent on this, so it's worth checking and correcting.
  • Try lower effort before higher. Leave Effort on High for most work and move to Extra or Max only when the problem warrants it.

Of course, none of this changes what actually creates value, which is your own AI fluency and your judgment about where to point these tools.

The models keep getting better at producing the answer, but deciding which question is worth asking, and what to do with it, is still entirely yours.