Loading the Elevenlabs Text to Speech AudioNative Player...

Anthropic shipped Claude Sonnet 5 on June 30 and made it the default model on the Free and Pro plans, the version most people using Claude are already on. Opus 5 followed three weeks later, on July 24. It's priced the same as the Opus 4.8 model it replaced, well above Sonnet 5. Anthropic pitched it as the strongest option on Pro and made it the new default on Max, its most expensive plan.

I love Claude models but since Opus 5 shipped I thought: is it worth the extra money over the free model already available? Anthropic's own benchmark charts don't really answer that, since they're built to show a model's ceiling, not what it does with an ordinary day-to-day workload.

So, I built six tasks that mirror that kind of ordinary work. This includes drafting a message, comparing two products, verifying a claim, revising advice after new information, a small coding job, and one odd little image prompt I run on every model I test. I put each one through both models, in fresh chats.

I also ran one task ten times per model for real cost numbers, since the $20 Pro subscription doesn't show token counts, the units Anthropic bills by. And I sent one full pair of answers to two people with no idea which model wrote what, to see whether any of this held up once nobody could see the label.

I Tested 5 AI Coding Tools, and Only Two Didn’t Fail
One of these five tools let me publish a page that leaked every guest’s name and email: Lovable, Replit, Antigravity, Cursor, and Claude Code.

Testing Claude Opus 5 vs Sonnet 5 on six everyday tasks

The first task was a customer-support email about canceling a subscription and asking for a refund. Both models wrote something I'd have sent. Sonnet answered exactly what I asked, but Opus wrote two versions I hadn't requested, adding a useful note about app-store refunds working differently from refunds through the streaming service itself.

Claude Opus 5's two unprompted email drafts
Screenshot: Damilare Odedina / Techloy

On a messy expense spreadsheet with one date that could be read two ways, Opus built the more careful script and flagged its own uncertainty in the code. It still guessed wrong on that date, but Sonnet's less careful guess turned out right.

A two-turn coding-subscription question produced a close result too. I asked both models to pick the best-value coding assistant for a small dev team, then revealed in a follow-up prompt that the team does embedded firmware work, not general app development. Both switched to the same right answer once I corrected them.

Only Opus went back afterward and found a real reason its own advice might fail. It flagged known installation problems with one of its recommended tool's plugins on the team's actual development environment, something neither Sonnet nor I had raised.

Subscribe for free to continue reading this article

Subscribe Subscribe

Already have an account? Log in