13 Comments
User's avatar
Adam M.J's avatar

Thanks, Karo. I have to say, your launch deep dives are some of the most interesting ones I read. You always help me understand what to really pay attention to.

Karo (Product with Attitude)'s avatar

That makes me very happy, thank you so much, Adam! 🤗

Law's avatar

Really interesting(I'm not a dev so excuse the question) , do you know how many lines of code it originally was before the 750k and how much it cost or the token count?

Karo (Product with Attitude)'s avatar

Great question! No, I don't know. Buy I'll try to find out! 🤗

Dr Jim Polk's avatar

Thanks for this. I love using AI for a multitude of tasks, routine and creative.

And, I don’t understand why we go so insane and give so much attention to these weekly benchmark improvements. They’re not life changing, or revolutionary. Just my take. Thanks.

Karo (Product with Attitude)'s avatar

Thank you so much for taking the time to read and comment dr Jim! 🤗 Yes, the benchmark theater can get ridiculous heheh. Especially when they're cited without context.

But for builders, they're worth watching for two reasons:

1-even small improvements can matter if they make previously annoying workflows reliable enough to use.

2 - those small jumps sometimes reveal a product shift underneath: less babysitting, better verification, etc.

And for me personally, benchmarks are useful because they give me a way to discuss critical AI literacy.

A benchmark is a receipt from one narrow test environment, so I always encourage everyone to test themselves. That’s where it becomes interesting.

ToxSec's avatar

awesome stuff Karo. i’m really liking the idea of less shortcuts and more honest. will have to test more, but my initial impressions are they are heading the right direction with this one.

Karo (Product with Attitude)'s avatar

Thank you for reading Toxsec! And yes, I agree, the change shows potential.

But since writing the article, I’ve been testing Opus 4.8 more, and I’ve already caught it doing the classic model thing: sounding very confident while being wrong.

I’ll write more about that soon. How has your experience been? Have you seen the same?

ToxSec's avatar

yeah a lot of my use cases are security, same experience. x is vulnerable, “are you sure?” “oh yeah no it’s not”.

unfortunate lol. maybe the mythos-class release they are teasing is better, but something tells me it will still have this.

Petar Dimov's avatar

The shift from better answers to better judgment + workflows feels like the real story here, not just another model upgrade

YT's avatar

I'd listen to these briefs on the regular. Digestible and comprehensive. Thanks again!

Anthropic continues to feel like that new restaurant down the street that is not disappointing and looks to be the "new spot" for a long time, and 4.8--classic Anthropic--feels like a new dish to try before we've even begun to really appreciate the rest of their knockout menu. Let's eat!

QUASAR's avatar

This is the part of the Opus 4.8 shift that feels most important to me:

not just stronger answers, but better judgment around when an answer should be trusted.

The real upgrade is not only capability.

It is uncertainty calibration, verification, effort control, and workflow orchestration.

For builders, that changes how we work with AI.

For education, it matters even more.

A model that can answer quickly is useful.

But a model that can notice flaws, slow down when the cost of being wrong is high, and support better human judgment is where the real value starts.

Speed helps.

Precision decides whether the speed is safe to trust.

User's avatar
Comment removed
May 30
Comment removed
Karo (Product with Attitude)'s avatar

Thank you for reading and taking the time to comment, Rohan! That’s exactly the question I’m circling too: when does the extra autonomy pay for itself?

So far, my threshold is rough, but I’m basing it on how much coordination effort is needed from me. For me, that matters more than execution time.

How about you?