Thanks, Karo. I have to say, your launch deep dives are some of the most interesting ones I read. You always help me understand what to really pay attention to.
Really interesting(I'm not a dev so excuse the question) , do you know how many lines of code it originally was before the 750k and how much it cost or the token count?
Thanks for this. I love using AI for a multitude of tasks, routine and creative.
And, I don’t understand why we go so insane and give so much attention to these weekly benchmark improvements. They’re not life changing, or revolutionary. Just my take. Thanks.
Thank you so much for taking the time to read and comment dr Jim! 🤗 Yes, the benchmark theater can get ridiculous heheh. Especially when they're cited without context.
But for builders, they're worth watching for two reasons:
1-even small improvements can matter if they make previously annoying workflows reliable enough to use.
2 - those small jumps sometimes reveal a product shift underneath: less babysitting, better verification, etc.
And for me personally, benchmarks are useful because they give me a way to discuss critical AI literacy.
A benchmark is a receipt from one narrow test environment, so I always encourage everyone to test themselves. That’s where it becomes interesting.
awesome stuff Karo. i’m really liking the idea of less shortcuts and more honest. will have to test more, but my initial impressions are they are heading the right direction with this one.
Thank you for reading Toxsec! And yes, I agree, the change shows potential.
But since writing the article, I’ve been testing Opus 4.8 more, and I’ve already caught it doing the classic model thing: sounding very confident while being wrong.
I’ll write more about that soon. How has your experience been? Have you seen the same?
I'd listen to these briefs on the regular. Digestible and comprehensive. Thanks again!
Anthropic continues to feel like that new restaurant down the street that is not disappointing and looks to be the "new spot" for a long time, and 4.8--classic Anthropic--feels like a new dish to try before we've even begun to really appreciate the rest of their knockout menu. Let's eat!
Thank you for reading and taking the time to comment, Rohan! That’s exactly the question I’m circling too: when does the extra autonomy pay for itself?
So far, my threshold is rough, but I’m basing it on how much coordination effort is needed from me. For me, that matters more than execution time.
Thanks, Karo. I have to say, your launch deep dives are some of the most interesting ones I read. You always help me understand what to really pay attention to.
That makes me very happy, thank you so much, Adam! 🤗
Really interesting(I'm not a dev so excuse the question) , do you know how many lines of code it originally was before the 750k and how much it cost or the token count?
Great question! No, I don't know. Buy I'll try to find out! 🤗
Thanks for this. I love using AI for a multitude of tasks, routine and creative.
And, I don’t understand why we go so insane and give so much attention to these weekly benchmark improvements. They’re not life changing, or revolutionary. Just my take. Thanks.
Thank you so much for taking the time to read and comment dr Jim! 🤗 Yes, the benchmark theater can get ridiculous heheh. Especially when they're cited without context.
But for builders, they're worth watching for two reasons:
1-even small improvements can matter if they make previously annoying workflows reliable enough to use.
2 - those small jumps sometimes reveal a product shift underneath: less babysitting, better verification, etc.
And for me personally, benchmarks are useful because they give me a way to discuss critical AI literacy.
A benchmark is a receipt from one narrow test environment, so I always encourage everyone to test themselves. That’s where it becomes interesting.
awesome stuff Karo. i’m really liking the idea of less shortcuts and more honest. will have to test more, but my initial impressions are they are heading the right direction with this one.
Thank you for reading Toxsec! And yes, I agree, the change shows potential.
But since writing the article, I’ve been testing Opus 4.8 more, and I’ve already caught it doing the classic model thing: sounding very confident while being wrong.
I’ll write more about that soon. How has your experience been? Have you seen the same?
yeah a lot of my use cases are security, same experience. x is vulnerable, “are you sure?” “oh yeah no it’s not”.
unfortunate lol. maybe the mythos-class release they are teasing is better, but something tells me it will still have this.
The shift from better answers to better judgment + workflows feels like the real story here, not just another model upgrade
I'd listen to these briefs on the regular. Digestible and comprehensive. Thanks again!
Anthropic continues to feel like that new restaurant down the street that is not disappointing and looks to be the "new spot" for a long time, and 4.8--classic Anthropic--feels like a new dish to try before we've even begun to really appreciate the rest of their knockout menu. Let's eat!
This is the part of the Opus 4.8 shift that feels most important to me:
not just stronger answers, but better judgment around when an answer should be trusted.
The real upgrade is not only capability.
It is uncertainty calibration, verification, effort control, and workflow orchestration.
For builders, that changes how we work with AI.
For education, it matters even more.
A model that can answer quickly is useful.
But a model that can notice flaws, slow down when the cost of being wrong is high, and support better human judgment is where the real value starts.
Speed helps.
Precision decides whether the speed is safe to trust.
Thank you for reading and taking the time to comment, Rohan! That’s exactly the question I’m circling too: when does the extra autonomy pay for itself?
So far, my threshold is rough, but I’m basing it on how much coordination effort is needed from me. For me, that matters more than execution time.
How about you?