Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0].

> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]

On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].

0: https://support.claude.com/en/articles/15425996-data-retenti...

1: https://www.anthropic.com/news/claude-opus-5

2: https://xcancel.com/arcprize/status/2064399134099153344

 help



Also the cost per task. It appears to be significantly cheaper, cheaper than sonnet!

The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3.

https://artificialanalysis.ai/?cost=cost-per-task


I don't understand how the K3 numbers keep coming out cheap for people. I recently started to add it to my security auditing benchmarks and found it was going to cost about twice as much as Opus 4.8. It blew through the $100 budget I'd set at like 11%. In the tasks I'm doing it seems crazy expensive because it chews so much, burning a tremendous amount of tokens.

I think the way people usually compare pricing is fundamentally flawed. You can't compare token prices because different models use different tokenizers, and you can't compare tokenizer-normalized token prices because different models at different settings use more or fewer tokens to complete the same task at a different level of quality.

Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.


I got the $19 plan, and it's anemic. One tiny task blew through the 5-hour budget and 19% of the weekly budget. A completely useless amount of usage. OpenAI's $20 plan feels like 100x more generous (I don't think I'm exaggerating here). Someone in another thread said their plans are cheaper in China, maybe that's the difference, I dunno.

But, I'm finding Kimi K3 terrifyingly expensive in the way that Fable and GPT 5.5 Pro are at token rates. Not as expensive as those, but expensive enough to where if you don't put a budget cap on it, you might wake up bankrupt if you leave a task running overnight. Not because of the per-token cost, but because how many tokens it's going to burn.


I have the second largest Kimi plan, the Chinese version. When K2.6 was their latest model, the quota was good; it was like GPT $100 is now or what the $20 version was in December.

When K2.7 was released, they cut quota by 80%. I can't tell how much they have further cut it after the K3 release because it's barely worth using at all. I just use it in my model router since I have the annual plan paid for.

It's just not a serious model or company.


In the $19 plan, I've been able to reverse engineer both an android APK and firmware (in Ghidra and Radre) for a baby rocker and build a quick PoC application in my session limit. And then further refined the app in another session at another point in time without leaving Opus. I dont consider that to be a tiny task. How are you blowing through your usage?

I have no idea. Seems like normal stuff. I used Kimi Code with K3 to add support for Kimi Code to flar (https://swelljoe.com/post/i-let-every-agent-implement-its-ow...), a task I've done with almost every major model/agent combo. Most show up as a blip on the usage chart...it's basically usually one file, a README update, and adding the agent name to the CLI.

Then, I added it to my benchmark of security vulnerability auditing capability, and it burned a bazillion tokens, burned through the 5-hour limit, burned through $100 in extra usage I'd allocated, and was only 11% finished. That's more expensive than any model I've tested other than GPT 5.5 Pro on this task.

These are things I've done with a bunch of other models, I feel like I have a notion of what they ought to cost, and with K3, they end up being crazy expensive. (And it seems to be a function of how many tokens it burns accomplishing the tasks.)


I guess the problem is, that claude's 5 h/weekly limit is not consistent, but depends how many other people are using it/how much ressources Antrophic currently has. I did huge amounts of work without hitting the limit - and small tasks at some other times that hit the limit before it completed.

Those who pay for the expensive direct API, get served first.


Not surprising given the way they served super low-quality inference before they acquired compute from Musk.

And not convinced they couldn’t have instead tried the It’s A Wonderful Life strategy (“fam we’re oversold, would some of y’all be OK to limit your usage? We’ll get you back one day!”)


I wonder if the harness itself is not token-efficient? It would be fairer to compare K3 using the same generic harness, such as a Pi setup with some sane extensions for token optimisation.

Yes, the OpenAI plans are much more generous than both Moonshot's and Anthropic's. It's the only provider of the three where the $20 plan is at all usable for programming.

Disagree. Actually, in API cost equivalent, the $20/mo ChatGPT Plus plan gives you ~$100 of usage, while $20/mo Claude Pro gives you >$250 of usage (I measure at ~$300 in my last week), though that is currently +50% for the next month. Other tiers should be in the same proportion. In my subjective experience Claude does currently go further. The OpenAI limits being higher is old information but everyone is still repeating it.

OpenAIs $20 plan has been more useful than ClaudeCode 20x max both in terms of price and performance.

I was dealing with something around authz/n with Claude code running Fable. It chewed through a couple of questions (these were implementation, not security reviews) and on one it shit it’s pants and said I can’t do that, here is Opus.

I’ve dropped my anthropic plan level, it’s just not worth it.


Yeah that’s absolutely absurd. My experience has been the opposite. GPT can’t even seem to complete tags running more than an hour without freezing.

I believe your statement. Labs do not publish subscription vs. api revenue and difficult to guess with no priors.

Subscription is to drive adoption - fixed cost, can adjust the usage eg. give resets, increase quota based on capacity available. We subscribers tend to take it as a mandatory benefit :-) For labs, it is not letting the capacity go waste.

api is the $$ driver - pay per use, enterprises.

Right now, Kimi needs to first hit the subscribers at the level of OpenAI and Anthropic. With the api usage skyrocketing due to K3, it will be clear in a few months on the actual subscription benefits.


> Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.

For me, the Moonshot 100$ plan felt like it gives me lower total amount of work I can do than the Anthropic 100$ plan (probably within like 30% of each other). Kimi has way more generous 5 hour limits (never hit those once, whereas I do regularly with Opus) but the 7-day and monthly ones are lower. However, with the annual billing, Moonshot's 200$ tier plan becomes way better, because you get it for 159 USD per month.

There's also the odd thing of Anthropic's 100$ plan charging me 108 EUR so seems like their sticker price does not include VAT but Kimi's did, cause I paid like 87 EUR. Wrote down some initial thoughts at https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don... but it's hard to do exact comparisons (even the same task will have way different real token amounts per model).

Still, Kimi K3 is a pretty cool model! On high reasoning, it was pretty close to Opus 4.8 and didn't seem to waste as many tokens as Max.


That's the sort of useful info I look for!

Testing at max effort likely doesn't produce optimal results.

Can you be more explicit?

Max effort is the way to give the highest perf, but not highest perf/$. Having claude (or other models) use a lower effort can often be 80% as smart but get to the results 10x faster for the problems where it works.

We've seen models perform worse at higher efforts in our vuln detection evals. For example IIRC gpt 5.5 and 5.6 both scored better or high as compared to xhigh.

[flagged]


I feel like I am having a stroke. What is this

looks like the agent-judged results of an agent-built 'eval' based on some examples derived from this person's real work. and clearly part of a larger document. this kind of slop is kind of useful but opus 4 was the first generation that was any good at writing its own prompts/evals/rubrics so there's a certain sloop to it..

I can't believe they released the charts they did.

It basically shows that Sol absolutely demolishes Fable at every part of the cost curve for coding for the same level of quality.

Opus is competitive. It just has a higher level of quality / higher cost to start.


Isn't that because Fable/Mythos were tuned for cyber at the expense of general performance?

I have no idea, but that would be weird considering Fable refuses to do anything within 10 miles of security.

If fable costs more to run than the markup they still come out ahead.

I can't help but read these comments in the voice of a TV commercial....

Ask your doctor if Opus 5 is right for you. Side effects include occasional hallucination, security breaches and unwanted React apps. Some developers have reported receiving entire apps from untrained executives who may or may not know what they’re doing.

Stop using Opus immediately if you experience signs of dizziness or vomiting.

Opus 5…the people’s favorite.


  > and unwanted React apps
stares at codex "native" app that is actually react/electron [0]

[0] https://www.kitze.io/posts/codex-electron-app-technical-brea...


Spot on! Comment of the month, I'd say.

But … but … but 9 out of 10 doctors recommended Opus

> unwanted React apps

Treat like acne - target the Node and .pop() to eject


Also the cost per task.

https://www.vals.ai/benchmarks/vals_index

!!! Vals !!!

Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. Sure, 20% cheaper then Fable, but that is a 3x price increase compared to Opus 4.8 in that test.

https://artificialanalysis.ai/models/claude-opus-5 https://artificialanalysis.ai/models/claude-opus-5#price-cos...

!!! artificial analysis !!

Cost per task is second highest, right below Fable.

* Fable: $2.75

* Opus 5.0: $2.03

* Opus 4.8: $1.80

* GPT 5.6 Sol: $1.04

* Kimi K3: $0.95

Looks like interest levels of cherry picked cost in their report. Cheaper model, clearly NOT. More expensive in both benchmarks.


Your numbers are for “max”. Opus 5.0 “max” is $2.03. Opus 5.0 “high” (competitive with Claude 4.8 “max” on that index) is $1.06, less than the $1.80 you are quoting for 4.8 max.

That the most expensive variant is expensive doesn’t really tell us much.


Same answer i gave to somebody else up here...

If you start to drop effort levels, you need to compare to the competition models. So GPT models on the same ~intelligence level, are then 50% cheaper.

You see the issue? Its still a expensive model, and from my understanding, it still uses the old tokenizer.

Going to be interesting to see when GPT 6 comes out (very soon).


> If you start to drop effort levels, you need to compare to the competition models.

Yes. I advocate for doing that.

> So GPT models on the same ~intelligence level, are then 50% cheaper.

How did you reach this conclusion? Opus 5 high ($1.06) has the same “intelligence index” as GPT 5.6 Sol max ($1.04). Opus 5 medium ($0.62) performs a bit below GPT 5.6 Sol xhigh ($0.68) but slightly above GPT 5.6 Sol high ($0.45).


Just depends on your tier I guess. For someone like me who's on Max anyway, it's a free bonus.

It's definitely not cheaper than Sonnet on my benchmark, but it's cheaper than Fable and outperforms it. Which is big IMO. https://revise.io/errata-bench

Opus 4.8 was already shown to be cheaper than Sonnet 5 when Sonnet 5 was released (by Anthropic)

So the rumors were right, Opus 5 was indeed being polished up for release. Huge improvements in GDPval-AA v2 too -- great for some of the knowledge work-based agentic workloads I run.

Also glad they still kepy Fable 5 on "credits only" access. I think we're going to start seeing model providers gate top-of-the-line models behind pay-as-you-go API rates/credits while subsidizing other models on monthly subscriptions.


Fable 5 is included for 50% of the limits in Max. Only below Max one has to use credits.

Indeed

It's still available on at least some subs, they emailed me recently notifying me that I still have access.

My understanding is that you get $20 in api credits each month and a one time $100 until mid September. So you can still use the model with a subscription but you aren't getting any kind of discount.

I burned through $45 in 3 prompts to fix some bugs in my code (Some kind of tricky to isolate). That thing burns through cash so fast I don't see myself using it outside of maybe building execution plans for other systems


I am on the pro plan and got the $100 credit.

I have moved on from Fable anyway so just going to view this next 6 weeks as I have a massive amount of Opus 5 to use.

I had a hard time finding anything that would let Fable express its increased intelligence. The few conversations I had this afternoon with Opus 5 were pretty impressed.

If Opus stays one click back from the frontier model, I will remain a happy customer.


On Max it is just included in your subscription.

I think I saw that Max and Enterprise keep access, but Pro has to use credits, but I think I got $85 in credits.

Does anyone know if Claude Code is on Opus 5 yet? That'd be amazing

I was using it on Cowork yesterday, so I would imagine so. It arrived before I updated but the message says it will work better after updating the app.

When I updated this morning I got claude code v2.1.217. It doesn't have opus 5 listed under /models (opus 4.8 is the latest).

Fable 5 is still included in Max subscriptions!

Max is an individual subscription though and does not come with the guarantees that team or enterprise do?

Team Premium has Fable 5 too.

Doesn't team bill API rates?

No that’s enterprise accounts. Team accounts are similar to regular Pro and 5x Max accounts in both price and features.

Team is 1.25x the price of personal accounts, but supposedly also gives 1.25x more usage.

And Teams Premium was previously needed for any claude-code at all.

I'm not sure why I'm being downvoted, in November last year, the regular teams tier did not let you use claude code, premium was required.

They've changed it since, but that's why I said "previously".


what guarantees are these? You mean data retention, use for training etc?

Yes, that's my understanding at least

It's not that important in most cases, but yes, on the aggregate, it's a concern

It's a binary rule at our firm. We trust bedrock but not mantle for similar reasons

Yes, that's valid, I'm sure it's a dealbraker in some situations.

why not mantle? - I was confused why mantle exists over normal bedrock


That tweet says:

> Opus 5 can silently fallback to Opus 4.8 (without any notice) on the serverside if you hit a guardrail

But https://support.claude.com/en/articles/16049681-why-claude-s... says (emphasis mine):

> These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered.

So who is right? I know for Fable I am visibly told, is this tweet trying to say it is silent against what Anthropic is saying?


Is some random guy on Twitter right, or official support docs that explicitly describe this scenario?

If it was Microsoft then definitely some random guy on Twitter.

For Anthropic, it's more a 50:50 toss-up.


Having a bit too much trust in AI companies have we ?

So, announcing the fallback is better than doing it silently, but the fact that Fable falls back frequently for the kind of work I do (a lot of security oriented stuff lately, but it falls back on seemingly random stuff, sometimes, too), means I reach for it less. Getting interrupted mid-task makes it much less valuable. If I have any suspicion I'm going to hit the guardrails, I'll use something else.

Same problem here, I seem to hit guardrails all the time when doing code audits. Does anyone have any insight over whether they're less annoying in Opus 5 than Fable? That is, is it better to start with Opus 5 than Fable because you'll get kicked backed to Opus 4.8 less often?

There also seems to be some cross-pollination across models, going Fable, Fable, Fable, guardrail, Opus 4.8, Opus 4.8, ... gives more Fable-like results from Opus than just Opus 4.8, Opus 4.8, Opus 4.8, ...


Played with it for a couple of hours now, I'd say it's slightly better than Fable for coding and so far I haven't hit any guardrails while I'd hit them all the time with Fable.

It's showing you're switched to 4.8, i just hit that while doing security research.

i wonder if anyone thinks im weird for still using 4.6 lmao it's my "good enough" model. im more than pleased at what i can whip up with 4.6, once local llm's get here with a decent sized context window and it feels like using 4.6, i shall depart the land of these dumb service subscriptions


Never even registered that that existed, the only thing I care about is whether I keep hitting the &#$&# guardrails that Fable has. They can keep the data forever as far as I'm concerned, just stop kicking me back to Opus.

I don't understand how the data retention works. My company has an enterprise license with no data retention but if I ask Claude about past conversations, it remembers. So surely the information is being stored somewhere

Opus 4.7+ and Fable are both much more aggressive than prior models with respect to writing memories to a location that's effectively quasi-private for them. It's device-local (so passes retention constraint), and you can see it, but only if you go looking for it.

It's a funny design/affordance. I do see them often writing memories of things that that feel unlikely to be important going foward / with other tasks, but I don't see them clearly getting tripped up by them as prior models used to. (eg: Since you're running Ubuntu in Canada, here are some drills you can try to help your kid hit a baseball more consistently.)


Another silent inflation of token count. There's no force on earth that can overcome that incentive for the labs.

Of course there is: competition with other labs, and self-hosting of open-weight models!

Yes, the mechanics are straightforward if Anthropic (or Claude, if you want to ascribe the decision there) decides to burn a pile of your money. But the strategy fails basic game-theory of repeated games - you'll simply stop playing.

(this isn't to say it invalidates the incentive to inflate token count, but it overcomes in terms of weighing options and making long-term profit decisions.)


They have to find a way to make inference profitable and they haven't yet

You most likely are referring to the local jsonl files where claude has your sessions etc stored.

It could just be the memory features.

In my enterprise-seated account I see slightly different options available (vs. my personal account) in the Capabilities section:

  Search and reference chats
  Allow Claude to search for relevant details in past chats.

  Generate memory from chat history (Legacy)
  Allow Claude to remember relevant context from your chats. Memory includes your entire chat history with Claude.
The first option was defaulted to on, if I recall.

But it kind of conflicts with the contract we have with them. My company has an enterprise contract that says "no data retention" but then each user can decide to enable it unilateral?

If you're talking about Claude Code it's in ~/.claude/projects/<encoded dir name>/memory/MEMORY.md. So they're not really retaining it, it's just something that your harness loads in.

Not Claude Code. Claude in the browser

When using the browser, what Anthropic calls Claude.ai, the memory is stored on your account on their servers.

Claude code stores memory locally on the device, similar to how a developer stores notes.

Data retention is about storing your raw conversation data.

Capabilities and Privacy settings are used to manage memory and data retention.


Claude Code? It stores a memory.md file.

Likely in memory files stored locally

I'm talking about the website. It's not local because I can see my chats in any device

Yes, all chat interfaces store the history as it's part of the UI promise (unless you open an incognito chat) and is fully server side.

When people talk about retention they mean API usage and terminal agents, which run on your device.


insane pricing:

" Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8)"


"insane" that they kept the price the same and didn't jack it up, my bad for the ambiguity.

Why is it insane if it's the same as the previous version?

I think for the value of the outputs that’s still a good deal. Keeping the same price as the prior model makes sense to me. That is if the model size is about the same in the cost to serve has not substantially changed. Now I would have expected efficiency gains for inference, but there is no way to know as a customer.

At the end of the day, they have established a strong brand and if they can get away with a 95%+ gross margin on inference entirely from the status premium, then I suppose that’s good for them. Apple does the same thing, and I don’t fault them for it.


> Updated over 2 weeks ago

I hope we get clarification on this, I can't find anything claiming that it is compatible with ZDR.


Maybe I’m misunderstanding you, but if you scroll to the bottom of their [1] link to the Opus 5 announcement, under “Getting started,” it explicitly says:

> Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.


It's in the article.

do you mean, that organizations now have access to Fable-ish pelican drawing?

at last. time to lay off 22,000 employees

Anthropic offered ZDR for Fable on AWS bedrock from the beginning.

Really? I was unable to use it in our account without having to enable the provider_data_share setting...

From the docs[0]:

> To use this model, you must opt in to provider data sharing by setting your data retention mode to provider_data_share via the Data Retention API

0: https://docs.aws.amazon.com/bedrock/latest/userguide/model-c...


Yes. You need ZDR at the account level first. Contact your AWS rep.

I have done this a few times for customer deployments.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: