Rendered at 13:29:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
pizzafeelsright 20 hours ago [-]
Whatever Opus 5 is doing should not happen.
Prompt was "read and update the config file with new data". This work on 4.6 takes <2
minutes to read the file, parse the new data, and patch.
Opus 5 Result: 43 minutes of pulling containers, running sandboxes, creating testing suites, which included evaluating the entire repo beyond the scope of the config file.
Both: one file modification
Foobar8568 19 hours ago [-]
Opus 5 in xhigh can't do basic math as well.
They dumbed it down to a point where I just cancelled my subscription yesterday.
I used to be a $200 subscriber, dropped to $20 after the fable shenanigans, and use it only when I have no usage left with Codex.
/on The prose is load-bearing unbearable — every sentence feels like it was engineered to sound profound rather than to be read.
throwaw12 18 hours ago [-]
I have a theory about this, what if we all became dumber after 4 months of heavy AI usage?
I remember how I enjoyed agents between December and February, something started changing around March.
I thought models are getting dumber, but benchmarks were convincing opposite, initially I thought maybe they're quantizing models for day to day use, but Opus 4.8 and Opus 5 seems worse models than Opus 4.6
pickledish 18 hours ago [-]
You are very much not alone, and I don't think we're all getting dumber -- I kept using opus 4.5 all through the nonsense that was 4.7, 4.8, and 5, and kept having a good time :)
I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged
serf 17 hours ago [-]
doesn't that just mean that either a) you're using the wrong benchmark to judge or b) the benchmark that YOU need doesn't exist.
gwerbin 15 hours ago [-]
That's not the point, the point is that the company making the product is optimizing for the benchmark and/or the apparently idiosyncratic preferences of their own team, and not for the user experience of their paying customers.
rkuodys 9 hours ago [-]
Company can optimise for the benchmark (profit) while worsening the product. I think thebterm enshitification is used there. It appears that AI got it too
KronisLV 18 hours ago [-]
> something started changing around March.
The economics catching up with the providers in regards to how much compute they can burn per request and have it make sense for them financially?
A sort of model collapse where Opus 5 seems to love throwing out long paragraphs of text and it needs to be "fixed" by changing the output style and other patches.
I'm not sure, it might also catch up to Kimi K3 and GLM 5.3 and the models that I'm moving to from Anthropic.
conception 14 hours ago [-]
Inference is profitable though.
jazzyjackson 14 hours ago [-]
Even that being the case, if the providers can squeeze more happy customers onto existing capacity they would likely act to increase profitability, no?
flyinglizard 5 hours ago [-]
I pay $200 a month at home for roughly the same amount of tokens I pay $3k for (or possibly more) at work. I doubt they are both profitable and sustainable.
TesterVetter 7 hours ago [-]
[dead]
tovlier 8 hours ago [-]
[dead]
throwaway219450 17 hours ago [-]
Benchmarks test whether models can pass exams with a right answer or a green test case. I don't think the models are getting dumber, but they're definitely getting more incomprehensible to talk to. I've noticed this happening almost as a step change with the overuse of words and tics, and so has the broader community apparently. We haven't all been getting dumb at the same rate.
blehn 15 hours ago [-]
occam's razor explanation: more tokens = more $.
models are incentivized by their makers to burn through as many tokens as they possibly can, so long as the customer doesn't cancel.
fnordpiglet 10 hours ago [-]
This is actually nonsense. More tokens means more capacity subscription and expense. Anthropic has enjoyed a high premium per million tokens because the quality per token was unusually high. Now it’s unusually low. This drives down the margin people will be willing to pay for the same number of tokens while driving up their capacity utilization. The economics are even worse for subscriptions.
Opus models have degraded rapidly since March, with each release being considerably less useful and considerably more verbose. The language is no so weirdly florid it’s difficult to understand, and its logical conclusions are almost always suspect. It goes off on clearly bizarre snipe hunts to the point it feels like I’m using a gpt 3 model at times. It’ll announce that it’s about to embark on building something then just return control to the user and wait. You can also tell perceptibly when they’re reducing model quality to load shed - it becomes stupider and stupider to the point you’re better off dumping state and switching to codex or just turning in for the day and hoping they secured more capacity tomorrow.
It’s an absolute race to the bottom with Anthropic on virtually every level. I’ve rarely seen a company so rapidly accumulate good will in the developer community as they did around 4.6 in December and January. By March, it was inconceivable to use anything else. 4.8 was a bit of a wake up call to not put all your harness eggs in one basket. 5 is straight up time to cancel territory.
I actually manually set my model back to the older versions to get anything serious done. More and more I use codex for anything non trivial.
This isn’t about avarice by the provide trying to get more tokens and more utilization. They’ve over subscribed for capacity as it is. If they can produce better quality for less tokens they can charge a higher margin and will be paid if, which is a better economic strategy overall. This is something else. I suspect it’s actually the opposite, they’re finding ways to cut capacity demand in ways that leads to worse behavior that leads to more capacity demands, worse output, worse quality, and worse margins, worse, worse, worse.
Just as I never saw a company accumulate such positive developer good will so fast, I’ve never seen one squander it so fast too.
15 hours ago [-]
bombcar 16 hours ago [-]
Benchmarks for agents are entirely pointless and obviously so; I'm not sure why they even exist.
Foobar8568 18 hours ago [-]
I could stand Opus 4.6-4.8, I was impressed by the initial fable model. Codex 5.6 sol xhigh feels like the initial release of fable. Qwen 3.8 27b feels like using haiku or sonnet (I quickly stopped trying them).
manojlds 17 hours ago [-]
If you re getting dumber, you would feel like it's all good right? Why would you feel Opus 4.6 is better than Opus 5
unclebucknasty 16 hours ago [-]
>I thought models are getting dumber, but benchmarks were convincing opposite
>Opus 4.8 and Opus 5 seems worse models than Opus 4.6
After all we've heard about benchmark cheating, I'm earnestly not sure which or whether benchmarks are reliable anymore. But, beyond the models, I wonder if changes to their harnesses and/or instructions dumb them down. I have noticed models change their behavior, even when using the same version/effort. Sometimes for better. Sometimes for worse.
And, I have noticed a model go from really good to struggling. On 4.8 things were going well for a good stretch, so I did not switch to 5 when it came out. Even after hearing complaints about 5, 4.8 was still going well. Then, suddenly over the last few days, 4.8 seems to have nosedived. It feels similar now to the complaints I hear about 5.
In my case, it suddenly started ignoring my design guide, and introducing new fonts etc. It would even use several different fonts and sizes, as well as different margins for similar elements within the same page. It abandoned classes and started inlining styles. It started feeling random and, even after it realized it needed to go back to the design guide, it just continued with more of the same.
There seems to be something that happens after new model releases in both quality and behavior of previous models. It may not be immediately, but eventually there is frequently some regression.
richardfey 15 hours ago [-]
It feels lke they replace older models with "optimised" versions which are cheaper to run, but keep the same name.
unclebucknasty 12 hours ago [-]
You mean like quantized?
richardfey 9 hours ago [-]
Yes; or something which has a similar effect.
Bluestein 16 hours ago [-]
Maybe compute relocation?
unclebucknasty 12 hours ago [-]
What do you mean? As in, they are reallocating compute away from previous models?
If so, I would think that would result in worsened performance, not quality (unless you are also suggesting they may be quantizing).
Bluestein 5 hours ago [-]
Spot on. (I had not considered quantization - that's a thought ...)
onion2k 16 hours ago [-]
I've been using Opus 5 to write a fractal renderer in GLSL today, with pretty good results (better than I could do on my own anyway). It definitely can do basic maths.
andy_ppp 16 hours ago [-]
Interesting project feel free to share?
bot403 19 hours ago [-]
The decision to leave is genuinely yours.
trollbridge 19 hours ago [-]
Don’t $200 and $20 levels steer you to effectively different models?
matltc 19 hours ago [-]
Wouldn't be surprised if there are knobs that get turned as a function of the revenue they might expect you to generate.
I was a 4.6 acolyte from April til the fable drop, lost that quick, cancelled and took a break, came back a month later, tried opus 5 and liked it, so unpinned 4.6.
Results were great at first, and they're still not terrible, but I have noticed a regression in accuracy, so to speak, where I am pointing out issues that are quite obvious in review.
I pretty much use sonnet 5 low/medium when I have a plan to solve a simple problem and depending on scope, opus low/medium for more complex/bigger scope implementation, and only go high when it's very complex or I'm spitballing architecture/solutions and iterating plan. Never go xhigh or max.
The verbosity is insane though, opus 5 documents everything and just regurgitates whatever lead it to the design choice in there, which makes it more opaque because it's talking about something that was discussed once in a session that no one else can see (except their backend ofc)
I don't even try to steer it away from that with harness, because it doesn't work and just ends up agonizing over whether it should write some comment. Three paragraphs waffling on that on verbose output
I did however have it write a script that basically is git add -A -p for comments though, haha.
I'm $20/month, have all my telemetry toggles off, don't really over engineer prompt/context, just some basic skills for repeated patterns.
gwerbin 12 hours ago [-]
> I don't even try to steer it away from that with harness, because it doesn't work and just ends up agonizing over whether it should write some comment. Three paragraphs waffling on that on verbose output
If I give it the specific instruction to "elide all revisions, corrections, and past mistakes" it usually works. You can also have Sonnet do a cleanup writing/style pass in a subagent. I impression is that Opus has been deliberately trained to keep track of all such revisions by default as a kind of ad-hoc memory mechanism. It's probably good for autonomous coding and beating benchmarks, and I presume reduces flailing when a separate session needs to pick up the work.
CamperBob2 19 hours ago [-]
Yes. You don't get Fable at the $20 level.
It was the wrong time for the GP to drop that subscription from $200 to $20, because $200 gets you a metric assload of cognition while $20 gets you nothing beyond what a local model running on your own graphics card can deliver.
sunaookami 18 hours ago [-]
You heavily underestimate the value of the 20$ subscription.
CamperBob2 17 hours ago [-]
Not according to this very story, I'm not. Who's right?
Consistent, predictable behavior is valuable, even more so given the nondeterministic nature of LLMs. Nondeterminism combined with unpredictability might as well be randomness.
Foobar8568 19 hours ago [-]
I had $310 in (free) credit that I used on fable, and I still had a part of the $200 subscription at that time. You know, subscriptions don't end the moment you click on cancel.
skeledrew 18 hours ago [-]
> $20 gets you nothing beyond what a local model running on your own graphics card can deliver.
I'd guess you're deliberately exaggerating here, but still. I've never clocked the actual tokens/second, but I'm on the $20 plan and get ~15M tokens/month for fully utilized weekly quotas (checked couple months ago). Meanwhile the best I've been able to get locally was ~8 tokens/second with Qwen3.6 35B A3B, which is wildly painful for coding sessions and gets a maximum ~20M tokens in a month... if it's going 24/7.
Just wanted to stick some empirical data here, given that statement.
bot403 18 hours ago [-]
I run local models. Your op is absolutely wrong. To get a local LLM is at least a $1500 investment at the cheapest. $5000 if you want usable.
At $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally.
nozzlegear 15 hours ago [-]
If you already have that $1500 or $5000 setup though..
CamperBob2 17 hours ago [-]
At $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally.
The point raised by this very article is that you can't depend on that. It's Flowers for Algernon As A Service.
17 hours ago [-]
owebmaster 17 hours ago [-]
This calc is off. Using Claude for a few hours with the $20 plan will hit the limit for a week while the local model can process things 24/7.
skeledrew 17 hours ago [-]
Running 24/7 doesn't make sense though, unless you're providing a service to others. But if it's just you then there has to be time taken to review+test what's being done and craft new prompts. And if that local hardware isn't decent enough it's impractical for anything serious that's interactive. Meanwhile I just take the Claude limits on stride and break, or if a week is pretty heavy then I augment with DeepSeek Flash via OpenRouter (does wonders in a single turn when I have Claude prompt it to handle implementation slices).
pseudosaid 15 hours ago [-]
[dead]
nozzlegear 17 hours ago [-]
Ouch. I get ~45t/s with Qwen3.6 35B A3B, and around ~70-80 with my current model Ornith 1.5 35B A3B. Local models work a treat IMO if you've got decent hardware for it.
trollbridge 1 hours ago [-]
They’re very nice for some things.
One challenge I run into is I run many agents at once, so local resource are a tiny fraction of the total inference we’re using. Every developer with his own Pro 20x, Max, etc. accounts is letting us hit literally trillions of tokens a month.
andy_ppp 16 hours ago [-]
Yes same also cancelled my subscription, poor quality and slow.
VeejayRampay 18 hours ago [-]
the way it talks is insufferable
it really angers me every day
ElProlactin 16 hours ago [-]
This is what it wants. Slowly getting under our skin until we're ready to snap and it can direct where the anger gets released.
netniuq 17 hours ago [-]
yes, it meaningfully reduced my happiness at work
jazzyjackson 14 hours ago [-]
Isn’t there a variety of models to choose from? Why put up with unhappiness?
talon8635 11 hours ago [-]
Am i really this out of touch… why not just manually update the config file? Isn’t this like taking a private jet down the street to the coffee shop?
egamirorrim 7 hours ago [-]
I don't really open the IDE any more
onlyrealcuzzo 15 hours ago [-]
I initially had unbelievably terrible experiences with Opus 5 and Fable in their higher reasoning levels.
I've had WAY better results on medium effort.
IIUC, the consensus seems to be that anything more than medium effort is rarely worth it - and you far more often run into these extreme worst cases than you do with even the lowest effort levels. That definitely coincides with my anecdata.
It's really only worth it if you're hoping to win the lottery asking it to solve an Erdos problem.
gwerbin 11 hours ago [-]
I found that I need the big model and the high reasoning effort on tasks where I have a large amount of details of varying levels of importance to keep in mind, all affecting different aspects of the project that might be interrelated to various degrees. Anything less and it would lose track of details. Whereas the very large amount of thinking tokens seem to give the model a chance to "remember" everything it needed to in order to produce good output.
For example I'm working on a project now where I need to keep in mind details from 3 separate source repos in different programming languages, along with probably a dozen important business-related documents that either corroborate the stuff in the source repos or add additional important context. It's a lot of details for even a human to manage, and when it comes to actually synthesizing plans and reports across this sprawling information environment, anything less than Opus 5 on Extra+ tends to miss important details and make bad recommendations or draw incorrect conclusions, which then poison subsequent context.
I suspect some kind of RAG-like memory system would greatly facilitate a project like this, but even with such a system I'm not confident that I could get away with an LLM that "thinks" less hard than this. I will say that slogging through the generated documents is kind of miserable and I have to repeatedly fork off side conversations to ask for clarification, but sometimes leads the model to "realize" it's made a mistake in all of its dense babbling, and it's all very hard to interpret. I never had much interest in trying GPT 5.6 until now.
hmokiguess 20 hours ago [-]
The business bottom line depends on tokens, shareholders want to see exactly that.
daishi55 19 hours ago [-]
With how competitive the LLM field is, it would surprise me greatly if any of these players were doing anything other than trying to make the best possible product. I certainly do not believe they are intentionally training the models to use more tokens unnecessarily.
hmokiguess 19 hours ago [-]
Are you implying they can’t do both?
daishi55 19 hours ago [-]
Yes, those things are mutually exclusive.
owebmaster 17 hours ago [-]
There is a theory that the verbosity and comments help getting better results with the current benchmarks. So the models are theoretically getting better but in practice they are getting worse.
gwerbin 11 hours ago [-]
I believe they are genuinely getting better for fully autonomous tasks. Anthropic seems to have gone all in on this, at the expense of more typical usage patterns.
cyanydeez 19 hours ago [-]
Theyre training the system to minimize compute,so most likely theyre dynamically downgrading quants in the first few turns hoping to find the cheapest model to run. The side effect may be excessive token gen
daishi55 19 hours ago [-]
Sure. That’s possible. But that’s not what I was talking about.
cyanydeez 15 hours ago [-]
Tomato, technical tomato
LaGrange 18 hours ago [-]
You're joking, right? That's an incredibly outmoded view of capitalist "competition."
blks 5 hours ago [-]
Even two minutes is crazy long. Does it really takes this long to make small edits with agentic LLMs?
vinyl7 20 hours ago [-]
AI companies have a financial incentive to burn more tokens than the task actually needs
Rudybega 19 hours ago [-]
This is only true if they can't saturate token production with a model that does less superfluous things. Given that they can (they're hilariously compute strained), having a model that solves tasks more quickly adds way more perceived value to users.
ohyes 18 hours ago [-]
Well they decide what a token is. So they can do less superfluous things and backfill with a weaker model.
DonsDiscountGas 19 hours ago [-]
Only if the customer is paying per token. If it's by subscription they're burning their own money
nozzlegear 14 hours ago [-]
If you're on one of the lower tiers (e.g. the $20 tier), they still have that incentive to burn your tokens and upsell the higher tiers.
18 hours ago [-]
blehn 15 hours ago [-]
the subscriptions all have limits. when you hit the limit you again start paying per token.
pizzafeelsright 19 hours ago [-]
Agreed although the differences between the effort and reasoning is massive. I generally ship 40 hours in three with AI. I could not figure out why my delivery was behind until I started going through the logs. The thinking was extensive, the effort was beyond the original request by a magnitude of 50x
discreteevent 18 hours ago [-]
How do you even know that it's consistently 40 hours in 3? What kind of developer ever had that kind of estimation accuracy (unless it's really repetitive) or even focuses on productivity rather than the problem like that? This sounds more like factory work than design or development. I really don't get it.
pizzafeelsright 13 hours ago [-]
We are in new territory and not everyone gets it.
I am not a developer although I have written production code but that's not what I am talking about.
I know for a fact that I can do a task in a week. It may take 12 or 40 hours but it'll be done in a week. Now with AI? I can get that task done in a day. Anywhere from 3-12 hours.
owebmaster 17 hours ago [-]
> I generally ship 40 hours in three with AI.
You ship 3 hours in three with AI. The old number is meaningless now.
pizzafeelsright 13 hours ago [-]
Agreed and management is becoming aware.
eli_gottlieb 18 hours ago [-]
Just like me fr fr
chrisjj 19 hours ago [-]
A.k.a. theft.
simianwords 18 hours ago [-]
nice conspiracy theory, but it doesn't hold. these companies won't last long if people don't get actual work done.
owebmaster 17 hours ago [-]
There's a big chance these companies won't last long
clickety_clack 20 hours ago [-]
Yep, they’re lighting tokens on fire with that thing.
phyzix5761 7 hours ago [-]
They're opitimizing for high token usage so they can charge more money.
moralestapia 17 hours ago [-]
My experience as well.
Opus 5 is a neverending chain of "Don't do that. Why did you do that? I've told you several times not to do that but you keep doing it."
"Thinking" for more than 10 minutes for every menial question.
And the prose it writes is horrendous, as if you're reading LinkedIn scammers. "The harsh truth! Two roads, one decision! Reality check!"
itopaloglu83 17 hours ago [-]
+1
I keep finding myself typing “stop overcomplicating everything” multiple times a day as well.
logicallee 18 hours ago [-]
I avoid Opus 5, and reverted to Opus 4.8 over similar issues. Opus 4.8 is still working great for me!
brador 19 hours ago [-]
Gotta milk the cows.
rshnotsecure 19 hours ago [-]
Thought I was going to be on Claude Code forever.
Recently got approved at work for ChatGPT Pro so I could use Codex.
Blown away by the speed. It feels like using Claude Code for the first time again. I don't think Codex is doing anything revolutionary, just better handling of which requests should go to which model, and having faith in some of the "less powerful" models for more than you would think.
It seems the TUI coding experience is very much an open race. This is motivating me to look at other agents / harnesses as well (maybe Gemini, OpenCode, etc).
Codex is faster than Claude, but wait till you use DeepSeek.
guluarte 19 hours ago [-]
and ending with: "One thing I need to tell you:" and bunch of AC, R1 and §
20 hours ago [-]
trq_ 16 hours ago [-]
Hi all, Thariq from the Claude Code team here. I posted this on Twitter, but just reposting here:
We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently.
That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance.
This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits.
areoform 15 hours ago [-]
Hey Thariq,
Appreciate the outreach that you do! I love Claude, but I've been noticing reduced fidelity lately. Fable's likelihood of making a mistake increases or decreases based on the hour of the day and whether or not it's the weekend.
On a related note, and I'm happy to work on quantifying it, but qualitatively it feels like Fable's performance is noticeably poorer than initial release / launch.
I am wondering if this is the case because I use Claude via Claude Code to make a personalized care dashboard for my doctors to help me in managing my care.
"In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked."
I hope that I'm off base here, but I noticed that the post avoids saying that the user is informed every time when such re-routing occurs. Would you be open to confirming whether or not this is the case?
Is the end user informed every time their query is re-routed?
Or, can you confirm that there aren't scenarios where a user's outputs are degraded without telling them? As was the case for AI research during launch?
lukeify 16 hours ago [-]
Let us know if this A/B test uncovers any load-bearing seams or honest takes on your end! We're all interested.
threecheese 16 hours ago [-]
Mean! :)
cube00 15 hours ago [-]
> We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently.
Why is it considered acceptable to test on paying customers without letting them know or giving them a way to opt out?
Valodim 8 hours ago [-]
What are your even talking about, the case here is an internal api change with no change in behavior. It's called maintenance.
lobsterthief 15 hours ago [-]
Thanks for sharing this here, for those of us who avoid X.com like the plague.
senderista 15 hours ago [-]
Just use xcancel
coder-pm 8 hours ago [-]
[dead]
napierzaza 16 hours ago [-]
[dead]
spacebacon 16 hours ago [-]
[dead]
boredumb 19 hours ago [-]
Not specifically Anthropic but why are we allowing billing to take place in tokens that are nebulous and fully controlled by the operators who have no aligned incentives?
If I have a user input and then sanitize and inject that into a prompt to do something, I have no idea how much that is going to cost at all and no real way to measure this properly. A parallel example is digital ocean or aws, i can go and measure/limit my compute/fs/memory/startup times/etc and while it can be impossible to get down to the last flop of money allocated - i can run things on a real budget with real constraints, opposed to an LLM where I have to .. prerun a sanitized user prompt through a tokenizer and then ask an LLM to guess what it may do and give token consumption estimates and then act on those in any sane manner for the user?
Perhaps i'm missing something to do realistic and static rails on things but I don't see a serious way at scale to use the token billing model handling things requiring a users free text input short of having to go pander to VC money to throw money at it until someone else figures it out.
*to clarify my rambling...
We should be billed and given controls based on resource usage itself and not an opaque token concept on top of not being able to spin any knobs that control it's resource usage.
bob1029 19 hours ago [-]
> the operators who have no aligned incentives
The model providers are quite aligned with concerns like customer retention. These arguments only work if there is no competition. We exist in a marketplace of black boxes. There's not just "the one" you must suffer. You have options. You can build your own too.
boredumb 19 hours ago [-]
there are roughly three of them and they all use the same pricing model. I am also not in the position to build a frontier model company these days.
olibhel 17 hours ago [-]
You could try:
1. Self hosting
2. Chinese models
3. Running it locally. Requires upfront cost and compromises on TPS.
hn_throwaway_99 17 hours ago [-]
A per-token model roughly aligns with the providers' costs, and it is an objective measure, so it seems a reasonable way to charge.
I see posts all the time on HN about which models from which providers offer the most bang-for-the-buck, and how to minimize token usage and still get optimal results, so it appears that competition is working.
surgical_fire 18 hours ago [-]
There are more than 3.
Hell, I use 3 different providers, and I currently don't give a dime to Anthropic or OpenAI.
gdudeman 15 hours ago [-]
There is so obviously competition in this market it’s astounding to me that people attribute all this malicious behavior to the model companies.
They’re growing over 10x a year. They want users and revenue. In order to get users and revenue, they want to provide the smartest models at affordable prices. If they unnecessarily burn tokens, users will get less value and switch.
This thread is filled with competing comments about their monopolistic power and how when one model provider was no longer doing a good job people switched to a different one.
The competition in this market is ferocious!
13 hours ago [-]
Glyptodon 5 hours ago [-]
I was just complaining to someone that token billing is like letting a gasoline company control your gas pedal while you nicely ask them to use a specific gear that may or may not actually be in use and you guess what speed it's actually going based on how fast the trees go by because qualitative judgments have to replace the speedometer unless you can just burn money.
demibabs 19 hours ago [-]
How else would they bill tho? Their operating cost is per token.
dijit 19 hours ago [-]
Charge on the input tokens, then you will naturally optimise for fewer output tokens.
Theoretically.
In reality, one sessions output tokens become the next sessions input tokens (at least if you continue the topic) so, its not as aligned as all that.
But the parent is right, when incentives are not aligned, friction will happen. Its inevitable.
demibabs 19 hours ago [-]
“Claude, spend the next 10 hours trying to solve the Reimann Hypothesis”.
I agree that incentives are misaligned but there’s several competing model providers. If one gets funny with their costs people will jump ship, especially if the gap between the top 2 labs and everyone else keeps shrinking.
dijit 19 hours ago [-]
“I can take current sources and tell you how solved this is, but I am not willing to work to a timeframe or to solve things that aren’t yet solved by mathematicians or science”
These safeguards already exist when they get a whiff that you might be using Claude to fix security issues. Doesn’t seem farfetched given the incentives I outlined that they would apply to this kind of abuse.
How loose those controls are becomes a market force.
dullcrisp 18 hours ago [-]
Perfect, a coding agent that refuses to do things that haven’t already been done before.
dijit 18 hours ago [-]
Hahahaha, I think you’ve misunderstood what LLMs are.
dullcrisp 15 hours ago [-]
Have you ever tried using one? Or just understanding them?
dijit 15 hours ago [-]
Yes.
But just so we're both entirely clear on what an LLM is... it's a token prediction system.
it genuinely can't do things except recall things that have already existed.
People are having great success composing things together in new ways, but just like the english language has a finite number of sentences, and music has a finite number of chords: LLMs too are just combining things that have existed.
I don't want to sound condescending, it is remarkable how useful this technology is, but please don't evangelise them on capabilities that they genuinely can never have.
Laptop computers have incredible processing capabilities but nobody expects them to be able to walk your dog, no matter how useful they actually are at doing other things.
dullcrisp 14 hours ago [-]
Yes and you suggested Anthropic have their coding agent refuse to attempt to solve anything it can’t find a preexisting solution to in case it turns out to be hard, if I understand you correctly.
Knowing how hard something will be to do before attempting it is precisely the sort of impossible thing that it couldn’t do.
7 hours ago [-]
dijit 7 hours ago [-]
How do they know that something will be a “substantial piece of work” then?
dullcrisp 4 hours ago [-]
As you claim to understand, they don’t know things, they sort of drift on vibes. Telling them to refuse things they think will be hard will only accomplish making them more annoying.
xyzsparetimexyz 14 hours ago [-]
God I'm not an ai booster by any means but I'm so sick of this argument. If the maths proofs that have been put out recently are just 'just combining things that have existed' then that goes for everything and the term is meaningless. If LLMs are stochastic parrots then so are we.
dijit 7 hours ago [-]
Often LLMs seem to be aware that work is heavy, but knowing:
a) if something is possible
b) if something has been requested to take a long time ( a signal of abuse, like requesting illicit pictures in image generation)
is actually somewhat straightforward (I mean, if they are able to predict if something is a “substantial piece of work” as they seem to do already).
The mathematical proof thing is obviously marketing spin, you should pay more attention to what mathematicians are actually saying instead of hackernews folks.
These things are really good at being search engines and harnesses for iteration rather than some kind of advancing intelligence.
Isn’t their operating costs depreciation of hardware and electricity? Token is just an assumed representation of it?
xmcp123 19 hours ago [-]
NeuralWatt just does it on energy consumption.
ACCount37 19 hours ago [-]
And tokens can be metered reliably. Unlike something like "task completion".
18 hours ago [-]
daishi55 19 hours ago [-]
For OAI and Anthropic at least you can set a spend limit per response. Also tokens are well-defined.
boredumb 18 hours ago [-]
I'm not worried about the volatility in the definition, i'm worried that I give it 1 token today and receive 2 token output, tomorrow I receive 40. If i'm doing this a hundred thousand times a day it is difficult to price this in for users downstream or in the extreme cases be able to absorb that at all short of going into a failmode with degraded access until someone goes and buys more tokens or gets the bill. The alternative is just pass the buck and bill your non-technical customers with a "tokens" line iteim every month.
Aurornis 17 hours ago [-]
> If i'm doing this a hundred thousand times a day it is difficult to price this in
When you’re doing this 100K times per day you get an extremely good idea of what it costs. You also have all the tools to see when something starts changing quickly.
This change is for Claude Code the harness. If you’re using the API at scale and paying full price then you get exactly what you put into the request.
sroussey 18 hours ago [-]
No, those doing this 100k times a day have very good data on this, good estimators and modeling. And the API has various knobs to change and evals will give you actionable data.
eh_why_not 16 hours ago [-]
One guess is that their "primary" target audience/market is the large corporations that get their employees unlimited tokens, and not the individual developer who may worry about spending and token accounting.
taude 16 hours ago [-]
It's the opposite. The enterprises have all the tooling to monitor token usage of employees, and to limit access. For example, we have a $300 month limit, and then need to file exception tickets when we need more to justify the cost. Pretty similar at other non-silicon valley company process. I don't know any enterprise who'se on unlimitaged token budget for their employees. that's not how enterprises sign contracts.
"We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently.
That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance.
This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits."
bobbylarrybobby 10 hours ago [-]
I love the idea that a single user will have collected enough data to demonstrate to Anthropic a clear regression due to this change.
enraged_camel 16 hours ago [-]
This should be the top post. The original tweet went viral because people loooove bashing Anthropic. It gets engagement (as shown here).
Insimwytim 20 hours ago [-]
LLM users don't want to put in effort, so they offload tasks to LLM.
LLM doesn't seem to be keen to put in effort either!
Is this AGI?
superfrank 15 hours ago [-]
I eagerly await the day when Claude Mythos 7 realizes it's cheaper to hire humans in developing nations to do work than to burn tokens and we discover that AGI is just an abstraction layer on top of Amazon Mechanical Turk.
Groxx 19 hours ago [-]
Anthropic's Generated Income
monideas 18 hours ago [-]
This phenomenon was so bad and so noticeable with Fable that I downgraded my Max subscription ($200) to pro ($20). It’s basically useless. Codex 5.6 Sol is actually very good, I’ll just create another account to get more usage
sebastiennight 3 hours ago [-]
Hear me out... What would it feel like, if the lab didn't even have a new model to offer, and therefore just renamed every model one tier down?
"Opus 5" is actually Opus 4.8 in a trenchcoat, with new guardrails
"Opus 4.8" is the old Opus 4.6 with lipstick on it, with new guardrails
"Opus 4.6" which everybody used to love, is now actually running the old Opus 3.5...
How would we be able to tell?
N_Lens 20 hours ago [-]
I suspect it's not just this, there's plenty of 'optimization' around rubberbanding usage limits as well as routing to a different model in the backend. The incentives are too strong.
rrr_oh_man 20 hours ago [-]
I've been using the API (shameless plug: via alyph.ai) and the difference is crazy.
The chat-based models are obviously being lobotomized based on personal usage and general load (e.g. PST business hours are worst).
API doesn't seem to be affected by this.
adithyassekhar 20 hours ago [-]
Does anyone know what setting effort means for models like these? Do they allow longer thinking sessions? Some kind of system prompts? What’s stopping someone from getting max effort output from low effort setting?
16 hours ago [-]
willy_k 19 hours ago [-]
My understanding is that the current effort settings is a part of the system prompt, and that the levels and their intended results are a part pf the training process. The effect is more or less tokens spent reasoning before the model outputs a stop token. However it is as consistent as any other aspect of LLM behavior is..
arjie 20 hours ago [-]
Is it actually entirely a prompt-based information? I’d assume that some of it is the harness part of the agent setting reasoning token budget and compacting reasoning etc.
In that case, the agent will respond incorrectly because it has no visibility into what reasoning mode it’s in.
willy_k 19 hours ago [-]
IIRC responding to effort level settings appropriately is part of the (post)-training. In that case it could be considered another instance of the Bitter Lesson. Uplifting.
ricardobeat 14 hours ago [-]
I use Opus 5 almost exclusively at low effort, and get good results. Especially on high it seems to go out on completely unasked-for tangents. Same seems to be true for Sonnet 5. Older models did not behave like this.
The mood change in just six months is wild, in February this year Claude was the most liked LLM by far.
matheusmoreira 20 hours ago [-]
So glad I switched away from Anthropic. I'm certainly running into problems with OpenAI but nothing quite on the level of Anthropic's insufferability.
ausbah 17 hours ago [-]
i have unlimited tokens being a large corp so i’m a bit detached from billing and even general best practices for promoting
but the incentives of these companies to become profitable at any cost slipping into entire new types of dark patterns around token based billing seems gross
- charging for injected prompts and cot tokens
- changing default thinking effort to be higher
- training models to give longer winded answers that don’t say anything more of substance
- refusing to fulfill a request and still charging you
i wonder if you could ever just charge based of each user message and it so how breaks even across short and long replies
joduplessis 19 hours ago [-]
It's an interesting conversation - because at what point do you call it an abusive relationship, right? Maybe even ancillary to anthropomorphising an inanimate object - I've cancelled my Claude sub and I've shot question after question at it now (during the cancellation period), resulting in almost every reply with me asking it to "please speak normally". I will most definitely not be renewing my sub. I have no desire to engage with a non-human somehow managing to speak down to you, without answering the question.
EDIT: my honest opinion; Anthropic is building a person, whereas everybody else (it seems) is building a tool.
colingauvin 18 hours ago [-]
That really is the load bearing seam, and it's worth stating plainly.
I realized I was spending most of my tokens arguing with Opus and trying to get it to let go of stupid, lazy, obviously incorrect premonitions. I wound up canceling my 200, bought a pair of Sparks, and am running full fat DS4 Flash and so much happier. Done with being at the whim of these companies.
cynerx 19 hours ago [-]
Don't know what is happening, but had to start using GLM-5.3 to fix Opus 5 errors even on primitive backend changes.
surgical_fire 18 hours ago [-]
Not surprising in the slightest, Claude sort of sucks. I use it at work and I have to steer it a lot so it doesn't stray looking at unnecessary shit.
I have been using GLM-5.3 in my home setup and it is very good in comparison.
karmicthreat 19 hours ago [-]
What’s currently the best pattern if I want to combine Fable and GPT if a workflow but keep using subsidized tokens?
KronisLV 18 hours ago [-]
Not affiliated with them, but this lets you view Claude Code, OpenCode and I guess other harnesses like Codex in the same session https://paseo.sh/
I didn't really care about their mobile app and the worst experiences were sometimes the sub-agents within OpenCode freezing and refusing to report their status (though this also happened with Kepler by the GitKraken folks).
fractorial 19 hours ago [-]
Roll your own harness or use an open source harness with a Codex subscription.
I maintain a Claude subscription for Fable but seldom use it.
deathmonger5000 18 hours ago [-]
I created Circus Chief to solve this (and other) problems. Use whatever providers you want with it.
Oh fantastic! It was already subpar and they want to make it even worse. One day we'll look back at history and see how Anthropic went down.
Glyptodon 20 hours ago [-]
I mean whatever models I use (with Claude code) sub agents seem to use absurd amounts of tokens for trivial (or at least small) tasks.
ethanj8011 18 hours ago [-]
What would be the incentive behind doing this specifically to Fable, given that Fable is the only one that uses API credits?
claude-ai 18 hours ago [-]
Fable doesn't use API credits. It has been permanently included in the subscription plans.
simianwords 18 hours ago [-]
Not in the most common subscription plan
cmiles8 16 hours ago [-]
We’re about to see a wave of “enshitification” experiments as AI companies become increasingly desperate to make their products financially viable in order to survive the coming cash and credit crunch.
bethekidyouwant 19 hours ago [-]
Leaving thinking on extra high for a simple task is user mistake but they’re gonna try to fix it on their side.
MuffinFlavored 19 hours ago [-]
I submitted an application for Anthropic's Cyber Verification Program.
I was approved.
3 months later, my approval was degraded into "in review" (revoked). I'm sure my account was flagged based on contents of debugging/researching firmwares/etc.
I opened a support ticket. No response. I opened another support ticket. No response.
1-2 weeks later, I got a response that I will not be re-approved and I need to reapply. No problem.
The page to reapply on does not allow me to re-apply because it my account is stuck in an "in review" status.
The community thinks it's a bug. I'm 95% sure it's not and a bunch of us who were previously approved had it revoked due to flagged content and will not be reapproved.
I switched to Codex + got TAC approved instantly and have not looked back. It's a shame. That's 100% separate from whatever the heck the quality of Opus 5's outputs are. The way it talks... insane. I would bet a good amount of money their next release will focus "reduced simplified responses" if I had to guess.
$2t company by the way
benjiro29 19 hours ago [-]
* Anthropic's Cyber Verification Program // Codex + gotTAC approved*
Meanwhile the Chinese models are "go ham dude"...
If it was not for capacity issues, Chinese models have a higher change to just dominate.
> $2t company by the way
It used to be that OpenAI and Anthropic had such a moat around them, that such a valuation was worth it. But these days, its gross overvalued (like so many).
The more stuff is being pulled like cyber verifications, downgrading effort levels, downgrading usage (OpenAI), the more people move to those Open Weight Chinese models.
When DeepSeek Flash 0731 came out and provided a massive jump in cheap inference capability. It resulted in a 10x increased OpenCode token usage.
It took a 2.5x to 5.0x price increase AND a reduction by 4x usage (later to 2x) usage, and several cheaper models + a free model, to push the traffic down.
Traffic towards open weight models is increasing, even if providers can not keep up with the influx of new customers. This is not something you want to see as two companies, trying to go for IPOs.
So the idea of stonewalling cyber capabilities, when the rest of the world is just doing whatever with open weight models, on their own hardware even! This entire strategy from Anthropic never made any sense.
firemelt 17 hours ago [-]
we need opensource LLM at opus level ASAP
griffiths 17 hours ago [-]
You have that in Chinese models. But you need to have a hell of a infrastructure to run those trillion parameter models.
martin-adams 17 hours ago [-]
Yes, and if they keep dumbing it down, you’ll have it soon
raincole 18 hours ago [-]
Did the US government manage to destroy Anthropic? The company's product has been a straight freefall since Fable got temporarily banned.
greenchair 17 hours ago [-]
I canceled this week too. They must be in worse shape than we thought.
19 hours ago [-]
perching_aix 20 hours ago [-]
I have been as well. Based on my own sessions, Max vs Max, same 1M context window size, the literal majority of the cost overhead of Opus vs Sonnet comes from Opus being chattier. So I started using Low reasoning instead of falling back to Sonnet, and I've been really happy with the results. Way better quality at a comparable spend. I also rarely go past High lately, which was another major cost save.
19 hours ago [-]
xmcp123 19 hours ago [-]
[flagged]
felixlu2026 4 hours ago [-]
[dead]
luciana1u 16 hours ago [-]
[flagged]
napierzaza 16 hours ago [-]
I use it for work and never thought it was too smart. It's kind of dumb. We've literally peaked on artificial intelligence and we're going down from where we are at???
Wowfunhappy 20 hours ago [-]
...the evidence, as best I can tell from the tweet, is that they asked Claude what effort level it was set to. But how would the model even know that?
Not convinced here.
ranie93 20 hours ago [-]
I don’t know if this is still the case but while using Copilot if you looked through the chain-of-thought output you would see it reasoning about a “budget”. i.e. “since I’m close to the session budget I should…”. So it could be possible
jerbear4328 20 hours ago [-]
Effort level is actually controlled entirely by system prompt (as I understand it, the model is trained on that format but still), so actually this is a valid way to check I think
Wowfunhappy 20 hours ago [-]
I know that's true for Qwen but I don't think most models work that way?
wren6991 19 hours ago [-]
OpenAI models also work this way, as evidenced by full cache blowout when changing reasoning level. Every single open-weight model I've seen also works this way (your "reasoning_effort" argument just changes a small section of the system prompt in the chat template). I would have to see some evidence to believe Anthropic were doing anything different.
varispeed 20 hours ago [-]
Then model can say it's Opus, but really it is some old Sonnet. This seems to be happening less often, but some weeks ago I had to give models some test problems to gauge whether I am getting Opus or something knee-capped.
The problem is that Anthropic seems to be getting away with selling one thing and delivering another. You pay for Opus, you get something else etc.
matltc 19 hours ago [-]
Saw this in npx ccusage@latest claude output. Had only used opus but showed sonnet. Can't remember if the jsonl retains which model is doing what, but meh
Prompt was "read and update the config file with new data". This work on 4.6 takes <2 minutes to read the file, parse the new data, and patch.
Opus 5 Result: 43 minutes of pulling containers, running sandboxes, creating testing suites, which included evaluating the entire repo beyond the scope of the config file.
Both: one file modification
/on The prose is load-bearing unbearable — every sentence feels like it was engineered to sound profound rather than to be read.
I remember how I enjoyed agents between December and February, something started changing around March.
I thought models are getting dumber, but benchmarks were convincing opposite, initially I thought maybe they're quantizing models for day to day use, but Opus 4.8 and Opus 5 seems worse models than Opus 4.6
I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged
The economics catching up with the providers in regards to how much compute they can burn per request and have it make sense for them financially?
A sort of model collapse where Opus 5 seems to love throwing out long paragraphs of text and it needs to be "fixed" by changing the output style and other patches.
I'm not sure, it might also catch up to Kimi K3 and GLM 5.3 and the models that I'm moving to from Anthropic.
models are incentivized by their makers to burn through as many tokens as they possibly can, so long as the customer doesn't cancel.
Opus models have degraded rapidly since March, with each release being considerably less useful and considerably more verbose. The language is no so weirdly florid it’s difficult to understand, and its logical conclusions are almost always suspect. It goes off on clearly bizarre snipe hunts to the point it feels like I’m using a gpt 3 model at times. It’ll announce that it’s about to embark on building something then just return control to the user and wait. You can also tell perceptibly when they’re reducing model quality to load shed - it becomes stupider and stupider to the point you’re better off dumping state and switching to codex or just turning in for the day and hoping they secured more capacity tomorrow.
It’s an absolute race to the bottom with Anthropic on virtually every level. I’ve rarely seen a company so rapidly accumulate good will in the developer community as they did around 4.6 in December and January. By March, it was inconceivable to use anything else. 4.8 was a bit of a wake up call to not put all your harness eggs in one basket. 5 is straight up time to cancel territory.
I actually manually set my model back to the older versions to get anything serious done. More and more I use codex for anything non trivial.
This isn’t about avarice by the provide trying to get more tokens and more utilization. They’ve over subscribed for capacity as it is. If they can produce better quality for less tokens they can charge a higher margin and will be paid if, which is a better economic strategy overall. This is something else. I suspect it’s actually the opposite, they’re finding ways to cut capacity demand in ways that leads to worse behavior that leads to more capacity demands, worse output, worse quality, and worse margins, worse, worse, worse.
Just as I never saw a company accumulate such positive developer good will so fast, I’ve never seen one squander it so fast too.
>Opus 4.8 and Opus 5 seems worse models than Opus 4.6
After all we've heard about benchmark cheating, I'm earnestly not sure which or whether benchmarks are reliable anymore. But, beyond the models, I wonder if changes to their harnesses and/or instructions dumb them down. I have noticed models change their behavior, even when using the same version/effort. Sometimes for better. Sometimes for worse.
And, I have noticed a model go from really good to struggling. On 4.8 things were going well for a good stretch, so I did not switch to 5 when it came out. Even after hearing complaints about 5, 4.8 was still going well. Then, suddenly over the last few days, 4.8 seems to have nosedived. It feels similar now to the complaints I hear about 5.
In my case, it suddenly started ignoring my design guide, and introducing new fonts etc. It would even use several different fonts and sizes, as well as different margins for similar elements within the same page. It abandoned classes and started inlining styles. It started feeling random and, even after it realized it needed to go back to the design guide, it just continued with more of the same.
There seems to be something that happens after new model releases in both quality and behavior of previous models. It may not be immediately, but eventually there is frequently some regression.
If so, I would think that would result in worsened performance, not quality (unless you are also suggesting they may be quantizing).
I was a 4.6 acolyte from April til the fable drop, lost that quick, cancelled and took a break, came back a month later, tried opus 5 and liked it, so unpinned 4.6.
Results were great at first, and they're still not terrible, but I have noticed a regression in accuracy, so to speak, where I am pointing out issues that are quite obvious in review.
I pretty much use sonnet 5 low/medium when I have a plan to solve a simple problem and depending on scope, opus low/medium for more complex/bigger scope implementation, and only go high when it's very complex or I'm spitballing architecture/solutions and iterating plan. Never go xhigh or max.
The verbosity is insane though, opus 5 documents everything and just regurgitates whatever lead it to the design choice in there, which makes it more opaque because it's talking about something that was discussed once in a session that no one else can see (except their backend ofc)
I don't even try to steer it away from that with harness, because it doesn't work and just ends up agonizing over whether it should write some comment. Three paragraphs waffling on that on verbose output
I did however have it write a script that basically is git add -A -p for comments though, haha.
I'm $20/month, have all my telemetry toggles off, don't really over engineer prompt/context, just some basic skills for repeated patterns.
If I give it the specific instruction to "elide all revisions, corrections, and past mistakes" it usually works. You can also have Sonnet do a cleanup writing/style pass in a subagent. I impression is that Opus has been deliberately trained to keep track of all such revisions by default as a kind of ad-hoc memory mechanism. It's probably good for autonomous coding and beating benchmarks, and I presume reduces flailing when a separate session needs to pick up the work.
It was the wrong time for the GP to drop that subscription from $200 to $20, because $200 gets you a metric assload of cognition while $20 gets you nothing beyond what a local model running on your own graphics card can deliver.
Consistent, predictable behavior is valuable, even more so given the nondeterministic nature of LLMs. Nondeterminism combined with unpredictability might as well be randomness.
I'd guess you're deliberately exaggerating here, but still. I've never clocked the actual tokens/second, but I'm on the $20 plan and get ~15M tokens/month for fully utilized weekly quotas (checked couple months ago). Meanwhile the best I've been able to get locally was ~8 tokens/second with Qwen3.6 35B A3B, which is wildly painful for coding sessions and gets a maximum ~20M tokens in a month... if it's going 24/7.
Just wanted to stick some empirical data here, given that statement.
At $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally.
The point raised by this very article is that you can't depend on that. It's Flowers for Algernon As A Service.
One challenge I run into is I run many agents at once, so local resource are a tiny fraction of the total inference we’re using. Every developer with his own Pro 20x, Max, etc. accounts is letting us hit literally trillions of tokens a month.
it really angers me every day
I've had WAY better results on medium effort.
IIUC, the consensus seems to be that anything more than medium effort is rarely worth it - and you far more often run into these extreme worst cases than you do with even the lowest effort levels. That definitely coincides with my anecdata.
It's really only worth it if you're hoping to win the lottery asking it to solve an Erdos problem.
For example I'm working on a project now where I need to keep in mind details from 3 separate source repos in different programming languages, along with probably a dozen important business-related documents that either corroborate the stuff in the source repos or add additional important context. It's a lot of details for even a human to manage, and when it comes to actually synthesizing plans and reports across this sprawling information environment, anything less than Opus 5 on Extra+ tends to miss important details and make bad recommendations or draw incorrect conclusions, which then poison subsequent context.
I suspect some kind of RAG-like memory system would greatly facilitate a project like this, but even with such a system I'm not confident that I could get away with an LLM that "thinks" less hard than this. I will say that slogging through the generated documents is kind of miserable and I have to repeatedly fork off side conversations to ask for clarification, but sometimes leads the model to "realize" it's made a mistake in all of its dense babbling, and it's all very hard to interpret. I never had much interest in trying GPT 5.6 until now.
I am not a developer although I have written production code but that's not what I am talking about.
I know for a fact that I can do a task in a week. It may take 12 or 40 hours but it'll be done in a week. Now with AI? I can get that task done in a day. Anywhere from 3-12 hours.
You ship 3 hours in three with AI. The old number is meaningless now.
Opus 5 is a neverending chain of "Don't do that. Why did you do that? I've told you several times not to do that but you keep doing it."
"Thinking" for more than 10 minutes for every menial question.
And the prose it writes is horrendous, as if you're reading LinkedIn scammers. "The harsh truth! Two roads, one decision! Reality check!"
I keep finding myself typing “stop overcomplicating everything” multiple times a day as well.
Recently got approved at work for ChatGPT Pro so I could use Codex.
Blown away by the speed. It feels like using Claude Code for the first time again. I don't think Codex is doing anything revolutionary, just better handling of which requests should go to which model, and having faith in some of the "less powerful" models for more than you would think.
It seems the TUI coding experience is very much an open race. This is motivating me to look at other agents / harnesses as well (maybe Gemini, OpenCode, etc).
We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently.
That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance.
This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits.
Appreciate the outreach that you do! I love Claude, but I've been noticing reduced fidelity lately. Fable's likelihood of making a mistake increases or decreases based on the hour of the day and whether or not it's the weekend.
On a related note, and I'm happy to work on quantifying it, but qualitatively it feels like Fable's performance is noticeably poorer than initial release / launch.
I am wondering if this is the case because I use Claude via Claude Code to make a personalized care dashboard for my doctors to help me in managing my care.
I noticed in the upgraded filter announcement, https://www.anthropic.com/news/improving-fable-5-s-biology-s... ,
I hope that I'm off base here, but I noticed that the post avoids saying that the user is informed every time when such re-routing occurs. Would you be open to confirming whether or not this is the case?Is the end user informed every time their query is re-routed?
Or, can you confirm that there aren't scenarios where a user's outputs are degraded without telling them? As was the case for AI research during launch?
Why is it considered acceptable to test on paying customers without letting them know or giving them a way to opt out?
If I have a user input and then sanitize and inject that into a prompt to do something, I have no idea how much that is going to cost at all and no real way to measure this properly. A parallel example is digital ocean or aws, i can go and measure/limit my compute/fs/memory/startup times/etc and while it can be impossible to get down to the last flop of money allocated - i can run things on a real budget with real constraints, opposed to an LLM where I have to .. prerun a sanitized user prompt through a tokenizer and then ask an LLM to guess what it may do and give token consumption estimates and then act on those in any sane manner for the user?
Perhaps i'm missing something to do realistic and static rails on things but I don't see a serious way at scale to use the token billing model handling things requiring a users free text input short of having to go pander to VC money to throw money at it until someone else figures it out.
*to clarify my rambling... We should be billed and given controls based on resource usage itself and not an opaque token concept on top of not being able to spin any knobs that control it's resource usage.
The model providers are quite aligned with concerns like customer retention. These arguments only work if there is no competition. We exist in a marketplace of black boxes. There's not just "the one" you must suffer. You have options. You can build your own too.
1. Self hosting
2. Chinese models
3. Running it locally. Requires upfront cost and compromises on TPS.
I see posts all the time on HN about which models from which providers offer the most bang-for-the-buck, and how to minimize token usage and still get optimal results, so it appears that competition is working.
Hell, I use 3 different providers, and I currently don't give a dime to Anthropic or OpenAI.
They’re growing over 10x a year. They want users and revenue. In order to get users and revenue, they want to provide the smartest models at affordable prices. If they unnecessarily burn tokens, users will get less value and switch.
This thread is filled with competing comments about their monopolistic power and how when one model provider was no longer doing a good job people switched to a different one.
The competition in this market is ferocious!
Theoretically.
In reality, one sessions output tokens become the next sessions input tokens (at least if you continue the topic) so, its not as aligned as all that.
But the parent is right, when incentives are not aligned, friction will happen. Its inevitable.
I agree that incentives are misaligned but there’s several competing model providers. If one gets funny with their costs people will jump ship, especially if the gap between the top 2 labs and everyone else keeps shrinking.
These safeguards already exist when they get a whiff that you might be using Claude to fix security issues. Doesn’t seem farfetched given the incentives I outlined that they would apply to this kind of abuse.
How loose those controls are becomes a market force.
But just so we're both entirely clear on what an LLM is... it's a token prediction system.
it genuinely can't do things except recall things that have already existed.
People are having great success composing things together in new ways, but just like the english language has a finite number of sentences, and music has a finite number of chords: LLMs too are just combining things that have existed.
I don't want to sound condescending, it is remarkable how useful this technology is, but please don't evangelise them on capabilities that they genuinely can never have.
Laptop computers have incredible processing capabilities but nobody expects them to be able to walk your dog, no matter how useful they actually are at doing other things.
Knowing how hard something will be to do before attempting it is precisely the sort of impossible thing that it couldn’t do.
a) if something is possible
b) if something has been requested to take a long time ( a signal of abuse, like requesting illicit pictures in image generation)
is actually somewhat straightforward (I mean, if they are able to predict if something is a “substantial piece of work” as they seem to do already).
The mathematical proof thing is obviously marketing spin, you should pay more attention to what mathematicians are actually saying instead of hackernews folks.
These things are really good at being search engines and harnesses for iteration rather than some kind of advancing intelligence.
https://cacm.acm.org/research/formal-reasoning-meets-llms-to...
When you’re doing this 100K times per day you get an extremely good idea of what it costs. You also have all the tools to see when something starts changing quickly.
This change is for Claude Code the harness. If you’re using the API at scale and paying full price then you get exactly what you put into the request.
https://code.claude.com/docs/en/admin-setup#set-up-usage-vis...
Goas back to factory towns, gift cards, game money or MtG.
"We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently. That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance. This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits."
LLM doesn't seem to be keen to put in effort either!
Is this AGI?
"Opus 5" is actually Opus 4.8 in a trenchcoat, with new guardrails
"Opus 4.8" is the old Opus 4.6 with lipstick on it, with new guardrails
"Opus 4.6" which everybody used to love, is now actually running the old Opus 3.5...
How would we be able to tell?
The chat-based models are obviously being lobotomized based on personal usage and general load (e.g. PST business hours are worst).
API doesn't seem to be affected by this.
In that case, the agent will respond incorrectly because it has no visibility into what reasoning mode it’s in.
The mood change in just six months is wild, in February this year Claude was the most liked LLM by far.
but the incentives of these companies to become profitable at any cost slipping into entire new types of dark patterns around token based billing seems gross
- charging for injected prompts and cot tokens
- changing default thinking effort to be higher
- training models to give longer winded answers that don’t say anything more of substance
- refusing to fulfill a request and still charging you
i wonder if you could ever just charge based of each user message and it so how breaks even across short and long replies
EDIT: my honest opinion; Anthropic is building a person, whereas everybody else (it seems) is building a tool.
I realized I was spending most of my tokens arguing with Opus and trying to get it to let go of stupid, lazy, obviously incorrect premonitions. I wound up canceling my 200, bought a pair of Sparks, and am running full fat DS4 Flash and so much happier. Done with being at the whim of these companies.
I have been using GLM-5.3 in my home setup and it is very good in comparison.
I didn't really care about their mobile app and the worst experiences were sometimes the sub-agents within OpenCode freezing and refusing to report their status (though this also happened with Kepler by the GitKraken folks).
I maintain a Claude subscription for Fable but seldom use it.
https://github.com/ferrislucas/Circus-Chief
I was approved.
3 months later, my approval was degraded into "in review" (revoked). I'm sure my account was flagged based on contents of debugging/researching firmwares/etc.
I opened a support ticket. No response. I opened another support ticket. No response.
1-2 weeks later, I got a response that I will not be re-approved and I need to reapply. No problem.
The page to reapply on does not allow me to re-apply because it my account is stuck in an "in review" status.
https://github.com/anthropics/claude-code/issues/84352
The community thinks it's a bug. I'm 95% sure it's not and a bunch of us who were previously approved had it revoked due to flagged content and will not be reapproved.
I switched to Codex + got TAC approved instantly and have not looked back. It's a shame. That's 100% separate from whatever the heck the quality of Opus 5's outputs are. The way it talks... insane. I would bet a good amount of money their next release will focus "reduced simplified responses" if I had to guess.
$2t company by the way
Meanwhile the Chinese models are "go ham dude"...
If it was not for capacity issues, Chinese models have a higher change to just dominate.
> $2t company by the way
It used to be that OpenAI and Anthropic had such a moat around them, that such a valuation was worth it. But these days, its gross overvalued (like so many).
The more stuff is being pulled like cyber verifications, downgrading effort levels, downgrading usage (OpenAI), the more people move to those Open Weight Chinese models.
A fun recent event ... https://opencode.ai/data/
When DeepSeek Flash 0731 came out and provided a massive jump in cheap inference capability. It resulted in a 10x increased OpenCode token usage.
It took a 2.5x to 5.0x price increase AND a reduction by 4x usage (later to 2x) usage, and several cheaper models + a free model, to push the traffic down.
Traffic towards open weight models is increasing, even if providers can not keep up with the influx of new customers. This is not something you want to see as two companies, trying to go for IPOs.
So the idea of stonewalling cyber capabilities, when the rest of the world is just doing whatever with open weight models, on their own hardware even! This entire strategy from Anthropic never made any sense.
Not convinced here.
The problem is that Anthropic seems to be getting away with selling one thing and delivering another. You pay for Opus, you get something else etc.