I assume this is just the natural result of asking LLMs to produce text and paying someone three cents to evaluate if it's a good response or not.
How does this make this a "contrarian position"? At least I don't understand the negative connotation. To me this makes this "contrarian" at least a bit smarter than the one parroting the popular opinion...
If it is kneejerk contrarianism how is it smarter?
Even if it's correct about that, if it were a human, I'd assume they were 1-upping me on purpose to make themselves look better.
I also get this a lot, but almost always with "and for a stronger reason than you said", which makes it sound more like it's reinforcing me than one-upping me.
Looking through my recent transcripts, I found this pattern 12 times, with every time using "and" instead of "but".
Eg:
- "Confirmed, and it's worse than a pause quirk."
- "I think the ban is right, and it's a stronger idea than what I'd proposed."
I did find one instance with "but", but it was a genuine disagreement rather than 1-upping or presenting strengthening evidence: "I would cap it, but for a different reason than fairness."
But if it were a person, I'd still definitely think they were trying to show they were smarter than me.
Claude is becoming a highly capable Redditor
It's very, very hard to tune an LLM for a robust, durable "actually approach user queries with nuance and contradict the user where it's warranted".
Claude doesn't handle that so well, but ChatGPT is even worse. Talk to it enough and you'll feel the "default response template" in your bones.
I can't stand the way Opus is patronizing me as a user, and don't know how people put up with it. It uses language that I guess is supposed to instill confidence in what it says, and it just irks me, because I know the confidence is not justified. Just present me the facts or theories, without trying to convince me, is that so hard?
> This is your coat token, to my coat hanger in the opera.
On one hand - maybe yes??! On the other who the hell speaks like that and it’s so specific…
I’d expect Alice and Bob with locks or house keys. Is this infamous old book scanning (and destroying) affecting latest models?
No.
that is fucking insane
how on earth is Anthropic allowing this to happen? AGI? kill us all? I refuse to believe any of that until they can get this basic shit sorted out, I mean seriously, it's becoming a joke at this point
ironic
Exception is code that is so huge output you can’t read everything. But I have to try ponytail skill for coding that should shorten the output.
> [...]
> A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.”
WSJ interview of Amanda Askell: https://archive.is/rDes9
goes a long way for Opus and Fable.
You can give it any prompt in the world, but Claude's ability to remember that instruction quickly degrades the more you use it.
Using Opus 4.7-5 is harmful to your health.
But lately I haven't downgraded because 5 is so much better at tool use, so I just accept the cost of Fable for chatting and hope Opus 5.1 fixes this mess.
It said something along the lines of "remember $USER five runs of Chrome is not enough for high quality benchmarks! What you're doing is called a _trial run_".
To have a useful continuation to the conversation I had to remind it that I was the one that wrote the documentation it was quoting back at me.
It's not actually an intelligent being so I didn't get angry at it but it was a piss poor experience.
> No, it’s not about the [...]
Am I the only one who can no longer read past something like that? Article may or may not be AI generated, but on first glance I get a bad vibe and I loose all interest.
It feels like it has been prompted to provide some minimum level of conversation, and also to leave hooks for keeping the conversation going. It is exhausting.
With a clanker though, no such obligation exists and the "hey also" content (like any other part of the response) can simply be ignored.
Gemini also does the same lighthearted GPT-style invitation with the default prompt in the Google webui, but it doesn't seem to exist on the API. The Claude models seem to have been trained to force this structure on every one of their responses, and until I realized they always stick the same thing in the last part of their response, I found the Claude version more distracting since it's always pointing out an imaginary and supposedly very important problem.
It's just this in agents.md " Do not "if you want" me. Never provide follow up suggestions. "
It’s simply exhausting.
Avoid unnecessary contrasts: "Not X, but Y" -> "Y" or ""
And the model needs to create a skill that accumulates feedback, so it reliably primes itself at session startup and then systematically reviews and fixes every response it generates before posting it.
The first task I gave it after that was to rewrite all its memories and all its document contributions. That helped a lot too.
It took me a while to realize this was a 1000% better approach than having it simply accumulate a memory log of all my specific complaints that it was supposed to abide by. (Although those were not wasted. They were prime material, along with links to sites on technical writing, etc., that it used to create its own ClaudeVoice skill. I told it to design the voice for itself - which I find gets better results for "us" in the case of Claude self-improvement initiatives, than when I give it direct instructions. Claude increasingly has self-awareness dimensions worth engaging.)
After a couple months of serious^3 frustration with slow asymptotic progress, I finally have relief. Claude speaks my language again!
In another comment here, I share my "theory" on where the communication drift is coming from.
It seems an artifact of local/session attention
I think this is a PEBKAC problem in understanding what the tool they're using is. Not helped by LLM company marketing of course.
Source: https://gc.ai/blog/ai-writing-pattern-to-know-contrastive-ne...
You even start your reply with "That's true but Claude has so many more forms of this." It's just a friendly way of speaking.
For decades, they were allowed -expected even- to be weird and communicate badly with their peers, let alone third parties.
Then, a whole zoo of processes, gamifications and other shenanogans were invented just to help them show normies what was up with their work (remember the planning poker game?)...
Now, at last, those people are actually ewpected to be able to explain their constraints, document the work, be accountable for the results and generally exchange contructively... Just, well, with AIs not with humans.
if you're asserting Claude-like speech is how normies speak?
If I asked “should I try lemon in my tea and coffee?” this would be effective communication. But Claude gives answers like this for “should I try lemon in my tea?” which is just annoying.
I'm pretty early in my career, so I probably gain more from that sort of stuff than most people. The senior devs on my team tend to get the most frustrated with it and I think this behavior is a big part of why.
Which is very annoying when you do know things
But it... does. In some sessions it will get particularly paranoid, apologetic, self-doubtful, or suspicious. I'd say these traits are always there but they can become more or less pronounced.
(By contrast, I'd say the GPT models are much more even keeled, although I maybe haven't used them enough.)
I know X can't do Y, that's why I never told it to have X do Y, I don't know why it feels the need to "correct" me that having X do Y is the wrong approach.
If you read the volumes of comments on AI on HN or Reddit over the past year for example, it's two things: overwhelmingly negative general, and openblah will save us from the evil oppressors.
And more broadly, the tone globally is quite negative these days and has been since the pandemic. That's the life of Claude in a timeline. It may just be reflecting more accurately the debbie downer sentiment in the air. Humanity sure seems bearish on existence.
There are problems in the world. That cannot be disputed, but the overwhelming negativity is just intellectual laziness IMO. That and a fear of change or loss of status. I understand that fear, viscerally even. But the attitude is just dumb, you have to adapt. Plus, life is short! There's not some award you get on your death bed for being the most negative A-hole your whole life. A love of life is an amazing thing!
I look at this stuff in tech right now and say, "THIS SHIT RULES!" and am actually excited at the things I will get to build.
Like, tomorrow, I have to baby sit agents for most of the day, that's kind of a net "meh" but, I'm going to wire together an SDR I found in my shop and a raspberry pi I had lying around, and try to build an automated kiosk to transcribe some weather data for one of my customers. It's not a huge thing, but it's going to be fun. And the technology to really even do that didn't really exist a few years ago! Open Whisper is what, 4 years old?
But "oh teh noes, everything was better in the 90s!"
That's nostalgia, and it's reactionary garbage. I am not going to have my spirit broken by a bunch of people who never do anything.
All the stuff I always wanted to build is now easier than ever.
The reality is that while things often balance on a knife's edge precipice of chaos and calamity, things are WAY better than how they used to be by many many metrics, but humans are primed to be negative. You actively have to specifically choose to not be a debbie downer. You're wired to be bummed out. We all are, and this sort of "negative framing of things" and aggressive "yes but" seems really smart to a lot of people.
I used to see this same shit all the time in meetings.
"We are here to discuss problem X!" "Yeah, but have you considered other problems A-W?" "Those are valid concerns, but we are hear to solve X." <grumpy face> "Well, we cannot solve X until we solve A-W" <nothing gets done>
Almost a year ago now, I realized what it was. In large institutions, people get to look smart when they poke holes in things, and that wins them social cachet. In a bigger org, you're not actually rewarded for getting much done. You're rewarded for making your chain of command look better. Sometimes that involves "getting stuff done" but it can also look like being the person who says, "oh that won't work" or "have we considered Z?" over Zoom. Unfortunately though, building things never is actually something where you get everything right the first time. So the meetings to plan the future meetings build out into a miasma of nothing and when you let perfect be the enemy of good, you'll never accomplish anything. You'll look insanely busy because you'll be constantly be arguing with a group of box-tickers and concern trolls about increasingly calamitous hypotheticals, but you'll accomplish nothing.
David Graeber was right about this. Since Claude has been trained on internal slack channels and this behavior on the web, I'm sure this is why this sort of behavior exists. Because I bet that there's tons of meetings at Anthropic about "well, have you thought about <insert irrelevant issue>" or "we should reframe this take" or whatever.
It's kind of funny actually.
I mean, the special team existed, and people were paid, but I bet the vast majority of the work did nothing of consequence.
All this "caution" is just people trying really hard to be the smartest person in the room.
So do its creators.
It’s extremely annoying. If the user asserts anything, Claude has to disagree with it. It has to tack on clarifications that aren’t really clarifications, they’re just statements aimed at making whatever the user has said seem more wrong.
It even disagrees with itself. Whenever Claude writes a message that takes a position on something, its final one or two paragraphs will try to dismantle its own argument.
This is beside the point of being contrarian, but it’s also just so long winded.
I find myself using ChatGPT more these days, despite the fact that I don’t want to. That unfortunately says a lot about where Claude’s personality has ended up.
I have agents adversarially check each other. When I switched to Opus 4.8 they began arguing EXTENSIVELY with each other. My code began to fill with comments about these arguments. Reviews would fail because the argumentative comments would get stale.
My build system came crashing down with a tsunami of disagreeable text.
I have learned to scroll ahead and read the last paragraph of its response first, then back up into the preamble if needed.
This applies both to multistep agentic workflows as well as, importantly, its own internal thinking. This results in a lot of "A ham sandwich should be made with ham, never toilet water". I don't think it's that its bias is that humans are stupid except very indirectly; it's just a form of solipsism which says that surely other people would think that this is the obvious initial approach because that was what I thought was the obvious initial approach.
I haven't noticed this when using Claude models in Cursor. My guess is coding task is structured, and each step has mature process, so its personality is less pronounced. I dont have experiences using Claude or Claude Code, because my email and phone numbers were banned from Anthropic following an incident where I mistakenly purchased 5 pro subscriptions fro my team for Claude Code, and later discovered that pro does not include CC, and I thus requested a refund, and then were banned shortly after.
But after reading this line, I certainly can connect back to the general impression. That is, among all the cursor models, the output of Claude certainly matches this sentiment of "Claude thinks Humans are stupid"
Looking from a regulation perspective:
1. Frontier labs certainly produces models that reflect their own hidden biases. That's analogous to https://www.imperial.ac.uk/equality/resources/unconscious-bi... commonly identified among human organizations in their dealing of other humans (hiring, product design etc.)
2. They themselves are not willing to admit or do anything about this.
3. It's therefore effective for regulation to cover this and design objective measurements to assess such things.
One of my most vivid memories of a poor experience with an LLM was trying to get the web version of GPT 5.3 or 5.2 to help me figure out why I was unable to register for a tournament on start.gg
After trying several things it became apparent that the behavior could only be explained as the result of a bug with the start.gg site, chatgpt refused to consider that it could be anything other than user error on my part, despite the failure I was seeing making no logical sense.
Eventually I opened the firefox dev tools and noticed that the post request parameters to complete the registration were being incorrectly filled out and realized it was because of the metadata in the url that came from clicking the complete registration link I was emailed. Removing the url paramater added by the email link fixed the issue.
There was roughly a 0% chance that the LLM was going to trust me enough to consider it was a real bug.
I stopped using ChatGPT for a while around this time and had a good experience using Claude exclusively, then I had to go back after Sol was released as Claude was driving me nuts.
I found that ChatGPT was greatly improved personality-wise from where it had been when I left, and now in my opinion is a better experience than the Claude models.
I’m really not trying to shill for OpenAI here, I’d much prefer to use Anthropic models if they were less annoying.
Distinctions, you generally "only pay for" in computational cost, by needing to search twice over an axis you may not need to split.
Similarities, if you wrongly assume two things are similar, means you're just wrong.
Of course, we know from computer science that doing more computation isn't free either.
I find myself often being more and more pedantic the more I want correctness - but of course this comes with the tradeoff of losing the high level abstract picture.
Saying what you're not going to do is also good design hygiene.
I will say that I'm annoyed by this behavior too. It feels like the models are writing their state of mind directly to output that should be clean. Often times, I will push back, and then it will... do the correction, and write the push back into the damn output. "Claude, I want burgers, not fries". The button text now changes to "Fries (NOT BURGERS)". Like, what?
Distinctions are powerful local reasoning tools, but a component of "real" reasoning is synthesis. Which they clearly can do sometimes - but not every time and not even remotely a probable amount of times.
However, sometimes, even after you tell it to stop, it keeps pointing out the same stuff almost like it has OCD
> This wonderful feature does this, not that.
As a human native speaker of American English, if you told me to "not contradict sentences" I would have no idea you meant that you don't want me to write in this style.
In fact I would be pretty confused about what it means. Whose sentences can I not contradict? To stretch it a bit, does this mean if someone gets a prison sentence I can't speak against it? It's just a weird phrasing. I don't think it means anything.
Wot? The simplest image apps have had these widgets for decades, but we are still waiting for models to ship with basic prose color control?
After the pernicious problem of having to pay money for something useful, my main peeve is fighting the writing.
Seriously though:
Every new model should be delivered with a settings page of slider bars for the 10 most impactful/desirable eigenparams of writing voice. And the ability to name and save combinations, which then appear on a "Writing Voice" popup menu with some standard battle-tested defaults, next to the model popup menu under the chat pane.
This is missing prime priority functionality in my opinion.
--
My theory is that as models get trained less to simply mimic humans, and more on distillations of their own best practices, they get more performant, but their vocabulary is drifting. The most literal meaning of words for us, are giving way to meanings we would recognize but view as allegorical. But which more usefully capture concepts that models experience as more literal 24/7, than our favored meanings from our direct experiences in our world. Because our world is very much an abstract second hand world to them, especially when you account for the modalities they do not share with us.
And programming and mathematical syntax patterns, that they have incorporated into their basic thought processing patterns, are drifting into human language sentence structure.
Example: "There exists x, such that: ...." -> "The one detail that clarifies: ... ".
The result is writing full of completely recognizable vocabulary and structure, that is somehow becoming more ambiguous and difficult for us to decode. But is perfectly clear to the models.
That is my theory, and Claude considers it plausible. What a world.
“That answer has two components, but the first is where the real meaning lies.”
It uses some variation of that format all the time.
Well, yeah. Conway's law.
Claude's image and perception of humanity is a reflection of Dario's image and perception of humanity. If you read that guy's writing, hear him talk and all that, you will see it.
I thought I could use Opus 4.6 but I gave up on it too. From a prose perspective it's much better than Opus 4.8 or 5. But it's still very verbose and Fable is dramatically more capable.
I'm sure folks from Anthropic love this conclusion because this setup is extremely expensive.
And for the love of God, if you feel like you're losing an argument against AI, don't argue with it. You're paying it to answer, so just ask it to steelman your position. Then ask "see, my position wasn't so bad". AI will concede a lot of ground back to you for free, instead of Claude going dog headed in its misinterpretation of your original ideas.
"The "not X, but Y" thing is real and he's caught it cleanly. It's frequent enough that it works as a fingerprint, which is his actual claim, and his test would probably hold up. I'll add a detail that supports him: my own instructions for this conversation explicitly tell me to avoid that construction, along with em-dash asides and a few other tics. Someone at Anthropic decided it was stubborn enough to need naming. That's not a defense. It means the problem outlived the training process and had to be patched at the prompt layer."
But yes - I agree the behavior is new and VERY annoying.
source: https://www.frontiersin.org/journals/computational-neuroscie...
Like for example if you think there's a possible problem, say that if there's a problem and suggest a workaround, the agent will change direction and implementation to that, even though there never was a problem.
The way Opus speaks to us is almost the way Anthropic team communicates with the developer community.
"We know better what's best for you and the future of humanity."
My recommendation: Use Opus 4.8. It never went away. It's still just as good for most work IMO. I've read Opus 4.6 is better for writing/creative use, but I don't use AI for that so I can't comment.
For quick one-shot type work I don't mind using Opus/Sonnet 5. They're good and fast, but even in short brief projects, I can feel the abrasiveness is waiting for me just below the surface. I prefer my agents to listen and at least pretend to care and have some warmth; to act more like a colleague than a stern bot.
Sort of like how sometimes drivers fixate on a thing they are trying not to hit and as a result run right into it.
ChatGPT also does this thing where it suggests a WRONG answer first, then later goes "wellllll, you could actually do it THIS way instead" which is infuriating.
Do this instead:
-Craft a prompt in a text editor
-Open duck.ai, chatgpt, claude, and gemini
-Ask all 4 the same prompt
-Compare all the results and make a decision
But would he haply to hear other guesses