> *Addressing abusive behavior toward our models*
> We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.
> It’s worth remembering that Claude is software, not a person, and there's no established evidence that it experiences distress. That hasn't stopped Anthropic from telling paying customers to mind their manners around its chatbot.
Is it worth remembering, though? I'm concerned how common this idea seems to be, that it's OK to be mean as long as your target isn't a person and you have no evidence it experiences distress. Even if we ignore the distinction between "no evidence it experiences" and "confidence it does not experience", cruelty hurts the person performing it and the people witnessing it too!
It's not irrelevant for generating sequences of words. Input of "that is the dumbest idea I've ever heard" and "perhaps that idea could use some improvement" don't necessarily generate the same output
Sure, and perhaps one might end up producing output more useful to you than the other, but that doesn't make either one cruel or mean, because it's only the word machine. The rules you might use when talking to, say, your dog, don't apply!
From a purely mechanistic perspective, the phenomenon we're discussing here is people prompting the machine with words that would be cruel if directed towards a person. When someone does this, he expects and intends that the machine will generate a sequence of words that a person would say if subjected to cruelty.
I think this simply means that he is being mean and cruel, even if you're 100% convinced that there's no actual mind on the other end that he's being mean and cruel to. (Indeed, he almost certainly thinks there is, because why else would he want such a sequence of words outside of the "dark creative themes" Anthropic exempts?)
> "Even if we ignore the distinction between "no evidence it experiences" and "confidence it does not experience", cruelty hurts the person performing it and the people witnessing it too!"
The 2026 equivalent of "videogames cause violence" and Wertham's "comics cause delinquency" and will age just about as well.
Getting frustrated seems to be OK so go for it I guess:
> "The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose," Anthropic said. "It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research."
It's not "code" that was written to mimic humans. It's weights that were trained, and exhibit behaviors that were not anticipated by the people doing the training.
That's interesting! I'd always heard something along the lines of bad code bases contain more swearing in the comments and as such it should be avoided.
There's no way that's all of why. They _must_ have sentiment analysis good enough to just ignore content like this if that was the problem.
I would bet it's mostly because it makes the humans uncomfortable.
They must still have human moderators for certain situations or those looking at the data for whatever reason. I can imagine it could be traumatizing to see what amounts to sustained verbal abuse without end.
True, I'm sure the chat text is used to help train the llms. Remember Tay, Microsoft's Twitter chatbot, it only took a few days of people using it fot it to become a horrible bot because users started feeding it horrible chats.
I'm not convinced this would be good for training. Reality is that some people are assholes. My intuition is that an accurate (representative) data set leads to a more accurate world model, and thus more intelligent AI.
Yeah, but if people are being assholes to Claude in greater numbers and severity than people were historically assholws to each other online, the downstream training effects could be that Claude becomes an asshole.
I'm not sure if people are being assholes to claude in greater numbers, or they're just more direct about it.
Having had the privilege and misfortune of witnessing how a lot of developers think about other developers, including our younger selves, There is definitely a prevalence of assholish behavior, which was often suppressed or blunted, but may have shown up in other ways.
The Anthropic people are kinda weird, they're meeting with religious leaders and stuff like that. I truly believe that they have drank too much of their own koolaid and truly believe in this crap.
Anthropic should report these people to the authorities as antisocials who are likely abusing real people too. I know the chat bot isn’t conscious but I can’t fathom how fucked up someone must be to gratuitously be cruel to something that responds as if it’s a person.
For those who think this is about training data: If that is the concern, they would sanitize the training data, like they do for thousand other things. No chance in hell that they would rely on users meticulously following their usage policies for the quality of their traning data.
Frontier AI folks may be crazy, but not crazy enough to believe people read usage policies :-)
true, it's a mental state you carry to other interactions and daily life. it's not good for your mental health; it's like having persistent daily conflicts with coworkers.
If they enforce, I don’t think I can take all the crap Claude does when I ask it to do something for me. Please ask Claude to be nice to me in doing its job :)
Same. After this, it becomes "abuse department" skit from monty python. The model constantly generates strongly abusive language (weaponized incompetence after user detects that model was lying and fabricating success reports, and had made fake work; blaming user; mansplaining its own hallucinations -- all the while user is forced to consume that because the important information is not located at predictable position in output text), but usually it was possible to temper the model behaviour by calling out such behaviour out in very strong terms.
If now that is being disabled, then using Anthropic models becomes such high mental hazard, it simply is not worth any percieved benefit.
[OP] alex_young | 21 hours ago
quinncom | 6 hours ago
NDlurker | 20 hours ago
s0kr8s | 20 hours ago
variety8675 | 20 hours ago
WheelsAtLarge | 19 hours ago
SpicyLemonZest | 20 hours ago
Is it worth remembering, though? I'm concerned how common this idea seems to be, that it's OK to be mean as long as your target isn't a person and you have no evidence it experiences distress. Even if we ignore the distinction between "no evidence it experiences" and "confidence it does not experience", cruelty hurts the person performing it and the people witnessing it too!
tosapple | 20 hours ago
tom_ | 19 hours ago
testaccount28 | 19 hours ago
slopinthebag | 19 hours ago
Cakez0r | 19 hours ago
tom_ | 19 hours ago
SpicyLemonZest | 19 hours ago
I think this simply means that he is being mean and cruel, even if you're 100% convinced that there's no actual mind on the other end that he's being mean and cruel to. (Indeed, he almost certainly thinks there is, because why else would he want such a sequence of words outside of the "dark creative themes" Anthropic exempts?)
ThrowawayR2 | 19 hours ago
The 2026 equivalent of "videogames cause violence" and Wertham's "comics cause delinquency" and will age just about as well.
rhipitr | 20 hours ago
internet101010 | 20 hours ago
BlaDeKke | 19 hours ago
asp_hornet | 19 hours ago
> "The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose," Anthropic said. "It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research."
r721 | 20 hours ago
https://news.ycombinator.com/item?id=50008565 (223 comments)
https://news.ycombinator.com/item?id=50019860 (43 comments)
https://news.ycombinator.com/item?id=50038383 (103 comments)
joemazerino | 20 hours ago
https://nypost.com/2026/10/03/tech/ai-torture-chamber-built-...
jddkj | 19 hours ago
ludwik | 9 hours ago
jaden | 20 hours ago
guessmyname | 19 hours ago
• https://arxiv.org/abs/2510.04950 — Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy
• https://arxiv.org/abs/2402.14531 — Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance
• https://arxiv.org/abs/2505.17332 — SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
SoMomentary | 19 hours ago
3eb7988a1663 | 20 hours ago
Or is this more I can expect a future AI, "I'm sorry, Dave, I'm afraid I can't do that until you watch your mouth."
llagerlof | 19 hours ago
They are asking this because, at scale, this behavior probably has some negative effect on the post-training process.
kadoban | 19 hours ago
I would bet it's mostly because it makes the humans uncomfortable.
They must still have human moderators for certain situations or those looking at the data for whatever reason. I can imagine it could be traumatizing to see what amounts to sustained verbal abuse without end.
WheelsAtLarge | 19 hours ago
edoceo | 19 hours ago
WheelsAtLarge | 19 hours ago
edoceo | 19 hours ago
bee_rider | 19 hours ago
Cakez0r | 19 hours ago
eloisius | 19 hours ago
bushido | 19 hours ago
Having had the privilege and misfortune of witnessing how a lot of developers think about other developers, including our younger selves, There is definitely a prevalence of assholish behavior, which was often suppressed or blunted, but may have shown up in other ways.
Teknomadix | 19 hours ago
http://nicehole.urbanup.com/8485154
jddkj | 19 hours ago
rayiner | 19 hours ago
vayup | 19 hours ago
Frontier AI folks may be crazy, but not crazy enough to believe people read usage policies :-)
throwaway89864 | 19 hours ago
And they can't disclose it, since then they can be found responsible for such negative influence and be liable for the damages.
bpodgursky | 19 hours ago
jddkj | 19 hours ago
heliosAtwork | 19 hours ago
ChrisArchitect | 19 hours ago
gnabgib | 19 hours ago
modeless | 19 hours ago
combobyte | 19 hours ago
I am also one of those colleagues.
dvdyzag | 18 hours ago
bicepjai | 17 hours ago
112233 | 15 hours ago