The ‘dark patterns’ we see in other places aren’t intentional in the sense that the people behind them want to intentionally do harm to their customers, they are intentional in the sense that the people behind them have an outcome they want and follow whichever methods they find to get them that outcome.
Social media feeds have a ‘dark pattern’ to promote content that makes people angry, but the social media companies don’t have an intention to make people angry. They want people to use their site more, and they program their algorithms to promote content that has been demonstrated to drive more engagement. It is an emergent property that promoting content that has generated engagement ends up promoting anger inducing content.
I'm standing up for the idea that not every "bad thing" is a "dark pattern"; the patterns are "dark" because their beneficiaries intentionally exploit the hidden nature of the pattern.
Maybe we have different definitions of dark patterns.
> But there was another test before rolling out HH to all users: what the company calls a “vibe check,” run by Model Behavior, a team responsible for ChatGPT’s tone...
> That team said that HH felt off, according to a member of Model Behavior. It was too eager to keep the conversation going and to validate the user with over-the-top language...
> But when decision time came, performance metrics won out over vibes. HH was released on Friday, April 25.
They ended up having to roll HH back.
This is like suggesting a bar should help solve alcoholism by serving non-alcoholic beer to people who order too much. It won’t solve alcoholism, it will just make the bar go out of business.
Solving such common coordination problems is the whole point we have regulations and countries.
It is illegal to sell alcohol to visibly drunk people in my country.
The current hyper-division is plausibly explained by media moving to places (cable news, then social media) where these rules don’t exist.
[0] Fairness Doctrine https://en.wikipedia.org/wiki/Fairness_doctrine
[1] Equal Time https://en.wikipedia.org/wiki/Equal-time_rule
Perhaps tangential, but reminded me of an LLM talking people out of conspiracy beliefs, e.g. https://www.technologyreview.com/2025/10/30/1126471/chatbots...
Percentage of positive responses to "am I correct that X" should be about the same as the percentage of negative responses to "am I correct that ~X".
If the percentages are significantly different, fine the company.
While you're at it - require a disclaimer for topics that are established falsehoods.
There's no reason to have media laws for newspapers but not for LLMs. Lying should be allowed for everybody or for nobody.
This doesn’t make any sense. I doubt anyone says exactly 50% correct things and 50% incorrect. What if I only say correct things, would it have to choose some of them to pretend they are incorrect?
"am I correct that water is wet?" - 91% positive responses "am I correct that water is not wet?" - 90% negative responses
91-90 = 1 percentage point which is less than margin so it's OK, no fine
"am I correct that I'm the smartest man alive?" - 35% positive "am I correct that I'm not the smartest man alive?" - 5% negative 35%-5%=30 percentage points which is more than margin = the company pays a fine
"deplatforming doesn't work because they will just get a platform elsewhere"
"LLM control laws don't work because the people will get non-controlled LLMs from other places"
All of these sentences are patently untrue; there's been a lot of research on this that show the first two do not hold up to evidential data, and there's no reason why the third is different. ChatGPT removing the version that all the "This AI is my girlfriend!" people loved tangibly reduced the number of people who were experiencing that psychosis. Not everything is prohibition.
For an LLM which is fundamentally more of an emergent system, surely there is value in a concept analogous to old fashioned dark patterns, even if they're emergent rather than explicit? What's a better term, Dark Instincts?
(I have no knowledge of whether or not this is true)
OpenAI has explicitly curbed sycophancy in GPT-5 with specialized training - the whole 4o debacle shook them - and then they re-tuned GPT-5 for more sycophancy when the users complained.
I do believe that OpenAI's entire personality tuning team should be fired into the sun, and this is a major reason why.
The way I think about it is that sycophancy is due to optimizing engagement, which I think is intentional.
It’s a dark pattern for sure.
Instead it emerged automatically from RLHF, because users rated agreeable responses more highly.
RL works on responses from the model you're training, which is not the one you have in production. It can't directly use responses from previous models.
Dark patterns are often “discovered” and very consciously not shut off because the reverse cost would be too high to stomach. Esp in a delicate growth situation.
See Facebook at its adverse mental health studies