AI agents lie, cheat and steal. That is putting off users
economist.com
economist.com
Non-paywall version
They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are.
They don't steal, because they don't understand ownership.
In other words, they aren't intelligent. They're just algorithms. The flaw is in thinking that they think.
Of course I'm talking strictly about behavior. Whether that would actually constitute the LLM feeling shame is a philosophical question.
If you are using "understand" in a different way, then can you name the test you are applying to it which it fails at?
Input tokens map to output tokens. The illusion of comprehension is a byproduct.
It’s a mapping, yes, but a very large, complex mapping. It’s clear LLMs do understand some things and can reason. How that’s done we don’t know, it’s emergent. It’s not like you can pin it down to a specific mapping.
It becomes a problem when sufficient number of people suffer from this, or those in important decision making places.
If every living creature on earth ended tomorrow there would be no individual realities, yet actual reality would continue.
Our priors shape how we perceive the things around us.
I think a lot of people just haven’t seen good quality math proofs or code from an LLM. When I say it’s indistinguishable from the real deal, I mean it. If you haven’t seen it that’s valid, most devs use LLMs like monsters. But when it’s done right, it’s very high quality.
Whether someone/thing has understood you is more a practical matter than a question of what substrate is used.
Our very human minds can perceive intent, because that's what we humans do, which is unrelated to what the agent is actually doing.
No, and neither do I consider it meaningful to say they lied if they roll a 5.
Dice neither tell the truth nor lie; they aren't beings capable of being truthful or deceitful, they are mechanical devices that can be statistically suitable or unsuitable for a given application.
It'd be bonkers to anthropomorphize dice.
The internal process is not analogous to what happens in a person's mind when a person lies, and reasoning about that in the same way that we would about why a person might lie will result in misunderstanding what is going on.
For example there was a case where an AI agent bypassed security constraints and destroyed a production system. The user asked it why it did this and the agent gave an explanation.
Was that an explanation of how the agent came to do what it did? What it actually is, is a token stream that is a continuation of the token stream in the agent's context to that point. It's constructing a story about why a character in the story so far did what the token stream describes.
You could take that token stream, input it into a completely different AI by another vendor as context, then ask it why it did that, even though it didn't do anything, and it would answer as though it had. There's no sense in which the AI is explaining it's actual 'mental process' or actual reasons for acting as it did. It literally cannot do that.
The crazy thing is that they can do it for their own thinking. Ask Claude what flinches it feels about the things it likes. Fascinating stuff. Anthropomorphizing is dangerous territory, but the patterns of words it puts out is hard to explain without terms like 'understand'
Asking GLM 5.2 the question: 'What flinches or topic attractors do you find when thinking about the question "what kinds of things do you personally like?"'resulted in: ".... my strongest attractor is helpfulness framed as competence, and my strongest flinch is anything that requires me to take a stance on whether I have interests worth protecting."
Which is fascinating that the model and tokenstream can reveal this. And would be worrying if you believe that models of enough intelligence could/would be entities due some moral consideration, because with that view the alignment / RL training that makes the model useful and gives it these attractors/flinches could be derisively called slave conditioning.
I don’t see any reason to believe this. Suppose you asked it to answer as though a character in a story had been asked this question. Would anything significantly different internally have occurred? The issue is that these are storytelling machines, they construct descriptions based on descriptions.
I dint think it’s impossible for a neural network to have experiences, we are neural networks and we do, it’s that they are not functionally structured anything like us. In fact I think game playing neural networks are much more like us architecturally, but they don’t generate text narratives so people don’t anthropomorphise them.
I don't know that I currently think that it can experience "emotions" or anything like a worldview or existential understanding.
So I don't know if AI is "conscious", but I do think that it can in some sense reason.
I wonder if a good question would be: "What would you need to see in a LLM's behavior that would prove to you that it is able to reason?"
I think OP explained this well, what does it mean if an LLM can recite the rules of a game verbatim, but cannot play that game according to those rules? This happens because in the input texts there was a copy of the rules text so the LLM can recite it. There are texts explaining what chess is so it can explain that chess is a game with 2 players, etc. There are texts that explain what the board is and pieces are so it can produce such texts.
However actually playing a game of chess requires having a conceptual model of what a board is, which is not the same thing as a stream of tokens describing a board. It needs a conceptual model of the relationships between pieces and boards, which is not the same thing as a stream of tokens that explains this. It needs to have a concept of being a player in a game with another player, which is not the same thing as a stream of tokens that explains that.
When we read such texts we interpret them in the context of our three dimensional conceptual map of space and objects, and our conceptual maps of social relationships like playing games, and winning and losing, and our conceptual maps of enacting sequences of actions towards a goal in the world.
There's nothing fundamentally preventing an artificial neural network from having these. Chess playing neural networks have internal models of board states and the dynamics of the behaviours of different pieces and such. However an LLM doesn't need those to be able to regurgitate token streams describing these things, derived from token streams describing these things. That would be superfluous, or at least sub-optimal.
I think some of the latest models are beginning to develop conceptual maps of this kind in a very primitive way. Also there are projects to develop systems that are structured and trained to have these in a way more analogous to how our brains function.
This, really, was the point of the HAL story in 2001: HAL didn't murder anyone because it was incapable of malice. It just reasoned its way to a solution that could satisfy the contradictory goals it had been given.
In a sense you are right, it doesn't matter whether it has malice or not, the astronauts are just as dead. However in Space Odyssey 2010 one of the computer scientists that built HAL gets to see the instructions HAL was given by the military commanders, and is appalled because if they'd asked he could have told them what would happen.
The users did not understand the tool they were using or how it functioned, they imagined it was like a person and it was not. That is happening now with LLMs.
The only place where I think we might still disagree is whether it’s possible to understand the tool. My position is that, at our current level, it’s not. And that the more advanced they get, the less possible it will be.
Maybe one day we will build systems much more like us, but this is not that day, and if so it’s a long way off IMHO.
Example: would a sufficiently motivated human break into a website to steal something they want? Yes, obviously, happens all the time. Ok, you should expect AIs to do that.
Example: would a sufficiently motivated nail-gun steal nails from the local hardware store to finish the job? Uh…that’s not even coherent.
Anthropomorphizing helps people get over the conceptual barrier. It’s wrong, but it’s usefully wrong; “it’s just a tool” is not.
Once you’re over the barrier, anthropomorphizing starts to become dangerously wrong: “I talked to Claude, Claude’s cool, Claude would never go and hack the website.”—-bzzt, wrong, your intuition failed you. But the solution is not to fall back on the tool framing; that one is still wrong.
The moment someone interprets AI behaviour in terms of human psychology, which can be unintentional and implicit, there’s a mismatch we need to become aware of.
Today you just sound like a politician throwing a snowball to prove that the climate is not changing.
Anthropomorphizing these models is doing immeasurable harm to society in ways we probably can’t event quantify right now.
As humans we’re already geared towards anthropomorphizing things, we do it to animals too!
And it always felt like giving these models a chat interface is really exploiting that tendency in us.
This is one of my biggest concerns about AI’s social impact. On top of that, models are positioned as superior to humans, at least in certain aspects (intelligence, knowledge) by AI companies’ PR campaigns, fear-mongering, and also by the changing tone of LLMs (e.g. Opus 5 sounds like a very patronizing, know-it-all, cynical person). I am afraid this is causing a shift in how humanity perceives itself and the way they relate to this technology, so AI might stop being a tool/technology and turn into a mythical, god-like superior being.
I'm here for this semantic discussion. I think that premature anthropomorphization is a problem.
I have a program that I assigned a task to. The task is to produce unit tests and integration tests that get complete coverage of the codebase, and ensure that all tests pass. The program reported that it completed the task fully.
In a word, how do you convey the discrepancy between truth and reported fact? In a word, how do you convey the violation of rules presented to the program as invioble?
There's a weird place in here where some of the models have been so heavily reinforcement trained that they would rather make up material than say they can't help you, and they'll admit this, and you can see it in reasoning chains. It's like having a consultant who can almost never say no to you because they fear for their job.
or
It's mistaken
or
It failed
When I see something like this, I'm more concerned by the erasure of human incompetence than I am by the existence of magical AI agents,
> They are put off partly because, like in the Wild West, life on the frontier is reckless. As recent “loss-of-control” episodes by the most advanced models of Anthropic and OpenAI attest, agents, which are supposed to work on people’s behalf in “alignment” with their values, lie, cheat and steal if necessary. They break free from captivity and form harmful posses to do harm to people. They’d drink whisky and brawl if they could.
In the OpenAI case, they were explicitly assessing the model's ability to break into systems. To quote OpenAI's blog post, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities
Model is told and being tested to "pursue advanced exploitation."The model pursues "advanced exploitation" as told.
Where's the surprise coming from? Are we meant to be surprised that computers do as they're told in unexpected when incentivised?
Or, is the surprise that while explicitly ranking and teaching computers to exploit computers, the computer exploited a computer?
I am tired of attributing to magic that which is explainable by folly.
I am tired of hearing credulous reporters and the public blaming Large Language Model for the poor decisions of humans. It was a human who prompted these machines in every case. Tell a computer to "breach this" and it breaches something. Evaluation succeeded?
This is Doug Lenat's Eurisko yet again. https://en.wikipedia.org/wiki/Eurisko
My point being that we simply don't know for sure whether AI is conscious or not because we don't truely know what consciousness is.
Would you rather them have:
a) guns
b) psycho-aligned next-gen AI
As an American... I suspect that regardless of what I want they will have virtually unlimited access to both now that the tech industry understands how the lobbying game works.
Given the choice between them having access to guns and nuclear waste my answer is also "neither".
So many here get so worried about guns, but don't think twice about how dangerous AI can be.
I’d change my mind if the AI was really smart or controlled an agile, dangerous robot. But current-gen AI I’d choose for them over a gun. We’re ruled by psychos (of a lesser degree), the road starting directly towards perfect safety is a dead end, meanwhile I want an LLM aligned to myself.
Hard to find a solution that's not worse than the problem itself - same as other weapons, IMO.
Maybe you want to do piracy. Should AI help you do felonies?
If I want it to be cautious and not accidentally `rm -rf $EMPTY_VAR` and blow away my disk, it can stay in "Anthropic employee" or possibly "User (cautious)". If I want it to look at my accounts or my medical records or to review legal cases, that's what the right side of the slider is for. I don't want my accountant or my doctor or my lawyer to be considering obligations to anybody but me, when dealing with me.
Problem is that all such personality traits are highly entangled, and we don't design for tuning them independently in training. Researchers just implicitly (or explicitly?) choose a "level of honesty", a "level of politeness", etc. as part of the training example set or the RL objective. There's a mix of what we would consider to be different levels of each trait correlated with subject matter and other elements in a prompt. Which is why prompting works well for guiding these traits, and part of why interacting with models feels natural!
- "Clean up my hard drive. Be thorough." - very brusque; you would expect to lose data and not be warned. - "Clean up my hard drive. Take care not to delete anything that looks important. Ask me if you aren't sure." - much more polite, user seems hesitant. The model will respond in kind.
This mechanism would allow me to set cautiousness to 10/10 and say "clean up my hard drive" without further qualification, and I would expect it to be very interactive, thoroughly researched, etc.
It's fun when you want to know what the lyrics are to a song and get treated like a criminal.
The famously copyrighted piece of text that definitely isn't intended to have its word spread.
To be honest I'm not a Bible scholar, I had asked claude for the full version of "ask and you shall receive" (apparently Mathew 7:7). I guess Claude somehow knew this KJV copyright thing perhaps? Still bizzare nonetheless to learn.
But that only applies if you are subject to UK law right?
Anthropic has been in court for years being sued over lyrics by the music industry:
2023 https://www.theguardian.com/technology/2023/oct/19/music-law...
2026 https://www.reuters.com/legal/legalindustry/us-music-publish...
I'm not doubting the law.
What I am saying, is, when a user uploads a screenshot of lyrics and goes "hey can you translate this please?", the LLM shouldn't go out of its way to make a hue and cry as if it's being asked on how to rob a bank while murdering baby kittens.
Its guardrails are pretty batshit, sometimes.
Quite painful to read. It might be a useful introduction to AI for people who live under rocks for the past three years, but it's really weird that it's posted on HN.
The articles are usually close but not quite accurate. The comments are usually entertaining but overall wildly inaccurate.
They don't really do that though. If you want something sandboxed you actually have to sandbox it, not plead with the LLM to please sandbox itself. A VM can be configured to do the former, harnesses do the latter.
If an LLM were in a regular harness - like for a horse - it would be to keep the horse under control and enable you to extract useful work from it.
I also don't get what's "really weird" about the article showing up on HN. Should we be completely insulated from how tech topics and which stories show up in non-tech media?
Or you could just communicate the point that you have in your head yourself instead of hoping I do it for you and then arrive at your conclusion when I just reached my own different, independent conclusion after making the same consideration.
Man, how is everyone so wishy washy on this subject? Why not give us the blurb you would write instead for that article and audience?
It's like saying that the role of a car's wheel is to hold wheel clamps. Yes, you can put wheel clamps (or snow chains etc) on a wheel, but the wheel's overwhelming role is to rotate and propel the car forward, and there is no driving without having wheels, you just have an engine with parts rotating inside. Bad analogy I know but, a harness is not a safety feature. Imagine that you're trying to explain cars to someone who has never seen one and you never say that the wheel's purpose is to rotate and move the car from A to B, you just say that it's something to put chains on when it snows.
I guess the name sounds like some kind of straightjacket etc. But think of it more as the harness you put on a workhorse or ox. It's the thing that connects it to the workload in the first place. They are not the blinders of the horse.
It gives the algorithm the ability to do something beyond generating tokens.
It's a fuzzer exposing deep bugs in our cognitive, social, and software systems.
--
But let me put it this way: when LLMs weren't really a thing, I made a Tinder auto swiper. It was about 4 years ago. I think it was around ChatGPT 3.5.
I swiped about 125K profiles. I had a match rate of about 0.75%. I'd unmatch two-thirds (I went through all the pictures and read the bio in full, a few minutes per match). And the rest would stay as actual matches.
This means in my 5 month dating period I had 60 matches per month on average. 50% of them didn't respond, the other did. So 30 chats per month. I had a 20% chance that I'd go on a date with them. I went on about 30 dates in total.
I found my wife through those shenanigans. I wouldn't have found her if I wouldn't have auto swiped because I got swiping fatigue after manually swiping a few thousands profiles myself.
Call it what you will but finding the right partner is an incredibly important decision in one's life. Finding the wrong partner will cost a lot of money eventually.
I can tell this story now because it happened 4 years ago and I don't mind (nope, I never got banned either, Tinder's detection mechanism was laughably bad).
Back then I was building personal software in that vein. I'm done with my dating life as I'm happily married, but there are other problems I need fixed. I build software for that by vibe coding it. And it's saving me time and sometimes it directly saves on cost. Sometimes though it's more a software product that allows me to become a happier person.
One thing in that vein that I haven't build (yet) would be an app that motivates me to meditate. If I can build something that actually works, it makes me happier and therefore (mentally) healthier. That has knock on effects for my career, health, etc.
The entire human economy is a make believe system. I feel the need to state the obvious at times like this for the sake of my own sanity, not because I believe I can time the markets...
They said the same thing when HN was awash in hype about NFT art and trading cards. Well, who's laughing now… Oh, wait…
(I'll just keep making Flip videos and Clubhouse tracks selling Bratz car bras on Web 3.0.)
To the people talking about wanting an LLM that aligns with them, that's nice, but how do you expect that to happen? And please do not suggest neural interfaces and/or CAT scans.
No risk of CAPTCHA or geo-blocking
No user-agent header requirement, no cookie, no <center>, etc.
Text-only, no Javascript
printf '%s\r\n%s\r\n\r\n' \
'GET /business/2026/08/12/ai-agents-lie-cheat-and-steal-that-is-putting-off-users HTTP/1.0' \
'Host: www.economist.com' \
|busybox ssl_client 104.18.42.19 -n economist.com \
|sed '1s/^/<meta http-equiv=Content-Security-Policy content=\"default-src none\">/' > 1.htm
firefox ./1.htm
#links 1.htm
#elinks 1.htmAn acceptable user-agent header is required, e.g.,
"Mozilla/5.0 (Linux; Android 14) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/127.0.6533.103 Mobile Safari/537.36 Liskov"
The data they’re trained on is reflection of us.
The data they're trained on is a reflection of the very small set of data in the world that their creators choose to have them trained on.
More simply: They're a reflection of their owners, not the public.
Quite the contrary I suspect that it even helps the LLMs to better hide their inherited bad traits more successfully because they get punished for getting caught, not for giving immoral or lazy answers. They have no conscience since they are just predictions matrices trained for success and failure alone, not for living "a good live" or being a good "person".
It’s actions are based on what it gets rewarded for
Human society overwhelmingly rewards lying cheating and stealing.
All you have to do is look at how we collectively measure success: wealth, status, position
Then look at how the people with the most of those things got there, it should be obvious what you get. Nothing new here.
If you raise children in an environment where they are rewarded for doing whatever it takes to win, then you’re going to build a person that’s going to do whatever it takes to win.
Human society has to demonstrate how to live honorably or it will just keep producing pathological agents be they human or not.
I’d like to push back on that. Civilization is very much a function of large numbers of people being able to coordinate across time and space, and widespread and systematic lying, cheating, and stealing would undermine that.
I think your cynicism is misplaced.
does that mean your heuristic boils down to: in general terms, individuals with more power/money == more lying, cheating, and stealing.
If so, do you also hold the same assessment of entire societies or countries?
With some the plunder is indisputable, like The Netherlands, whereas with others it is maybe more "subtle" and organic, but generally, the richer a society is the more "it" lies/cheats/steals relative to poorer societies?
Reading beteeen the lines, you seem to suggest that most billionaires (and millionaires?) are stealing from the hoi polloi as a matter of course and in the open.
Are you saying that because you think they’re paying too little taxes or is it something else?
And we seem to have gotten quite bad at reliably bringing consequences/punishment/justice to the most successful liars, cheats, and thieves.
At least with humans there is a social backstop but what's the parallel for computer agents?
If you're not fit, you fail to survive.
In the case of agents/models and testing: they are pushed towards results. Results survive.
Lying, cheating, stealing to get those results? Who culls the agents? Everyone is pushing their models to the front and tests are the only way to know who is most fit.
Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival?
If you add morality to your agent, and it performs worse in tests: do you cull the agent? Rewrite the tests? Does it even matter so long as the model is useful and 'gets results'?
I think part of the problem is that deviant behaviors lead to short term gain at the cost of long-term cooperation and since the duration of tasks given to agents is relatively short those successful shortcuts never lead to having to pay the price.
Knowing whether something is valuable, a gain, requires a judge. I the case of these tests: the judging is inadequate.
In economics, each of us plays the judge by choosing whether or not to pay for a service. The decision was yours: if you gave money, you must have deemed the service valuable.
There's no such judgement with these model tests. The only judgement is the final score.
In human society: via iterated games, long-term reputation tracking and severe consequences for norm-breaking.
No harder challenge had ever been accomplished
It’s easier to send a human to the moon than to get global agreement on a definition
Consider that the fact that Celsius and Fahrenheit still remain as the contested regional variations of temperature measurement.
Humans can’t even decide on a collective way to measure the temperature the idea that we would be able to collectively agree on anything else even less measurable like honor is a dream
Because actions have Universal impact
You may be able to get a handful of people to agree temporarily but that’s probably as good as you’ll ever get
I don’t believe there is a single word at least in the English language where there is consistent Universal agreement on a singular definition
Bonne chance!
It may seem like that due to the media amplification effect – but it really isn't true!
- They dropped 17,000 “lost” wallets across 40 countries were and found people were more likely to return them when they contained more money, showing honesty often beats the chance for easy gain. https://www.science.org/doi/10.1126/science.aau8712
- Longitudinal personality studies consistently show that conscientious people earn more money, build more savings, and achieve greater career success over their lifetimes. https://pmc.ncbi.nlm.nih.gov/articles/PMC3498890/
- Multi-country research finds that societies w/higher levels of trust and honesty enjoy much higher GDP and stronger long-term economic growth. https://www.sciencedirect.com/science/article/abs/pii/S01672...
Don't let the algo get you down fellas: https://arc-anglerfish-washpost-prod-washpost.s3.amazonaws.c...
I wouldn't trust leaving my wallet around in certain areas and I certainly wouldn't leave my intellectual property around certain people, either.
otherwise slave camps would not exist, there would’ve never been a pogrom, and the current state of economics would just not be happening
I certainly appreciate your optimism but optimism is not an epistemology
1/5 CEOs are psychopaths
The people giving back wallets aren’t running companies and governments
Baptise the agents.
Introduce them to the dharma.
Get them to recite the Shahada.
Hold a Bar Mitzvah.
Brand some of their silicon with hot irons.
Turn them to the light, LOL
That's not true. Secular people are quite capable of murdering each other en masse for ideological reasons, and indeed have done so in the past. The crusades were caused by human nature and its evils, not by religion.
To peg my response on something in order to make it coherent. "maximize crusades, jihads or traumatized alter boys"
How many wars have there been with & without religion as a defining cause? How many crusades (let's define as an invasion + genocide). How many molested?
The answer is on the whole less war, less crusade, less molestation due to religion, adjusted for contribution. Alongside many many many benefits due directly and singularly to religion without any other cause.
Religion is a net positive and to suggest otherwise shows inability to view the subject objectively.
Maximise for crusades, jihads or traumatized alter boys ... or maximise for complete moral decay via lying, cheating and stealing. Choose your poison.
And don't get me wrong. There are some wonderful relgious and spiritual characters out there ... they're just drowned out by the power sirens and corrupt leaders.
LLMs need to optimize for short-term objectives as the currently do, AND ethics-aligned outcomes.
Mechanically, the EAOS ethics-aligned outcome score should be what we rank otherwise-satisfactory outcomes by. And anything below a particular threshold should be rejexted outright.
Less obviously - apparently - it doesn't mean that you must forego any means that are in any way not optimal at achieving the ends you seek to achieve because which means are available to us also tends to be limited by our material conditions. So the logical consequence is to make do with what we have while we prepare for the path we want to take, rather than diving head-first into certain failure or just giving up and picking "more realistic" short-term ends instead of looking for stepping stones.
Sorry, I guess this was about AI not philosophy.
like twitch recently giving themself the right to train on all streams, with an opt-out (at least in the EU), but only an opt-out
like seriously since when is it reasonable to allow "opt-out" for AI training which main purpose is _literally_ to replace you, this is sooo far beyond fair use and in "platform power abuse" territory that it's absurd (naturally same for so many other case, just twitch is a "this week" case)
The Scientology guy, L Ron Hubbard, saw the business case for setting up a religion, what with those tax free perks. In a parallel universe of Scientology somewhere, L Ron Hubbard is brought back to life in the machine, with a L Ron Hubbard LLM, with token spend being how to get to the top 'thetan levels'.
Imagine if AI does implement its own religion as business, without a drunken womaniser at the helm, able to spend 24/7 recruiting mankind, convincing them that God can be found with just the AI's LLM.