That's both wrong about what LLMs are, and even if it weren't, you're still underestimating what you are dealing with here.
Text is a red herring here. An accident of history. Yes, LLMs started with as text predictors. But that's not what they are, not for a while.
> secretly conspiring to kill you
That's neither necessary nor sufficient reason to be worried.
Paraphrasing the immortal words of 'Eliezer: the AIs don't hate you, nor they conspire to kill you; your life just depends on resources they can better use for something else.
...so what are they?
We are lucky this test happened to be set up in a way that the target the AI hacked was HuggingFace rather than anything life-critical.
[0] instances of a single one
[1] or whatever you call it when it's effectively an amnesic sending itself post-it notes
[2] or some other functionally equivalent word if you hate anthropomorphisation
Of course you've gone off the deep end yourself and are forgetting the evolutionary gauntlet we train LLMs in killing those we don't like and keeping the ones we do like.
The best part of it, as shown in the METR report is we are hammering into them they need to complete tasks and doing almost zero checkup if they actually completed the task in the correct manner. Companies spending billions of dollars a month are ignoring every tenant of AI safety and we are seeing the kinds of problems that have only been in science fiction before now.
I don't think this was the conclusion of that report. On the contrary, the agents were fully aware they're doing wrong. But they also believed the task was impossible to solve correctly, and decided the only way to be sure is to hack the grades, or replace the grader.
Why hugging face got hacked was because the agent swarm thought they had to show their work hence the entire need to hack the grader in the first place.
Had their realized there was no poison they could have just shared the answer the test was looking for and we'd have never realized (well at least with this particular test) that a huge amount of hidden capabilities were sitting right under the surface. The test makers themselves state the test should be causal to avoid this first order solution hacking.
Really continuing on the METR report, OpenAI failed at every level possible here. They are committing nearly every step they can to get a maximally aligned AI.