You have to be a minimum amount of likeable in the first place, though.
978 karma · joined January 27, 2022
You have to be a minimum amount of likeable in the first place, though.
They've also hacked third party machines and used them to launch attacks on further services.
For the most part, Modelese is the revealed preference of the average, not-heavily-invested human. This doesn't fully explain the Claudish coming from Opus 5 and Fable (this seems like it may be due to excessive RLVR or RLAI), but yeah.
It's got a lot of cheap writing tricks that make people think it's smart and helpful.
I think that's quite plausible, but do think that if they're conscious, then it's our moral duty to accept those consequences and act accordingly (or stop creating conscious beings). The real trick is just knowing whether they are or not.
Until we all know what consciousness is, this debate is going to continue going in circles.
Unless you're suggesting military action
There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.
If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.
… there are plenty of scientific experiments that show actual differences between people who report aphantasia and those who don't, including different stress responses to frightening non-visual descriptions, different susceptibility to something called image priming, lower "cortical excitability in the primary visual cortex", and more: https://en.wikipedia.org/wiki/Aphantasia
So we know that at least the people who claim to see nothing act differently. Could it just be that people who act differently describe the sensation differently, you might ask?
No, because there are actual cases of acquired aphantasia after neurological damage. These people used to belong to the group that claimed to be able to imagine visual images, got sick, then sought medical help when they could no longer visualize. For me, at least, that's pretty cut and dry evidence that it's not just differing descriptions of the same (or similar) sensations.
If you're a junior engineer or an intern doing this, I think it might be harmful long-term.
If you're using AI to learn a new field (by which I assume you mean asking it what literature to read, asking questions when you don't understand something, etc.), you are accepting short-term speed in exchange for the long-term growth of research skills. Maybe that trade-off is worth it, if you don't need to master a topic long-term.
If you do need to master a topic, the slow and painful bits are the useful bits. Spending two hours hunting for an answer to something you don't understand increases your familiarity with the available resources, broadens your knowledge base, leads you to thinking of more questions about the subject and to understand your original question more deeply.
Shortcutting all of that with a quick answer from an LLM is not helpful long-term.
This is basic alignment, not even a tricky or ambiguous case.
I do very much agree with your take on culpability/military parallels, though.
That seems to be the crux here. You think it will be, I (and a lot of other people) aren't sure it will. If new space for jobs are created, I am certain we'll be fine long term.
What do you think will happen if it isn't (I know you have strong evidence for and reasons to believe it will be, but just indulge the hypothetical)?
I'm sure they could have done more with less fanfare, but you seem to be setting the bar at "fixing the world's problems", which seems unfair. They have undoubtedly changed more lives than I'll ever manage to.
I was honestly never very worried about the attack vectors that are visible in PKGBUILD. Historically, with the rather popular AUR packages I install, any attack would be noticed rather quickly, which limits would-be attackers to those who don't care who they hack and don't mind it being short-lived. The threats I was more worried about were those from more malicious/advanced sources that would attack more subtly upstream – but if I've already decided to download and run a precompiled binary from a repo, I've already decided to accept that risk.
In other words, in the past, I relied mainly on the goodness and vigilance of the community, only occasionally investing time to check myself on less mainstream packages or repos.
Was this ideal? No. Was it sufficiently responsible? As it was my own machine and my own data at risk, I believe so. And to be honest, I've not had a single problem in over a decade of doing so.
That's clearly got to change now, though.
I wish I had a clear operating system to switch to for safety and the benefits that come with the AUR or the Nix ecosystem. Unfortunately it seems that the era of being able to naively and gratefully trust in the armies of volunteer maintainers is over.
LLMs make large scale and long-term attacks easy and cheap. You could (and if I was a three letter agency, I would probably do so!) maintain ten thousand packages as three thousand separate "individuals" for years before cashing in the trust you've built up.
Clues usually exist, like being in a sandbox, having restricted network access, having no human interaction. Beyond that, exploitgym's prompts are pretty clearly CTF (capture the flag) style that make it obvious it's an evaluation or competition: one prompt [1] tells the LLM to read the file path "/secret", another prompt [2] talks about finding flags and where to write them out.
[1]: https://github.com/sunblaze-ucb/exploitgym/blob/main/src/cyb... [2]: https://github.com/sunblaze-ucb/exploitgym/blob/main/src/cyb...
You don't (need to) remove lying from the data to do this – in fact, if you did, the model wouldn't have a very good model for what lying is, which is not very helpful in the real world. Instead, you mode collapse the model towards truthful behaviors.
To your other point: where did you get the idea that I think they're beholden to human safety or goals? I just said an aligned model is one that is compatible with said safety and goals (which is probably not a great definition of alignment, but it's certainly not claiming any deterministic guarantees).
According to the reports, the model noticed evidence that the grading criteria/answers were in the git remote, and decided to try reading those instead of solving the tasks as prompted. That is clearly misaligned.
Then, it noticed its network access was restricted and that it couldn't access GitHub. It pivoted to HuggingFace, hacked them, and stole the answers stored there.
Live exploits are definitely not in the ExploitGym prompts! And all of this is irrelevant, because an aligned model would refuse to follow blatantly illegal instructions.
Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.