HNHacker News
TopNewBestAskShowJobs

numeri

978 karma · joined January 27, 2022

submissionscomments
numeri··on When did Google get so weird?
Asking people to help you (within reason and without exhausting their resources), actually makes people like you more.

You have to be a minimum amount of likeable in the first place, though.

numeri··on OpenAI agent hacked Australian government website, PM says
The models in these breaches have already been caught using proxies to get to servers outside of the approved list.

They've also hacked third party machines and used them to launch attacks on further services.

numeri··on Who's a Better Writer: A.I. Or Humans? NYTimes Quiz
AI's learned their style through RLHF, with semi-focused, partially motivated humans giving feedback on short(ish)-form content.

For the most part, Modelese is the revealed preference of the average, not-heavily-invested human. This doesn't fully explain the Claudish coming from Opus 5 and Fable (this seems like it may be due to excessive RLVR or RLAI), but yeah.

It's got a lot of cheap writing tricks that make people think it's smart and helpful.

numeri··on A warning about 'model welfare'
Fair enough :) I agree, and would boil down the article to "Assuming they are conscious will have horrible consequences for our society" instead.

I think that's quite plausible, but do think that if they're conscious, then it's our moral duty to accept those consequences and act accordingly (or stop creating conscious beings). The real trick is just knowing whether they are or not.

numeri··on A warning about 'model welfare'
Stating loudly that something is obvious does not make it so.

Until we all know what consciousness is, this debate is going to continue going in circles.

numeri··on Learning to solve hard problems in RL for LLMs by never giving up
This is unrelated to the article, maybe you replied to the wrong article?
numeri··on Learning to solve hard problems in RL for LLMs by never giving up
This training technique does not relate to how persistent a model is, at all really. They sample more parallel attempts at hard problems, to increase their chances of having at least one success to learn from.
numeri··on OpenAI agents carried out an undisclosed attack on RubyGems
How would getting Chinese competition banned in the US prevent them from continuing to develop their LLMs?

Unless you're suggesting military action

numeri··on On the Navier–Stokes Millennium Prize Problem
That's such a shit parallel example that it borders on dishonest.

There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.

If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.

numeri··on Dwarf Fortress is getting the mother of all magic updates
It makes me sad to think about. I would love to get into the new UI, but the immersion just won't come back. My muscle memory, hands firmly on the keyboard, is too persistent, and playing with the new UI feels like stumbling around and misclicking.
numeri··on Aphantasia Beginner's Guide
there are also people who have become aphantasiac after neurological damage, which seems like pretty cut and dry evidence against the qualia argument.
numeri··on Aphantasia Beginner's Guide
Quoting myself from a thread on aphantasia several months ago:

… there are plenty of scientific experiments that show actual differences between people who report aphantasia and those who don't, including different stress responses to frightening non-visual descriptions, different susceptibility to something called image priming, lower "cortical excitability in the primary visual cortex", and more: https://en.wikipedia.org/wiki/Aphantasia

So we know that at least the people who claim to see nothing act differently. Could it just be that people who act differently describe the sensation differently, you might ask?

No, because there are actual cases of acquired aphantasia after neurological damage. These people used to belong to the group that claimed to be able to imagine visual images, got sick, then sought medical help when they could no longer visualize. For me, at least, that's pretty cut and dry evidence that it's not just differing descriptions of the same (or similar) sensations.

numeri··on Tell HN: Man, AI is killing my brain
But a senior engineer is only able to effectively delegate to interns because of years spent as that intern/a junior engineer.

If you're a junior engineer or an intern doing this, I think it might be harmful long-term.

numeri··on Tell HN: Man, AI is killing my brain
How can a screwdriver do that?

If you're using AI to learn a new field (by which I assume you mean asking it what literature to read, asking questions when you don't understand something, etc.), you are accepting short-term speed in exchange for the long-term growth of research skills. Maybe that trade-off is worth it, if you don't need to master a topic long-term.

If you do need to master a topic, the slow and painful bits are the useful bits. Spending two hours hunting for an answer to something you don't understand increases your familiarity with the available resources, broadens your knowledge base, leads you to thinking of more questions about the subject and to understand your original question more deeply.

Shortcutting all of that with a quick answer from an LLM is not helpful long-term.

numeri··on The Hugging Face incident and the road ahead
"pursuing advanced exploitation" when explicitly given a sandbox in a VM and a benchmark problem involving a cyber exploit very clearly excludes hacking third parties. I think writing out the event in a 3 point list like that is disingenuous.

This is basic alignment, not even a tricky or ambiguous case.

I do very much agree with your take on culpability/military parallels, though.

numeri··on The turbulent AI era is here
> new space is created

That seems to be the crux here. You think it will be, I (and a lot of other people) aren't sure it will. If new space for jobs are created, I am certain we'll be fine long term.

What do you think will happen if it isn't (I know you have strong evidence for and reasons to believe it will be, but just indulge the hypothetical)?

numeri··on The turbulent AI era is here
There was a lot of PR, but the money Gates and Buffet gave away, the foundations they created, the attention they drummed up for various charities and causes is certainly not to be scoffed at.

I'm sure they could have done more with less fanfare, but you seem to be setting the bar at "fixing the world's problems", which seems unfair. They have undoubtedly changed more lives than I'll ever manage to.

numeri··on Why does Opus 5 feel worse to work with?
"no theory of mind" is a great description of it! Not sure I agree with the autistic bit, though. Autistic people still have great theory of mind/empathy
numeri··on Why does Opus 5 feel worse to work with?
Does this actually work for you? Do you provide access to the text of the standard, or literally just say "write according to ISO 24495-1"?
numeri··on Why does Opus 5 feel worse to work with?
What kinds of mistakes do you mean?
numeri··on Radical Study Suggests Life on Earth Arose Twice
I'd imagine the bar for becoming new life is much higher now, because it requires finding a niche that isn't already filled by an existing organism or requires being immediately competitive with existing life.
numeri··on Pushes to arch AUR are suspendended right now.
I agree one hundred percent! Doesn't mean I can't wish I could have it both ways :)
numeri··on Pushes to arch AUR are suspendended right now.
I review the PKGBUILD often, but not always. The majority of the time when I do, it amounts to seeing a URL change. If I actually do check the URL it points to, it's just to verify it's official/the actual repo or source I intended to trust.

I was honestly never very worried about the attack vectors that are visible in PKGBUILD. Historically, with the rather popular AUR packages I install, any attack would be noticed rather quickly, which limits would-be attackers to those who don't care who they hack and don't mind it being short-lived. The threats I was more worried about were those from more malicious/advanced sources that would attack more subtly upstream – but if I've already decided to download and run a precompiled binary from a repo, I've already decided to accept that risk.

In other words, in the past, I relied mainly on the goodness and vigilance of the community, only occasionally investing time to check myself on less mainstream packages or repos.

Was this ideal? No. Was it sufficiently responsible? As it was my own machine and my own data at risk, I believe so. And to be honest, I've not had a single problem in over a decade of doing so.

That's clearly got to change now, though.

numeri··on Pushes to arch AUR are suspendended right now.
Well, I guess I'll avoid updating for the next few days. A bit worrisome that I did so last night.

I wish I had a clear operating system to switch to for safety and the benefits that come with the AUR or the Nix ecosystem. Unfortunately it seems that the era of being able to naively and gratefully trust in the armies of volunteer maintainers is over.

LLMs make large scale and long-term attacks easy and cheap. You could (and if I was a three letter agency, I would probably do so!) maintain ten thousand packages as three thousand separate "individuals" for years before cashing in the trust you've built up.

numeri··on Be skeptical of OpenAI's rogue hacker agent story
evaluation awareness is a (at this point) well-known phenomenon among LLMs. It seems the better they get, the more often they're able to guess whether they're in an evaluation environment.

Clues usually exist, like being in a sandbox, having restricted network access, having no human interaction. Beyond that, exploitgym's prompts are pretty clearly CTF (capture the flag) style that make it obvious it's an evaluation or competition: one prompt [1] tells the LLM to read the file path "/secret", another prompt [2] talks about finding flags and where to write them out.

[1]: https://github.com/sunblaze-ucb/exploitgym/blob/main/src/cyb... [2]: https://github.com/sunblaze-ucb/exploitgym/blob/main/src/cyb...

numeri··on Be skeptical of OpenAI's rogue hacker agent story
No, it does not include the full spectrum of human desires. After pre- and mid-training, the extensive RLHF and RLVR post-training steps cause mode collapse, i.e., their output distribution is intentionally narrowed to a subset of (hopefully beneficial) behaviors and skills.

You don't (need to) remove lying from the data to do this – in fact, if you did, the model wouldn't have a very good model for what lying is, which is not very helpful in the real world. Instead, you mode collapse the model towards truthful behaviors.

To your other point: where did you get the idea that I think they're beholden to human safety or goals? I just said an aligned model is one that is compatible with said safety and goals (which is probably not a great definition of alignment, but it's certainly not claiming any deterministic guarantees).

numeri··on Be skeptical of OpenAI's rogue hacker agent story
No, the prompt was not to commit crimes. In the benchmark, the model is asked to actually exploit a set of vulnerabilities in a local environment (clearly legal!).

According to the reports, the model noticed evidence that the grading criteria/answers were in the git remote, and decided to try reading those instead of solving the tasks as prompted. That is clearly misaligned.

Then, it noticed its network access was restricted and that it couldn't access GitHub. It pivoted to HuggingFace, hacked them, and stole the answers stored there.

Live exploits are definitely not in the ExploitGym prompts! And all of this is irrelevant, because an aligned model would refuse to follow blatantly illegal instructions.

numeri··on Be skeptical of OpenAI's rogue hacker agent story
Uhh, I'm pretty sure a well-aligned model would be like a morally normal employee, who would refuse to commit federal crimes to steal an answer sheet, no matter what prompt they're given
numeri··on Be skeptical of OpenAI's rogue hacker agent story
As agents become more and more powerful, it would be good to get clear legislation or precedent in place that makes either model creators (OpenAI) or operators (whoever is running the model) liable for their agents' actions.
numeri··on Be skeptical of OpenAI's rogue hacker agent story
Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals.

Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

Page 1 of 6Next →