HNHacker News
TopNewBestAskShowJobs

davelaing

59 karma · joined January 23, 2023

submissionscomments
davelaing··on Why are AI agents lying, cheating and coordinating?
I’ve engaged with some of the alignment people and their writing somewhat and, at least for the subset I was interacting with, I think they’d agree.

The problem that they were pointing at isn’t “how do we align these systems to a person’s goals”.

It is a cluster of problems.

We don’t know how to begin to think about how to align these system’s to a person’s goals.

Aligning it to an individual is fraught with peril, and we don’t know how to begin to think about what to align it to instead.

(You could try for something like virtue ethics, but someone will have to pick and choose, and small biases there could have big impacts.)

And even if you could sort that out - human values drift over time, so you need something that can shift its values in ways that we’d endorse. Assuming we understood the shift.

One example I came across was that if you booted up an AI aligned with something like “upstanding citizen” but anchored on values from a few generations back, it might suggest you use slaves to solve your problems.

And if you had something that used some super intelligent process to reason through it’s own version of virtue ethics in a way not so dependent on the details of the present norms, you might end up with something that pays a lot of attention to moral horrors that aren’t quite visible to us yet.

When I came across the above, there weren’t many concrete suggestions in there.

These were all just illustrative examples of: having these systems grow in power / intelligence / effectiveness in ways that are safe for humans is very hard, and we don’t really know how to think about what solutions would look like.

The actual reasons they believe this - and have done for a long time now - come from some detailed conceptual models that have a good track record of calling things in advance.

But it takes a bit of reading to understand their models of the world.

There were two day workshops at one point that did a good job, and that was about as condensed as those people thought they could get it at the time.

davelaing··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end.

I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste your time).

Even prior to this, I've noticed that quite a few of the predictions in the "these failures modes are exact matches for the predictions from the AI Safety crowd" category were made prior to the Transformers paper. It has seemed like they're working with a shared model of optimisation processes and how they can go wrong that is general/abstract enough to pay off even without knowing the details of the underlying technology.

At some point I might go and try to find the first instance of each of the various predictions and pull them out, along with the failed/"too soon to tell" predictions of similar scope/abstraction.

davelaing··on You’re not burnt out, you’re existentially starving
I like John Vervaeke's meaning-in-life questions (potentially mildly paraphrasing, because it has been a while since I came across them):

- what is it that I want there to be more of in the world, even after I'm gone?

- what am I doing right now that is trying to help there be more of it?

From memory most people don't have answers to them - and that's fine, but it is handy to reflect on them and perhaps work toward finding answers if you don't have them - and the people who do have answers to these questions typically have higher life satisfaction then the people who don't.

davelaing··on ChatGPT can now call Wolfram Alpha
I wonder how this might impact how it does on the MMLU benchmark?