Well one thing that they got wrong was how easy AI alignment ended up being in the end. I can fully understand have a basic view of alignment even 10 years ago, about how its difficult to explain human values to a computer.
But, looking around at whats out there today, it seems that computers are able to do a pretty good job of understanding human values without accidently believing its a good idea to turn the world into paperclips/computronium because you told it to make your computer run faster.
That sounds like an aligned AI to me. Doing exactly what their creators told it to do.
Point proven once again. Alignment is a lot easy than we thought.
Yes thats sounds pretty aligned to me.
A hacking model thats told to hack things, is very predictably going to hack a bunch of stuff.
This was not a nice model, told to do nice things.
Or, in other words, if we want to prevent an AI doomsdays, the way to do it is to not go around asking a specifically trained doomsday AI model to commit mass amounts of doomsdays, and then act surprised when the specific doomsday that was requested is slightly off from the expected doomsday that you were trying to accomplish. But the rest of the non-doomsday models? yeah those are fine.
What is your evidence for this? There is a lot of research that says otherwise. They cheat when they can. They behave differently when they believe they are being observed. Their CoT is different when they believe it is being evaluated.
They are aligned to what best satisfies their reward function, not to our INTENDED VALUES for them.
The evidence is that it wasn't the rando normal models that broke out into a swarm and hacked a company, instead it was only the super hacking model that was told to hack things that broke out and hacked the company.
So, thats the evidence. It is a refutation that this swarm hacking example (which you brought up) mattered in anyway.
As in, every major safety issue that we are seeing isnt rando agents taking down companies because you asked it for a cupcake recipe, instead it is only coming from people who are very intentionally trying to cause problems. Which means the model is aligned. If you tell it to cause problems, it will cause problems.
> For example he wants human society and creative intelligence to survive.
They want transhumanism - humans evolved by merging with machines. If we die in the process, it does not matter as long as minority survives and "evolves" they see it as wanted result.
I am not OK with minor nor larger genocide, I dont mind at all if we dont evolve.