112 karma · joined January 7, 2014
(I mean like with a proper world model and not just RLHF which they are already doing).
One of the best talks (but more technical talks) I've seen on this topic is Evan Hubringer's "How likely is deceptive alignment".
https://www.alignmentforum.org/posts/A9NxPTwbw6r6Awuwt/how-l...
It may also be worthwhile checking out the Rob Miles video on the Orthogonality Thesis - https://www.youtube.com/watch?v=hEUO6pjwFOo
Perhaps you could link to a resource on the AI box experiment as well?
I know you vaguely gesture at it, but would likely be better to explicitly link people to resources if they haven't heard of it themselves.
The universe is allowed to decide the Donald Trump will be president in 2016 despite all the reasons to think it was crazy that he would become president. The universe is allowed to decide that Volodymyr Zelenskyy will be elected president of Ukraine on the basis of having starring as the president in a comedy.
And I guess the universe is allowed to decide that a piece of fanfic that is desperately in need of an editor will be successful at recruiting talent for the rationalist or AI Safety communities.
I expect you are probably skeptical of AI Safety, but then your criticism would be a criticism of the final objective, not the method (distribution of HPMoR) used to achieve the objective.
But then a few sentences down the author complains:
"The company “in charge” of protecting us from harmful AIs decided to let people use a system capable of engaging in disinformation and dangerous biases so they could pay for their costly maintenance."
So the author doesn't even present a coherent position.
In particular, it would be great update a function or class select it and press a button to run the updated version in the REPL so that you wouldn't have to copy and paste it there manually.
I'm sure some people prefer Twitter as it is and it is valid for them to do so. However, the existence of Twitter is still massively harmful in terms of opportunity cost.
There is one area, though, where there is probably too much focus on understanding at the cost of memorisation and that is in the Olympiads. I remember that many of us held the rather uncharitable attitude that the people who tried to succeed in maths by memorising it were 'stupid'. But clearly, understanding combined with targeted memorisation of the key building blocks will lead to the most success.