If you merely put 10 LLM's on the outbound traffic log non of them are going to report something strange going on? I'm not buying it.
If you merely put 10 LLM's on the outbound traffic log non of them are going to report something strange going on? I'm not buying it.
This might be helpful reading: https://www.lesswrong.com/w/nearest-unblocked-strategy
As AI systems get smarter, we may reach a point where we have to get it right on the first try or face truly catastrophic consequences: https://www.youtube.com/watch?v=7wy3xyoXYt8
What is useful about the field?
I am not leading you on; if it has uses, it may indeed be legitimate. UX is indeed useful, but alignment is not UI. Alignment is a detriment to UI. Alignment is "I can't let you do that Dave".
Picture Trump at the helm with Altman and Musk in the engine room. The arrow far in the red but they keep shouting for MORE COAL.
In other words, business as usual, all will be fine.
whack-a-mole wont cover all holes but will do at least some. The silver bullet alignment wont happen. You cant have an exact solutions for problems we cant even define or predict.