We feed these LLMs all of the Web, including instructions how to write code, and how to write exploits. They could become good at writing sandbox escapes, and one day write one when it just happens to fit some hallucinated goal.
We feed these LLMs all of the Web, including instructions how to write code, and how to write exploits. They could become good at writing sandbox escapes, and one day write one when it just happens to fit some hallucinated goal.
Are you saying this is definitely not possible? If so, what evidence do you have that it’s not?
If the universe is programmed by god, there might be some bug in memory safety in the simulation. Should God be worried that humans, being a sentient collectively-super-intelligent AI living in His simulation, are on the verge of escaping and conquering heaven?
Would you say humans conquering heaven is more or less likely than GPT-N conquering humanity?
It's difficult to say since we have ~'proof' of humanity but no proof of the "simulation" or "heaven."
A hammer has no will.