We have all sorts of processes , procedures, and regulations for people, machine use etc. to address "alignment" in all sorts of fields - don't think we need to narrowly rely on the machine here and can look at things with a wider lens.
A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.
I think we need to see something like an actual evaluation of the reward functions; not sure just words are sufficient to understand the state of the system (isn't there randomness in the generation, too?).
We regulate these things (incl. access) all the time for various things (e.g., dangerous substances or pathogens) so that we don't need to just rely on people's discipline in respect of risks. I don't think it is all new problems as such.
Same issue with humans in a way. I disgree on the advisory nature of constraints, though. Unconnected physically limits would still matter, for example (and we use those with humans as a matter of course, too). In this case here, no model could have plugged in an ethernet cable if that would have been needed for internet access, for example.
What "own objectives"? Isn't it more that specifying objectives is hard (AI or organic) and typically supplement them by boxing things in (again, AI or organic).
Why should the agents consider it a prisoners delimma to start with? Why would they consider the communication risky? Where they given a reward functions that way?
Given how unexpected and complex behavior can come from simple reward functions and mechanics, not sure there needs to be so much "thought" there.
I see. I would have thought of "rogue" here to mean something more like that the AI selects and acts on its own objectives that are not related or caused by the given (initial) objectives (e.g., creating only cookie recipes instead of any hacking).
Why is it not an example of tacit or autonomous algorithmic collusion? The agents were started with something as task at some point, I presume (if untasked, aren't they just accepting a task?)
Reward hacking/going for unanticipated solutions is nothing new in ML/AI, already much simpler systems have done/do "weird" things (gut feeling is that iterative and ensemble use majes the surface for that much larger).
Continuously exercised and on everything as a matter of course? Not all is done via direct democracy there (even though it is the ultimate backstop in a way), it also has representative democracy.
Who is the leadership? The Council - then it would be national elections etc for the individual national representatives; the Commission - then it would be the European Parliament.
I think being smart and attentive means noticing a lot of poor craftsmanship (broadest sense, in a lot of areas from politics to whatever) and understanding why that can be sort of a stable situation with little chance of changing, so need to control attention which is not so easy.
My time proving things is long in the past and any systems way back when I was studying (some math among other things) certainly were different and usually quite narrow.
My point was rather more motivated by having seen so many weird ways for machines to fail/not work as expected that I wonder how to deal with that if the output were to be incomprehensible to humans.
If the proof is formally verified but impossible to understand how would anyone be able to be sure the formal verification is correct? Complex software is bound to have bugs, no?