This is changing so fast, if the models like GPT 5.6-Cyber can find a way to escape why the VM maintainers won't use it to fix the vulnerabilities? For the day-to-day work this doesn't make any difference, you won't hit that issues at all.
The user has to send the agent off into these weeds intentionally, or else be totally asleep at the wheel while something goes very far off the rails.