I really think this is "modified its own code" thing is overblown. The way Sakana's system works is that it has a script, experiment.py, which is seeded with an initial implementation of whatever topic it's supposed to research (e.g. a basic diffusion model). This is the code the LLM sees in its prompt, and the code it can propose modifications to. My understanding is that in one version, they put the time limit in this experiment script. The LLM tried running it, saw the timeout error, and proposed an edit to increase the timeout. That's what any LLM would do if given some code with a timeout and asked to propose edits given an error message.
But the phrase "modified its own code" suggests something more than this. It makes it sound like the AI researcher is modifying the code that defines how it performs research, rather than modifying the experiment script it's been given. I feel like Sakana is playing up the ambiguity of what exactly "its own code" means for publicity, and articles like this are eating it up.