> How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do. I haven't given this much thought, so maybe I'm missing something, but I don't see how this follows. One possible solution: to avoid overfitting, can it not just make a copy of itself, modify the copy, and empirically check if the model performs better? That's essentially what humans are currently doing when designing AIs.
Regarding the X-risk vs D-risk: I think how one weighs these risks partially depends on what one thinks the capabilities of the models are. Call me a boot-licker, but if the models get smart enough to explain, in detail, to any psychopath, how to construct a bomb or synthesize a deadly virus, I don't think benefits society to distribute them widely. Therefore, to argue for widespread distribution you have to argue that either (i) the models aren't that capable or (ii) the guardrails are robust enough to prevent them from being used in catastrophic ways by bad actors. I think we may be rapidly approaching a time where neither of these hold. Having said that, I certainly agree that the D-risk is also real.