How does this analogy work when in 1980 you knew that nukes and people using them were a real thing, but we've never seen a frontier model "hack into core inference infrastructure to upload weights to other GPUs to survive"?
We have seen models hack into core ML infrastructure. I’m unsure why the extra self-perpetuating shape seems this magical bridge too far, rather than simply a prompt away?
It's absolutely relevant if we're discussing SOTA models with huge computing requirements on large clusters, that are plausibly tightly coupled to the physical datacenter architecture they were developed for to optimize performance.
This sounds like magical thinking. The right couple of AWS or DO keys get you an environment entirely capable of hosting GLM 5.3. I'm sure Fable and Sol run _best_ in their own special aquariums, but "optimized for a particular cluster" != "can only run on that cluster". There's no reason to think you can't run them if you have the weights on your local university's cluster.