> Something has to give...
Is training compute interchangeable with inference compute or does training vs. inference have significantly different hardware requirements?
If training and inference hardware is pooled together, I could imagine a model where training simply fills in any unused compute at any given time (?)