This is insanely cool, what the hell.
How co-designed are these optimizations with the model itself? I'd imagine you can't just stick post-training adapters onto existing architectures for these things, or am I wrong?
I really want to explore the inference space, but it seems like many of the inference optimizations are coming from model-hardware codesign. I don't seem to recall many generic "inference engine" optimizations since prefill/decode disagg a year ago.
This matters for me since I want to break in but the bar seems to be understanding the actual theory of the training process now too given the codesign happening, and I'm not the richest guy on the block lol