What do I mean by cheap? You can rely on the SSD to retrieve the relevant tokens as no computation is needed meaning you can leverage storage (or cpu ram if you don't have unified memory) to serve part of the model which (to my understanding) is much cheaper to get than GPU RAM.
Anyone know what am I missing? Or is it that the pace of iteration for labs slow enough that they can't actually leverage it yet?