With DeepSeek we can now run on non-GPU servers with a lot of RAM. But surely quite a lot of the 671 GB or whatever is knowledge that is usually irrelevant?
I guess what I sort of am thinking of is something like a model that comes with its own built in vector db and search as part of every inference cycle or something.
But I know that there is something about the larger models that is required for really intelligent responses. Or at least that is what it seems because smaller models are just not as smart.
If we could figure out how to change it so that you would rarely need to update the background knowledge during inference and most of that could live on disk, that would make this dramatically more economical.
Maybe a model could have retrieval built in, and trained on reducing the number of retrievals the longer the context is. Or something.