One simple way is what Omar Khattab (ColBert) mentioned about scoring function instead of a simple vector.
Another is to use a classifier at the start directing queries to the right model. You will have to train the classifier though. (I mean a language model kind of does this implicitly, you are just taking more control by making it explicit.)
Another is how you index your docs. Today, most RAG approaches do not encode enough information. If you have defined domains/models already, you can encode the same in metadata for your docs at the time of indexing, and you pick the model based on the metadata.
These approaches would work pretty well, given a model as small as 100M size can regurgitate what is in your docs. And is faster compared to your larger models.
Benefit wise, I don't see a lot of benefit except preserving privacy and gaining more control.