link for curious: https://generative.ink/meta/block-multiverse/
link for curious: https://generative.ink/meta/block-multiverse/
This way you could combine e.g. a 7B parameter model with a 70B parameter model, and get the quality of the larger model while most of the time you are only running the small model.
Edit: You could also store the full probabilities for each token, and then the classifier could detect if it had gone down a bad path, and then unwind the tokens and pick a different path.
Better yet: have many smaller models that can when confused call upon larger models and the larger models could then pick the most appropriate smaller 'expert' model or, alternatively themselves escalate. Sort of a supervisor tree for language models.
https://huggingface.co/blog/vivien/optimal-lossy-variant-of-...