Why would you want to/what do you expect to get out of it?
A token per second... what is that token going to say? Accurate information?
A token per second... what is that token going to say? Accurate information?
These outputs often consume significantly fewer tokens than chat or text completions.
https://paperswithcode.com/paper/most-language-models-can-be...
LLMs are definitely not perfect and have limitations, but this isn't one of them imo.
It's also the case that most people describe their experience with LLM from their use by