Also, what's the best way to benchmark a model to compare it with others? Are there any tools to use off-the-shelf to do that?
Also, what's the best way to benchmark a model to compare it with others? Are there any tools to use off-the-shelf to do that?
You would have to confirm with someone deeper in the ecosystem, but I think you should be able to run this new model as is against a llamafile?
My recent work optimizing CPU evaluation https://justine.lol/matmul/ may have come at just the right time. Mixtral 8x7b always worked best at Q5_K_M and higher, which is 31GB. So unless you've got 4x GeForce RTX 4090's in your computer, CPU inference is going to be the best chance you've got at running 8x22b at top fidelity.
Really easy to search huggingface for new models to test directly in the app.
I’m sure they are already working on it.
That person is a hero, super bummed!
https://api.together.xyz/playground/language/mistralai/Mixtr...