What's the huge difference between the two pelicans riding bicycles? Was one running locally the small version vs the pretty good one running the bigger one thru the API?
Thanks, Morgan
What's the huge difference between the two pelicans riding bicycles? Was one running locally the small version vs the pretty good one running the bigger one thru the API?
Thanks, Morgan
Mistral's API defaults to `magistral-medium-2506` right now, which is running with full precision, no quantization.
It literally only makes everything worse and more convoluted with zero benefits.
It’s usually either because the context size is set very low by default or they didn’t realize that they weren’t running the full model (ollama uses the distilled version in place of the full version but names it after the full version).
There’s also been some controversy over not giving proper credit to llama.cpp which ollama is/was a wrapper around.
I've never used ollama, but perhaps you mean quantized and not distilled? Or do they actually use distilled versions?
Just use llama.cpp directly.