What's the huge difference between the two pelicans riding bicycles? Was one running locally the small version vs the pretty good one running the bigger one thru the API?
Thanks, Morgan
Mistral's API defaults to `magistral-medium-2506` right now, which is running with full precision, no quantization.
It literally only makes everything worse and more convoluted with zero benefits.
It’s usually either because the context size is set very low by default or they didn’t realize that they weren’t running the full model (ollama uses the distilled version in place of the full version but names it after the full version).
There’s also been some controversy over not giving proper credit to llama.cpp which ollama is/was a wrapper around.
I've never used ollama, but perhaps you mean quantized and not distilled? Or do they actually use distilled versions?
Just use llama.cpp directly.
but then someone found that, at least for distilled models,
> correct traces do not necessarily imply that the model outputs the correct final solution. Similarly, we find a low correlation between correct final solutions and intermediate trace correctness
https://arxiv.org/pdf/2505.13792
ie. the conclusion doesn't necessarily follow from the reasoning. So is there still value in seeing the reasoning? There may be useful information in the reasoning, but I'm not sure it can be interpreted by humans as a typical human chain of reasoning, maybe it should be interpreted more as a loud multi-party discussion on the relevant subject which may have informed the conclusion but not necessarily lead to it.
OTOH, considering the effects of automation fatigue vs human oversight, I guess it's unlikely anyone will ever look at the reasoning in practice, except to summarily verify that it's there and tick the boxes on some form.