> Probably better for them to misunderstand and run 7b than run 671b. [...] if you don't like how things are done on Ollama, you can run your own object registry, like HF does.
I use Ollama for prototyping and then move what I can to a vLLM set up
I think the main point that turned me off was how they have their custom way of storing weights/metadata on disk, which makes it too complicated to share models between applications, I much prefer to be able to use the same weights across all applications I use, as some of them end up being like 50GB.
I ended up using llama.cpp directly (since I am a developer) for prototyping and recommending LM Studio for people who want to run local models but aren't developers.
But again, if you find Ollama useful, I don't think there is any reasons for dropping it immediately.
ollama has their own way of releasing their models.
when you download r1 you get 7b.
this is due to not everyone is able to run 671b.
if its missleading then more likely due to user not reading.
I'm not super convinced by their argument to blame users for not reading, but after all it is their project so.DS has more or less been ignored for a very long time before this.
You think Ollama is purposefully using misleading naming because they're mad about DeepSeek? What benefit would there be for Ollama to be misleading in this way?
There is no benefit I think.
If you want the actual "lighter version" of the model the usual way, i.e. third-party quants, there's a bunch of "dynamic quants" of the bona fide (non-distilled) R1 here: https://unsloth.ai/blog/deepseekr1-dynamic. The smallest of them is just able to barely run on a beefy desktop, at less than 1 token per second.
Not that I don't believe you (I do, and I think I've seen them correct this before too), but you happen to have specific examples when this happened?
More recently deepseek 2 had a space after the assistant turn, causing issues with output quality and language https://www.reddit.com/r/LocalLLaMA/comments/1dko6rp/if_your...
How do you get accurate information on the template structure?