eh? doesn't the distilled+quantized version of the model fit on a high-end consumer grade gpu?
eh? doesn't the distilled+quantized version of the model fit on a high-end consumer grade gpu?
Sure, you could say that only running the 600+b model is running "the real thing"...
You can pretend it's R1, and if it works for your purpose that's fine, but it won't perform anywhere near the same as the real model, and any tests performed on it are not representative of the real model.
For everyone who doesn’t build LLMs themselves, “running a Llama:7B model fined-tuned on DeepSeek.” _is_ using Deepseek mostly on account of all the tools and files being named DeepSeek and the tutorials that are aimed as casual users all are titled with equivalents of “How to use DeepSeek locally”
Most people confuse mass and weight, that does not mean weight and mass are the same thing.