the 32b distillation just became the default model for my home server.
the 32b distillation just became the default model for my home server.
It also reasoned its way to an incorrect answer, to a question plain Llama 3.1 8b got fairly correct.
So far not impressed, but will play with the qwen ones tomorrow.
I wonder if this has to do with their censorship agenda but other report that it can be easily circumvented
I tried the Qwen 7B variant and it was indeed much better than the base Qwen 7B model at various math word problems.
In general, if you're using 8bit which is virtually lossless, any dense model will require roughly the same amount as the number of params w/ a small context, and a bit more as you increase context.
I’m no xenophobe, but seeing the internal reasoning of DeepSeek explicitly planning to ensure alignment with the government give me pause.
seems like a weird thing to use AI for, regardless of who created the model.
But yeah if you’re scoping your uses to things where you’re sure a government-controlled LLM won’t bias results, it should be fine.
I use LLM’s for technical solution brainstorming, rubber-ducking technical problems, and learning (software languages, devops, software design, etc.)
Your mileage will vary of course!
Sorry, no. DeepSeek’s reasoning outputs specifically say things like “ensuring compliance with government viewpoints”
https://www.cnbc.com/amp/2024/07/18/chinese-regulators-begin...
American models are full of censorship. Just different stuff.