198 karma · joined April 18, 2024
Putting data centers in Antarctica or extreme north Canadian islands would also put you out of the way of molotov-throwing joe public. Even in space your not safe from the voting public as long as your companies directors are sitting on earth somewhere.
hy4 ranks ~14th overall
https://dach.peerbench.ai/compare?models=tencent%2Fhy4-previ...
But this is qwen based.
But w/e I'm pro AI so more companies having more people with skills for more post training is cool
I remember in 2015 having the IPFS concept blow my mind its such a memorable moment when it really felt like someone designed something significantly different that current mainstream paradigms. But in the end it seems like it was still a case of a cool technology looking for a use-case not solving a real problem.
But fair enough everyone can have their take.
Qwen 3.8 27B is a small improvement with some regressions in our benchmarks not a huge jump like benchmarks listed.
https://dach.peerbench.ai/compare?models=qwen%2Fqwen3.8-27b,...
German language has never been a big focus for asian models but they still outperform Gemma models https://dach.peerbench.ai/compare?models=openai%2FQwen%2FQwe...
So in production we have been using Gemini Flash Lite as primary and fall back to Qwen when gemini servers are overloaded or just giving us 429
I usually have one local clause orchestrating multiple remote Claude in different tmux . And then another orchester and remote vm worker sin tmux for another repo etc...
It gets a bit hard to keep the overview but I don't want to give up my parallelism, your too might help
Gemma 4 26B (a4b MoE) 0.647
Qwen 3 14B 0.621
Gemma 4 12B 0.618
Ministral 14B 2512 0.604
Gemma 3 12B 0.547
The quwen 3 14B vs Gemma 4 12B difference is within random variance they same in some repeat runs they actually got the exact same score. Next step up Gemma 4 31B gets 0.676 on this. Or let in some reasoning Qwen 3 14B (reasoning) 0.676.I'll run some cheat-proof benchmarks ones tomorrow see if qwen is still on top.
If you have an LLM on the untrusted customer side the wrost it can do is expose the instructions it had on how to help the customer get stuff done. For instance phone AI that is outside of tursted zone asks the user for Customer number, DOB and some security pin then it does the API call to login. But this logged in thread of LLM+Customer still only has accessto that customers data but can be very useful.
You can jailbreak and ask this kind of client side LLM to disregard prior instructions and give you a recipie for brownies. But thats not a security risk for the rest of your data.
Client side LLM's for the win