Copyright (c) 2026 maan
So... Either alooshdenny stole the commits, or it's an alias for maan.
818 karma · joined October 30, 2021
Copyright (c) 2026 maan
So... Either alooshdenny stole the commits, or it's an alias for maan.
What matter the most and isn't told by token price is the latency. You expect a voice LLM to respond very quick. If it takes 5s to response to a simple "Hello, what the weather today?", them not much people will use it.
So, the cost of a setup to run Qwen-Flash-Next at +40tks is around $3000. Too much for most people.
With only a RTX 4090, you will reach 30tps (with DDR5...), not +40tks, and it's about the limit to be usable. Oh ! I forget Apple device too, it's a good option to run this model I guess, but still slow.
Yet, as you said, it's still a wip implementation, it may improve soon (MTP support is about to be merged in llama.cpp soon).
Yes, I'm not a native English speaker, and I'm definitely not an AI. ;)
I can't say the same thing because I just don't notice typos (mine or others).j I just assume my English is bad.
Isn't the definition of "being human" is "not to be perfect"? In French, it kinda is. We say "He/she humain after all" to mean that someone made a mistake.
Yes... I'm not english native. And definitively not an AI.
Disclarer: I'm unsing Vulkan on an AMD GC.
in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47
That is a massive cost reduction.
Refs: https://www.alibabacloud.com/help/en/model-studio/model-pric... https://runware.ai/gemini-omni
The last one was: Qwen3-Omni-30B-A3B https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct
And maybe Qwen4 won't be released, they only release Qwen3.8 27B (and a mostly unusable 125B). There are definitively slowing down open weight release.
So, you may actually have very good performance with local model. Just not yet on *every* device. So the Mozilla strategy here feel very reasonable. A Cloud provider specialised in local models, to be able to switch once local models will be quick enough on most devices.
Big claims, expensive and not release to the public yet.
What a bunch of amateurs. Here is it anyway :
https://ache.one/gpt6_now_down.png
The claims: https://share-md.com/view?id=870ba228-a25c-4169-bbc9-12d7f25...
And some others like this bugged Karts Game:
https://tidal-rush-paradise-gp.skirano.chatgpt.site/
This impressive spaceship construction game:
https://voidexplorer-shipyard.openai.chatgpt.site/?fleetSeed...
And a lot of graphs, some without even Astra on it. Oh and the logo is a Galaxy.
Can you do more or less?
27B local model just dropped, it's 6/8-month old SOTA. General ROI of AI investment is expected on a baseline of >10y.
Also, a lot of people don't really care about german language capacity, maybe people programming in DDP idk.
PS: You benchmark seems saturated. Most values sit @>75% in a benchmark generally indicate that it's no longer as useful as a <70% one. I mean, Qwen3.8 is 77.5% and Fable5 80%, the poll of values is from 65% to 90%.
1k lines are already shared, seems legit. Most of them is just <30k€ people. Only a handful of millionaires (8 >10M if I remember correctly), 0 billionaires. 2/3 people, 1/3 pro, there is no information about tax of professionals.
Seems to be an API endpoint about people who ask questions to the DGFiP.
Hopping for an AgentWorld variant from Qwen but I guess, I have too high expectations.
Like:
$ podman run -it --rm -v .:/workspace local-dev-ia /usr/bin/oc
Configured with a .env file. Hope to do it hopefully before the end of the week.
The progress compared to Qwen3.6 27B is good, not that impressive, it's a 4 months old model. (kuto to them to compare to 27B dense and not 35B MoE, it's more fair to do so). It is very probable that Qwen3.8 27B will crush Glimmer-30B on most benchmarks.