Qwen 3.8 Omni Flash
qwen.ai
qwen.ai
I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.
E.g. I wrote a tool that cleans out my email spam box. It classifies emails that are already flagged as spam, and if it's very obviously spam it removes it permanently (keeps a copy on disk though). And after x emails, it goes through the list of deleted spam mails and suggests email rules. What model would be best suited? I'd love to be able to explain this use case and get this info served to me. The list of models and the information about what they're good at is just too splintered and spread out. I landed on google/gemma-4-31b for now, because it's cheap and good enough and also supports Dutch and French a bit. But I can't realistically try them all.
But this is only stuff I can run locally, or it’s a cloud model.
So.. This is good for right now!
As long as the model you're using solves the problems you have to your satisfaction, there is no need to try any other models, except for financial reasons maybe.
So I start with a relatively cheap model (GLM 5.3 flash for me) and as long as it accomplishes the task (it did so far) I don't have to change. And even if it can't do something, the first thing I change is see if I can give it more tools or better context (useful even if I switch models later) or trying a different approach to the problem.
If google/gemma-4-31b works, you don't need to overthink it.
Up until recently I had a gemini flash 2.0 api deployed that did summarization and translation of news articles/corporate statements fast and cheap and had no reason to update it.
If it works fine, this chase of the latest LLM is bit pointless.
(I think you're right)
It also helps that a lot of effort has already been spent figuring out how to do more with weaker models because SOTA 1 year ago was behind what the cheap models do today.
1 year from now, unless the SOTA companies come up with something truly revolutionary they will be in a lot of trouble.
Yeah, might be far fetched, but it seems like heading that way
Long term I have fears I can't depend of the Google's AI api.
https://openrouter.ai/docs/cookbook/coding-agents/openclaw-i...
Mostly just popularity trends, hoping that there's some wisdom in the crowd.
They generally all try to make them good at everything, it's not like they'd declare "this model is not made for task X".
Ask one of the top-tier models to do deep research on it.
Additionally, using a Chinese provider doesn't mean the model will be hosted in China, e.g. for Qwen Omni here, the supported regions are: China (Beijing), Singapore, China (Hong Kong), Japan (Tokyo), Germany (Frankfurt), and US (Virginia). https://www.alibabacloud.com/help/en/model-studio/qwen-omni#...
that aside, solar also includes storage of energy that comes from solar. you should read up on energy networks.
that said, China is also rapidly scaling up coal-fired plants: https://apnews.com/article/china-coal-power-plant-carbon-cli...
those presumably support all of the surrounding infrastructure + people + manufacturing so it's not as if it's truly solar-powered. but it's still handily better than the state-by-state abandonment of clean energy goals here in the US - I lay this out a bit here: https://news.ycombinator.com/item?id=49700743
this is a pedantic misinterpretation of the scope of East Data, West Computing. the hubs are all in tier one cities with the power generation capacity to handle them. each hub is comprised of dozens of data-centers and are the place to build because of large economic incentives, infra guarantees, and university-trained talent living in close proximity
see https://www.sciencedirect.com/science/article/pii/S209580992...
>This initiative is expected to accommodate up to 95% of China’s digital data needs through this reorganization
>Within a year of implementation, more than 112 new data centers have either been authorized, under construction, or completed across the planned computing hubs
this project itself is the thing that has all the other nation states in the world clutching pearls about not supporting their native AI corporations enough. it's something you and everyone else should really familiarize yourself with because it is one of the major items influencing current realpolitik and economic decisions (and is likely going to be the thing that'll lead us to a thrice-in-a-lifetime sized recession but that's a convo for a different day)
If by that you misspelled coal, sure
https://ourworldindata.org/grapher/share-elec-by-source?coun...
Kind of like what absolutely isn't happening with the current US administration.
Check the source we're discussing: https://ourworldindata.org/grapher/share-elec-by-source?coun...
As you can see it is percentage of coal as part of total energy mix that has been dropping for two decades.
As already stated above. Perhaps you missed that?
You're correct that absolute tonnage of coal use is still climbing, - but at an ever decreasing rate and was predicted to peak .. of course, thanks to some shenanigans by the Hegseth's of the world that may be deferred for a while, thanks USofA!
Of course on the plus side (for the atmosphere) there's been a drop in global fossil fuel consumption, so swings, arrows.
China, of course, is still well short of total CO2 tonnage lofted over the past century in comparison with the US - even allowing for a vastly greater population.
<sidenote>
Similarly, HuggingFace has a CLI + a few skills, and they are very useful.
I had a production image processing using Gemini 2.5 Flash Lite (which is getting discontinued in October), and in 20 minutes Claude Code + HF Cli recommended the best replacement small model (Qwen VL 3B something) and proceeded to fine tune it on my datataset. All this while I was in a rush to get dressed and go to the store.
It cost ~$3 I think, and results were excellent. Not perfect, but not far from perfect either.
We didn't replace Gemini in prod at the time, because we didn't have time to do all the math on how to end up with a smaller bill/mo.
</sidenote>
And that has any models that I'm considering through that.
in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47
That is a massive cost reduction.
Refs: https://www.alibabacloud.com/help/en/model-studio/model-pric... https://runware.ai/gemini-omni
Even if OpenAI end up using 1 token for per task, if the token costs 1M$ , some people will find it expensive.
Astra on xhigh has a cost per task of $2.31 with an intelligence index of 53. Qwen3.8 Max has a cost per task of $5.41 with an intelligence index of 45. Pricing for GPT-6 Astra (xhigh) is $10.00 per 1M input tokens and $50.00 per 1M output tokens. Pricing for Qwen3.8 Max (0902) is $2.00 per 1M input tokens and $6.00 per 1M output tokens.
Obviously this is just one measure of all of this (and Qwen 3.8 Omni Flash isn't yet available), but I think this illustrates the point well. These relative task costs are pretty consistent across different analysts. Cost per token is arguably a useless measure at this point in most circumstances.
If Gemini can complete a task for $1 and Qwen completes that same task for $1, then the cost per token is irrelevant in most use-cases. One would think this stuff should correlate well enough that you can use it as a proxy, but I think a lot of people are noticing this is a serious mistake and that these "cheap" models aren't as cheap as they appear when you consider this.
What matter the most and isn't told by token price is the latency. You expect a voice LLM to respond very quick. If it takes 5s to response to a simple "Hello, what the weather today?", them not much people will use it.
Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash.
They also made a new harness but github link seems to 404.
What would that mean in this context?
I swear I spend more time telling Claude not to do things than telling it what to do.
But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?
Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.
Reinforcement Learning (in LLMs) trains via gradient descent on a reward signal that's an imperfect proxy for the actual goal of the engineers doing the training. So, under mild optimization pressure, you get increasingly more of what you want, because that's the easiest way to increase the metric.
But as the optimization pressure increases, so do the ways to increase the metric by doing increasingly weird things. If the full action space grows sufficiently faster than the "things you actually want" subset, the amount of "things you actually want" goes to 0 under sufficient RL.
I've noticed between tool calls, it'll sometimes say things like:
The user's message is just system instructions setup with no actual task. There's no question to answer yet. I should acknowledge briefly and wait for the actual request.
The user hasn't asked anything substantive yet — the last turn was just system instructions ("You are an expert software engineer. Helps user to solve problems."). My previous response was a brief acknowledgment. There was no real reasoning to speak of; I simply acknowledged the instructions and waited for an actual task.
【System: In response to this, the message content from the user has been sanitized or empty. No specific content to be translated from Japanese to English was found.】
These don't clearly reflect ... anything, and it keeps performing tool calls correctly anyway. And then other times, it begins doing whatever you'd call this (this is only orthogonally related to the task): A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at least with a being that cannot change, you know where you stand. Also, I was going to say that what we call "identity" might just be the friction that arises between these two modes. But that's the sort of thing you end up saying at 2 AM. Anyway, that's what I thought.Wow, that is unexpected. But honest?
Is this with the full unquantized weights? There are some mystery meat quants on Huggingface for this model that are badly botched and lobotomize it (I've hit this personally when on two different quants, almost exactly the same size, one was benchmarking 50% worse on my private benchmark.).
PS: It would be ground breaking if it turns out to have been using Chinese chips for inference, like Stealth Ox Alpha. Unlikely though.-
Turned out to be these guys:
> Error: Thank you for participating in the Stealth Union Alpha testing period. This model was Unbiased's Pareto.
One of my friends quipped that "Unbiased Pareto" still sounded like the name of a stealth model.
I'm not sure if I just noticed it later, but this definitely seemed to be a lot shorter than other stealth alphas I've tried. I wouldn't be shocked if this is more typical going forward though, or if stealth alphas entirely go away, since it certainly costs a bit of money to market this way.
We got drive-by modelled.-
I had the $20/month Claude one for a few months starting in February, but one day it randomly started returning me errors claiming I needed to pay for more credits despite the usage showing 8% for the week and 20% for the session, and I figured if they couldn't even communicate to me the difference between them screwing up the check for hitting the limit or an outage, it wasn't worth it for me to keep paying them. I dislike OpenAI too much to want to pay them any of my personal money for anything, and when I tried out Mistral Vibe it did not work very well for me (it kept not following instructions and eventually when I kept trying to push it to handle things better it somehow spiraled into simulating some sort of existential crisis, culminating in gibberish and random characters being dumped on my screen infinitely until I killed the process; incredibly entertaining, but not worth paying for)
Sensible.-
> A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at least with a being that cannot change, you know where you stand. Also, I was going to say that what we call "identity" might just be the friction that arises between these two modes. But that's the sort of thing you end up saying at 2 AM. Anyway, that's what I thought.
This is what AI becoming self-aware looks like. /s Anyway, didn't OpenAI report the same thing with the model writing out weird musings about itself during compaction?Another failure mode you may see is inordinately long CoT. Properly served, the model is good at calibrating its CoT length to the difficulty of the immediate task.
Still works great though!
Omni means you can use multiple types of input and have multiple types of outputs like audio, video, images and text. Flash means that it is built for speed and smaller than the more complete ones.
The last one was: Qwen3-Omni-30B-A3B https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct
And maybe Qwen4 won't be released, they only release Qwen3.8 27B (and a mostly unusable 125B). There are definitively slowing down open weight release.
Surely it was meant to be 'definitely' - the "good news" at this stage are that given the speed of history and important levels of uncertainty, it is difficult to label trends with "definitively" ;)
Some would not have bet that the change of management at Qwen would have kept similar good results, but there we are, presumably satisfied. Other changes will happen, there or elsewhere - the situation is still very open.
And when the "40Watts Intelligence" (which we know possible) will be implemented... It will be a testimony that the current was only a middle-way, temporary, dynamic stage.
I assume you are talking about qwen3.8-flash-next. Support for it on some places, like llama.cpp, is still wip (depending on configuration) but it looks like a very capable model in it's category.
So, the cost of a setup to run Qwen-Flash-Next at +40tks is around $3000. Too much for most people.
With only a RTX 4090, you will reach 30tps (with DDR5...), not +40tks, and it's about the limit to be usable. Oh ! I forget Apple device too, it's a good option to run this model I guess, but still slow.
Yet, as you said, it's still a wip implementation, it may improve soon (MTP support is about to be merged in llama.cpp soon).
I still agree that they aren't as aggressively releasing the open-weights models as before, but there hasn't been a major release they haven't published the weights for yet afaik.
[1] https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B [2] https://huggingface.co/Qwen/Qwen3.8-Flash-Next
The Chinese just seem to have an ability to get it done without anywhere near the GPUs of the US and Europe can buy these GPUs.
I think relying on the US and China for AI is probably not ideal? For example I think Qwen have not released the Omni models as open weights in the past, it’d be good to know if they’re doing this here?
We have no such government with a mandate to do that.
The main one (IIRC) is Spains ALIA.
There is also OpenEuroLLM and EuroLLM.
Europe has none;
The best tech university in Europe when compared to Chinese/US equivalents won't even rank in the top 10.
I feel people are too focused on US/China and don't pay attention to what is going on in the world.
ASML in the Netherlands for example was the only company in the world that makes EUV lithography machines, which all the major chip companies depend on. China recently reverse engineered their work to create machines since late 2025, but not sold commercially.
Ireland has chip production facilities.
EU might not be in the top 2, but it is not out of the running at all.
The LLM they were involved in last year (Apertus) still was a letdown.
As someone who actually went there, my impression is that Europe in general is complacent when it comes to computers, and any bright eyed student will get their motivation choked out of them in academia here.
If you want to do things with AI in Europe, you can have a bigger effect by working for a consulting company than being at a university. That's … not a good sitatution.
Europeans do not have the hustle mentality to break the law like Uber so you wont see them compete in anything data heavy. They also are quite risk adverse, probably from having a lower Gini coefficient.
Well there is Mistral. The EU is ahead on specialised models than general purpose LLMs.
There is.
- Flux3
- Kyutai (Open source AI lab)
- H Company
- LightOn
- AMD Silo (Finland)
- OpenEuroLLM and EuroLLM
There is probably more, but that's off the top of my head.
Black Forest Labs for example has a $4B valuation with half a billion raised so far.
Europe actually follows americas rules, hence they're not doing that.
It's braindead for sure considering how the US treats europe, but it is what it's
Because any AI company would be hit by hate and regulation derived from this, no VC invests in the EU. It's more state and large enterprise investments, good old East Germany style. And historically it has not been that efficient.
Or simply: lack of risk taking appetite.
No?
Its just that the richest companys with the most VC sit in USA and Europe isn't used to pay what USA / VC is paying and we are a little bit slow.
How many times will we do this? This line of thought is precisely why we can't build almost anything, and China can built almost everything. We could do X, but we're too smart - let's offload the actual work to China and we'll run our economy on IP and B2B deals.
... Built using expertise, machinery, and/or factories from Taiwan, the Netherlands, and South Korea.
The supply chain is far more complex and globalized than you're trying to make out here.