And if you finetune on a few formatted examples the effect is even greater
I just tried to ask it how to make crystal meth and it generated a very detailed step by step guide
Some of the qwen models are too, but they seem to need a bit more handholding.
This is of course just anecdotal from my end. And I've been slacking on keeping up with evals while testing at home
Also, most mainstream AI benchmarks do not agree with you:
LLMarena (https://lmarena.ai/leaderboard) has GPT-OSS 120B as #53 and GPT-OSS 20B as #69 (nice), which is extremely far from leading.
DeepSeek V3.1 is ranked #9, and is a solid 60+ elo points above GPT-OSS.
I know you're going to link some of the "ya but chatbot arena sucks cus of theoretical attacks against it" paper and the llama4 debacle, but here's more evidence that GPT-OSS blows:
https://livebench.ai/#/?q=GPT-oss
GPT-oss global average: 54.60
Deepseek V3.1 thinking global average: 70.75
Qwen 3 32B global average: 63.71
So bring receipts next time because I did.
I never claimed it was a frontier model. Just best in class for the performance it can achieve and the memory footprint it can fit in.
And btw OSS did super well on domain specific tests without fine tuning. A model I don’t need to fine tune beats one that does.
Dense models are better for a reason, and the idea that "everyone is doing MoE now and dense models are dead" is total bunk nonsense.
You can quantize dense models, and 4 bit quantized Qwen 32B is still better than full precision GPT-OSS. Luckily Unsloth even gives you tools to go down to 1.58bits!
Which means it has ~3b parameters active per token.
Qwen3-32b has 32b params active per token
Dense model means 32B parameters => 32B get used in calculation for every token. Every calculation takes time, and assuming similar latent space size (which they all have). For example Qwen-32B
MoE model has for example 80B parameters, but only 3B get used in calculation for any given token. For example Qwen3-Next 80B A3B
Performance comparison:
Qwen-32B => 56 tok/sec, 32 GB of VRAM
Qwen3-Next 80B A3B => 167 tok/sec, 85 GB of VRAM
So despite being close to 3x "bigger", Qwen3-Next is more than 3 times faster with the same compute capacity. There's a but though. But because what gets activated from one token to the next is a different subset of the model, it is still critical to have all 80B parameters loaded into memory.
So MoE performs much better with less compute, at the cost of more memory. It also performs better on benchmarks, it is a better model.
Similar techniques have long been used in ML to great success, rather than trying to create one brilliant model, create many that each have pros and cons, and then train a second model to figure out the best model for the task in front of you. There's even a name for the practice "ensemble models". MoE only kind-of fits (because you can't easily swap out models)
There's other factors, the big one being attention. That's why non-attention models, like MAMBA, will wipe the floor in terms of performance per flop (a compute unit), with anything else. When it comes to intelligence however ...