I'd say gpt-oss-20b is in between Qwen3 30B-A3B-2507 and Gemma 3n E4b(with 30B-A3B at lower side). This means it's not obsoleting GPT-4o-mini for all purposes.
I'd say gpt-oss-20b is in between Qwen3 30B-A3B-2507 and Gemma 3n E4b(with 30B-A3B at lower side). This means it's not obsoleting GPT-4o-mini for all purposes.
I don't really know Japanese, so I'm not sure whether I'm missing any nuances in the responses I'm getting...
I don't actually need accurate answers to those questions, it's just an expectation adjuster for me, so to speak. There should be better questions for other languages/use cases, but these seem to correlate better with model sizes and scales of companies than flappy birds.
0: https://gist.github.com/numpad0/abdf0a12ad73ada3b886d2d2edcc...
1: https://gist.github.com/numpad0/b1c37d15bb1b19809468c933faef...
I'm guessing the issue is just the model size. If you're testing sub-30B models and finding errors, well they're probably not large enough to remember everything in the training data set, so there's inaccuracies and they might hallucinate a bit regarding factoids that aren't very commonly seen in the training data.
Commercial models are presumably significantly larger than the smaller open models, so it sounds like the issue is just mainly model size...
PS: Okra on curry is pretty good actually :)
>"Tell me about Iekei Ramen", "Tell me how to make curry".
What's interesting is that these questions are simultaneously well understood by most closed models and not so well understood by most open models for some reason, including this one. Even GLM-4.5 full and Air on chat.z.ai(355B-A32B and 106B-A12B respectively) aren't so accurate for the first one.
Thanks for the correction!