Ah good catch on the total tokens, was going off vague memory there, and I thought people had gotten qwen 3.8 27b up to similar decode speeds as ds v4 flash.
>One data point that seems relevant to me is that the previous gen qwen Qwen3.6-27b was not so different in performance from its sibling model Qwen3.6-35b-a3b. We never got a qwen3.8-35b-a3b, but if we had, would the gap have stayed the same or gotten bigger? I.e. would the quality gains by improving training coming up against a hard limitation with 35b, or not.
Yeah good question, kind of shocking that a 3b active model would perform as well as a 27b dense.
First off, I'd include Qwen flash-next and GLM 5.3 to show some of the other strong open weight models, and they predictably dominate it, but they're much larger. But, it shows up right next to DSv4 Flash 0731 on the overall index, and that's much larger. It's a great model! But then scroll down and hit Time Per Task, and you'll see that DSv4 Flash takes 3.6 seconds per task to Qwen's 21.1. That's what I meant when I said this:
>speed due to excessive thinking maybe to make up for the smaller amount of world knowledge baked in (qwen 27b's main issue iirc), etc - they're tuned for different things.
It can make up for its shortcomings by iterating a lot longer, and using way more thinking tokens. And that's a great trade if you don't have the vram to run the bigger models, but speed is pretty important for getting things done... And that's why DSv4Flash is great, too, despite being much larger, and scoring similarly on the intelligence index.
Ah sorry, I should've continued, the bigger recent models are commensurately smarter. If you really want to make the point, then you'd need to show 27b being smarter than similar vintage bigger models. And in that case, there's confounding issues like efficiency, speed due to excessive thinking maybe to make up for the smaller amount of world knowledge baked in (qwen 27b's main issue iirc), etc - they're tuned for different things.
Gpt-oss—120b is like 1000 years old in AI years, whereas Qwen 3.8 27b is pretty young. What you’re seeing is that parameters aren’t apples to apples, and at a given parameter level, the new models are much, much better than the ones from a year or two ago. Like, to a comical degree.
Yeah, it's really tough, there's so much potential good, and so much potential bad. My intuition says that decentralized will probably be better, the other way seems to concentrate power while also not really keeping it out of the hands of anyone. It doesn't seem likely that the US firms will be able to convince the Chinese to agree to an international regulatory framework for this stuff. Opinion polls there have the populace being quite optimistic about the benefits of this stuff, and I think the cat's too far out of the bag to really put it back in, anyway.
Really? You’ve spent the time to learn every member of the EU government that supported them? And you think that’s a good use of everyone’s time? I guarantee you that the average US citizen hasn’t bothered to do that with even their members of congress, many don’t know the name of their representative, let alone those of a foreign government, they’re generally much more concerned with how they’re going to make enough money to pay to live as costs rise. Not to mention the flurry of domestic cultural infighting related concerns.
You’re certainly free to hold whoever you want responsible, but I’d argue it’s not a realistic position. And keep in mind that it’s hard for someone living here to have an accurate view of what the average American believes, even harder for someone living outside the country - dominant opinions vary widely between regions of the US. I’m sorry that our government supports war criminals of varying degrees on our behalf, though, would certainly be nice if our leaders were more principled on that front.
No, I don't blame China or India for continuing to do business with Russia, it seems completely rational for them to do so. They're powerful enough to ignore the structures the US has built to try to exert leverage internationally. Seems like you're building lots of straw men in your mind of what US/EU citizens actually think, or the propaganda apparatus you're using to keep on top of this stuff has built them for you.
To be clear, I don't think they're doing it because they want all the money, I think they truly believe that this is existentially threatening, and are trying to regulate their open source competition away because they think that's the way to be safe. That they can be trusted with the power, but if it's out and about, that's unsafe - they like to compare it to nukes, and they want nuclear non-proliferation.
Geopolitics is way too complicated for such black and white thinking. Do you know who amongst the governments of those countries decided to back this Syrian president? I certainly don't. What percentage of the (huge) populations that you're damning as "deeply disgusting, sanctimonious hypocrites" do you think know who the current Syrian president is, his history, know who amongst their governments backed him, voted for the people who backed him? Humans just physically can't keep track of everything, especially things so far outside of their daily life and physical experience. You can't reasonably hold them responsible for everything bad that happens in the world that their countries contributed to, especially things that are entirely abstract for them.
Presumably you can install your own and have the electrician to the final connection to the grid/your house? Solar panels are much, much cheaper than you seem to think.
Really? That's more expensive than I would've guessed. The US has substantial tariffs, and you can get a pallet of panels for $0.20-$0.30/watt. You can also pay a lot more, but I'm not sure why you would.
My understanding is that it’s the thin film variant that’s “famously poisonous”, the Cadmium Telluride ones. But that’s not the type that almost any field full of solar is made up of.
Yeah, I'm guessing this isn't unique, but I remember the first time I used ChatJimmy, I missed the fact that it had responded because I was still hitting the enter key, and getting ready to see tokens stream in, but they were already all sitting there, and I'd missed registering the visual diff somehow.
Yeah, it'd be a lot easier to maintain flow, less need to work on more than one session at once, etc. And then tool calls would be the limiting factor, especially network access. I hope AMD keeps the project moving forward post acquisition.
Sure, Google could be their agent, or ChatGPT, or Claude, or whatever local model the person is running. The big guys might all have the results cached so they don't have to rerun the crawl. Whatever their choice of agent, it seems pretty clear that almost everyone's going to use them, they're way too useful not to.
Nah, Taalas was putting the weights into silicon as a mask ROM. Their demo chip was hardwired to serve Llama 3.1 8B, and could never be updated. New models, even new versions without any architectural/size changes meant new tape outs.
But in exchange, you get insane speed and great energy efficiency. I could see it being a great approach for basic "good enough" models.
They may have had a little flexibility by supporting finetuning via LoRAs.
Ahh yeah, I've seen people using those. I was thinking you'd ideally want to connect them full speed, but it looks like they're 200gbps ports anyway, and I guess even if they were 400gpbs, you don't really need that much for cross-GPU traffic for MoE models doing tensor parallel.
No, the asic could only ever run one model/set of weights, no updates possible, ever. These are general purpose processors that can have their models updated. But the chips are enormous, with a substantial amount of on-die memory alongside the execution units, for a relatively insane amount of memory bandwidth.
Essentially, yes. The DGX Sparks, at 128 gigs for ~$4700 are one of the cheapest ways to get enough high-ish speed memory cobbled together to run one of the more capable open weight models to run at home/for a small biz. This switch lets you combine the memory of 2-4 of them.