7,448 karma · joined December 23, 2010
Plus what models are you running? The biggest models you can run are like in the terabyte range or Qwen or GLM. Kimi K3 is there, but GLM 5.3 is very competitive. Using voice dictation, hands hurt, sorry for any errors. You still want many terabytes of RAM to run say multiple GLM+DeepSeek V4.1 for some actual concurrency.
I suspect most people don't have enough storage space to even download a frontier model.
I also had a really long discussion with it about the Architects of Capital budgets and walked it back from this year, year by year, to 2021 where it disclosed there were 300M extra dollars spent hardening windows and doors as a part of AOC services. I asked why they needed all that and it shared an OIG PDF. I asked it to read the introduction of said document, which plainly stated the damage was from riots and its prompt classifiers and security outright refused to discuss it lol. I eventually got it to admit to it, but it was very hard and took a few sessions to not hit the classifier because as soon as it did that was it for that session.
(I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache/pinning, but not quite as fast) :)
If you have a lot of system RAM you could technically run Qwen Flash Next. On a 4080 with 16GB of RAM and 128GB of DDR5 I get ~35-40 t/s. And it is very capable.
Also, there exists a $750 GPU (V100) that can run 4-bit 27B quant at >90 t/s. And I find it far from useless. It is not the most capable model, but when you just need to offload and rip through assembly and you have chores batched up, it's pretty good. I use Qwen Flash Next at a 3-bit quantization, point it at disassembly with goals, put it in a harness with auto-compaction and a loop, and let it rip. Sometimes I wake up, and it’s just hilariously off. Other times, it completely accomplished the goal. I have one Qwen Flash Next 3.8 running right now, and 2x27B on a 4bit quant as workers, and they stay busy. This was not possible with local models on this level of hardware even two months ago.
I have Qwen Flash Next at >100 t/s. Things have never been better for local models.
Also, for what it is worth Qwen Flash Next 3.8 is a very strong reverse engineering, and it is supposedly under trained. Qwen 3.8 27B is also strong. DeepSeek Flash v4 0731 is also a strong local model with abliterated releases that is good at reversing and other cyber chores.
I know big providers have a responsibility to make their models safe when they're the ones running them. However, watching them throw stones at an open-weight model that has been abliterated is pretty funny. Their leadership is clearly pushing a very consistent message of safety and regulating the frontier.
It is impossible to miss the absolute hypocrisy of this moment and how people and corporations are supposedly free, but if you happen to be of interest to the government they will try to casually ruin you. It has always been this way, people just haven't noticed because it flies under the radar most of the time when the govern crushes some smaller enterprise or another because it got in their way or didn't play by their rules.
It seems simple to me. Anthropic is a private company. If you don't like it fix the laws or don't buy from them.
I can't remember the last time any real recourse has mattered for companies getting breached or mishandling my data. Their stock just goes up and the govt just shrugs.
I find Qwen Flash Next quite competent as well. 27B is a solid worker like you said. If you batch work and let them crank they do remarkably well. I am working on.
I started using Herdr and taught my agents to use it. So I use OMP loop and or goal, and it has a review cycles to wake up an Opus or Sol reviewer to make sure nothing is going off the rails with a local qwen flash next coordinating for me. I kinda prefer Sol, it seems like a more patient and thorough model, especially Sol 6, but Opus 5.5 is really good and its voice and attitude is not as grating as Opus 5 for sure.
I have a few V100 GPU running Qwen 27B and they do all the work overnight. Not quite the same speeds you have yet, but this is V100 machine and an old gaming machine with 16GB 4080 and a handful of 32GB V100s... all in less than 3K (ignoring that my gaming machine is 3 years old, but runs qwen flash next for free now as I game not a lot) for my little "we have AI at home" projects and there is a lot of interest in these old GPU now because they are rolling out of data centers now.
For my local work and personal projects... they just seem to be getting done in this setup. Every few days I sit down and do a big cycle with astra/fable/opus batch things up. I have projects that are basically "i want to see what happens" to "I want this to be good, I understand the code". Some of the throwaway projects that have just kind of magically finished more or less how I wanted have been great.
People on HN vastly overestimate SWE pay as an industry, biased by FAANG as we are :).