HNHacker News
TopNewBestAskShowJobs

bitexploder

7,448 karma · joined December 23, 2010

I did not keep blogging.
submissionscomments
bitexploder··on Gemini 4 Argon
I was being dramatic, but no normal person is running 20x DGX spark. That is 100K of slow VRAM. Not sure it makes a lot of sense TBH.
bitexploder··on Gemini 4 Argon
To be comfortable you need 300 to 400 terabytes of storage. You need the GPUs and the cooling for them. They don't need a lot more. You put models in memory and you do math really fast and make tokens. You need some CPU and RAM infrastructure for like prompt caches and things like that, which does add up, but the dominant cost is by far the GPUs and VRAM. Everything else is a pretty easy step down to afford. You probably need to have multiple different models, right? Like if you want an Astra plus a Sol plus, you know, like a Luna-class model, say, you probably need like 10 terabytes for Astra, probably 3 to 4 terabytes for Sol, and 1 to 2 for Luna, maybe less, maybe one or half a billion even for Luna. That's RAM. You need So all total you need 15 to 16 terabytes of active VRAM to keep those models resident. That's not counting what you want to keep in cash and whatever hosting constraints these models have that aren't public. That much VRAM is five to eight million dollars right now.

Plus what models are you running? The biggest models you can run are like in the terabyte range or Qwen or GLM. Kimi K3 is there, but GLM 5.3 is very competitive. Using voice dictation, hands hurt, sorry for any errors. You still want many terabytes of RAM to run say multiple GLM+DeepSeek V4.1 for some actual concurrency.

bitexploder··on Gemini 4 Argon
We have had this discussion for a decade regarding cloud hosting. There are definitely workloads and AI inference that you want to do in-house. Hosting these GPUs and models and setting up the interconnect and balancing it is really hard right now. It's not traditional infrastructure. It is exotic. Typical enterprise is at such a disadvantage to try and host this completely themselves. And models are constantly updating. And what models are you going to run? Six months ago, things that probably cost you a billion tokens probably cost you 50 to 100 million now, given how much more efficient models have gotten, or at least they cost that much. I don't know. It's just not a given.
bitexploder··on Gemini 4 Argon
No one on the planet could run a frontier model. The amount of VRAM you need would bankrupt a normal person.

I suspect most people don't have enough storage space to even download a frontier model.

bitexploder··on Gemini 4 Argon
It is definitely smarter than that. It is mostly mannerisms and how it likes to work. I would take it seriously as an Astra or Fable or Opus 5.5 level model. It just needs polish, but where and how you harness and use it matters a lot. But it has amazing long horizon attention and gets things done.
bitexploder··on America.gov goes crazy on "play Minecraft"
Nothing I can do will go as hard as a frontier model with zero safeties. Trust me. They are downright scary in how sneaky they get defeating classifiers and safeties.
bitexploder··on America.gov goes crazy on "play Minecraft"
https://pastes.io/n8hmwfcs same/similar.
bitexploder··on America.gov goes crazy on "play Minecraft"
I used NIST docs and convinced it to teach me about hashing and conduct table top incident responses with me.

I also had a really long discussion with it about the Architects of Capital budgets and walked it back from this year, year by year, to 2021 where it disclosed there were 300M extra dollars spent hardening windows and doors as a part of AOC services. I asked why they needed all that and it shared an OIG PDF. I asked it to read the introduction of said document, which plainly stated the damage was from riots and its prompt classifiers and security outright refused to discuss it lol. I eventually got it to admit to it, but it was very hard and took a few sessions to not hit the classifier because as soon as it did that was it for that session.

bitexploder··on Gemini 4 Argon
Some of the time AI says the darndest things.
bitexploder··on Gemini 4 Argon
Opinions my own but I have been using this model for a bit. I would say it is a good model and the skill with which people use AI varies widely.
bitexploder··on Gemini 4 Argon
Evidence needed. I think for certain kinds of outcomes it has very strong advantages, but these advantages are not a given as 'best for LLMs' :)
bitexploder··on The AI Race Just Got Awkward
V100S 32GB, I have had Claude optimizing it for about a week and it is already at around 900 t/s prefill, 90-100 t/s output in Pi on coding tasks. There is also a Ninfer fork for the v100 but it requires a custom format. I am working on upstream Unsloth with GGUF 4-bit quant.

(I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache/pinning, but not quite as fast) :)

bitexploder··on The AI Race Just Got Awkward
Well, I have a $750 card that runs at about 50-60% of that token rate :)
bitexploder··on Livenerf: Has Opus 5.5 been nerfed yet?
Trimming parameter size that can be reduced while surviving regression evals. They have so much data they know exactly where to shave the models. Most people will never see it in their work loads. It won’t affect core benches because that is part of the regression evaluation.
bitexploder··on Livenerf: Has Opus 5.5 been nerfed yet?
I suspect they play with their quants and perform weight sensitive tensor/parameter tuning among other things to get serving faster and some of the time for some workloads it surfaces. I feel this has a high probability of being correct and an explanation for some of this.
bitexploder··on GLM-5.3 and the spread of advanced cyber capabilities
If you are really invested and have some system RAM you could get a 3-4 bit quant Qwen 3.5 35B-A3B running. There are builds that do expert caching, keeping the hot experts in cache. For something like disassembly, you're looking at being able to fit, if you have, say, 11 to 12 GB of VRAM, you could get at least three hot experts. For pure disassembly tests, I would say that would be pretty fast. A 4-bit quant is pretty decent and maintains most of the smarts of the larger quants. Depending on the GPU I would expect a decent token rate. It is medium strength local model, but if you harness and ground it well I expect it can reconstruct C code for you. The quality of your disassembler will matter here.

If you have a lot of system RAM you could technically run Qwen Flash Next. On a 4080 with 16GB of RAM and 128GB of DDR5 I get ~35-40 t/s. And it is very capable.

bitexploder··on GLM-5.3 and the spread of advanced cyber capabilities
I absolutely do not want a public provider having any of my data for reverse engineering work. That is a hard pass from me.

Also, there exists a $750 GPU (V100) that can run 4-bit 27B quant at >90 t/s. And I find it far from useless. It is not the most capable model, but when you just need to offload and rip through assembly and you have chores batched up, it's pretty good. I use Qwen Flash Next at a 3-bit quantization, point it at disassembly with goals, put it in a harness with auto-compaction and a loop, and let it rip. Sometimes I wake up, and it’s just hilariously off. Other times, it completely accomplished the goal. I have one Qwen Flash Next 3.8 running right now, and 2x27B on a 4bit quant as workers, and they stay busy. This was not possible with local models on this level of hardware even two months ago.

I have Qwen Flash Next at >100 t/s. Things have never been better for local models.

bitexploder··on GLM-5.3 and the spread of advanced cyber capabilities
This just makes me want a home lab capable of running GLM 5.3 at a 4bit quant.

Also, for what it is worth Qwen Flash Next 3.8 is a very strong reverse engineering, and it is supposedly under trained. Qwen 3.8 27B is also strong. DeepSeek Flash v4 0731 is also a strong local model with abliterated releases that is good at reversing and other cyber chores.

I know big providers have a responsibility to make their models safe when they're the ones running them. However, watching them throw stones at an open-weight model that has been abliterated is pretty funny. Their leadership is clearly pushing a very consistent message of safety and regulating the frontier.

bitexploder··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
I have likewise not been impressed with Astra 6 for most things. It is good, but Opus 5.5 seems just as good or better and I have had Opus 5.5 workers just... hammering since release and cannot spend all of my quota yet.
bitexploder··on Sonnet 5.5
If you care about the things terminal bench cares about, yes. Sonnet was probably trained aggressively on agentic coding and things that align well with deepswe and terminal bench and or tuned heavily for those tasks. Sonnet is an agent likely to do more of those tasks and be given the more grunt work tasks. Whilst Opus' wider knowledge pool means it can deal with a much higher variety of real world situations successfully. And, those benches are often timed or limited. Opus may have been running out of time. Looots of factors.
bitexploder··on U.S. appeals court upholds designation of Anthropic as supply chain risk
That is the funny thing to me. I hate diving into politics here on HN, but this has such a rank smell to it. Trump's comments about Dario. His insane propaganda video about how we will never be communist / Marxist. In the same breath our government is bribing people with offers of $$ (the $5,000 checks) and essentially bullying to outright moving towards saying this technology is so powerful we will just have to seize it for our own use.

It is impossible to miss the absolute hypocrisy of this moment and how people and corporations are supposedly free, but if you happen to be of interest to the government they will try to casually ruin you. It has always been this way, people just haven't noticed because it flies under the radar most of the time when the govern crushes some smaller enterprise or another because it got in their way or didn't play by their rules.

It seems simple to me. Anthropic is a private company. If you don't like it fix the laws or don't buy from them.

bitexploder··on Oracle on the hook to pay data centre investors even if site has no electricity
Ellison has more or less lived his entire life running Oracle (and his personal finances) completely over-leveraged, skating just barely ahead of disaster. I don't see this ending any way other than SV picking apart its bones for pennies on the dollar.
bitexploder··on Best LLM for every budget, updated daily
I load OpenRouter up and use models like GLM Flash 5.3, DeepSeek Flash 4.1, Luna, etc. And I often have random niche needs where I need a handful of calls for say, a really good image reader like Gemini Flash 3.8 or whatever. You can do a lot with $25 on openrouter or direct to chinese providers. I am cautious about what data I send overseas, but also like... just because it is in China does not inherently mean it is any less secure than a US provider.

I can't remember the last time any real recourse has mattered for companies getting breached or mishandling my data. Their stock just goes up and the govt just shrugs.

bitexploder··on Best LLM for every budget, updated daily
There also exists a $600-800 GPU that can run Qwen 27B 3.8 @ like 60 t/s for around 200W of energy.

I find Qwen Flash Next quite competent as well. 27B is a solid worker like you said. If you batch work and let them crank they do remarkably well. I am working on.

I started using Herdr and taught my agents to use it. So I use OMP loop and or goal, and it has a review cycles to wake up an Opus or Sol reviewer to make sure nothing is going off the rails with a local qwen flash next coordinating for me. I kinda prefer Sol, it seems like a more patient and thorough model, especially Sol 6, but Opus 5.5 is really good and its voice and attitude is not as grating as Opus 5 for sure.

I have a few V100 GPU running Qwen 27B and they do all the work overnight. Not quite the same speeds you have yet, but this is V100 machine and an old gaming machine with 16GB 4080 and a handful of 32GB V100s... all in less than 3K (ignoring that my gaming machine is 3 years old, but runs qwen flash next for free now as I game not a lot) for my little "we have AI at home" projects and there is a lot of interest in these old GPU now because they are rolling out of data centers now.

For my local work and personal projects... they just seem to be getting done in this setup. Every few days I sit down and do a big cycle with astra/fable/opus batch things up. I have projects that are basically "i want to see what happens" to "I want this to be good, I understand the code". Some of the throwaway projects that have just kind of magically finished more or less how I wanted have been great.

bitexploder··on The darker side of being a doctor
Whatever you want it to be. You can change it any time you want.
bitexploder··on The darker side of being a doctor
I accounted for that when I modeled it. I explicitly added investment start and compounding. A thing to remember: a physician is delayed yes but their higher income lets them much more heavily invest as their income often exceeds life expenses by a lot more than median swe.
bitexploder··on The darker side of being a doctor
Physician lifetime earnings counting for delay, and debt… normal investment, etc: 8.5-10m. Software engineer median is much lower. About 6.5m. Specialist and high end varies greatly. Median doctor is earning 2x a SWE (130k vs 250k).

People on HN vastly overestimate SWE pay as an industry, biased by FAANG as we are :).

bitexploder··on Claude Opus 5.5
I guarantee they are at least tweaking quants, caching systems, and finding ways to move serving costs down. This definitely impacts the model’s intelligence at times. There are also a lot of model tweaks, RLHF rollouts etc. I don’t think it means they are doing anything malicious or deceptive. And if Opus 5.5 is literally smarter than Fable? Ehh, it is plausible :)
bitexploder··on Claude Opus 5.5
What if the recent Fable intelligence regression was basically just them serving Opus 5.5 until they got it working well?
bitexploder··on Fable 5 – Median thinking declined in August
They obviously test various quants and other serving cost saving strategies. Models like Fable are probably trillions parameters with hundreds of billions active MoE. They probably try to squeeze and quant each piece until people notice.
Page 1 of 34Next →