HNHacker News
TopNewBestAskShowJobs

DiabloD3

47,548 karma · joined September 12, 2010

Blog: http://adterrasperaspera.com/ Email: diablod3@gmail.com
submissionscomments
DiabloD3··on Everything I own, owned
They can.... which everyone bitched about: they would do their cleaning cycle, no matter if you were busy or not.
DiabloD3··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
It does depending on the technique.
DiabloD3··on Hook, hold, harvest and hide: Meta's alleged strategy laid out in first week
Its the same way "embrace, extend, and extinguish" was coined. "Embrace and extend", without extinguish/exterminate, was used in Microsoft corp docs that were handed to the DOJ during the famous antitrust case, but it was the DOJ that turned it into EEE after realizing what "embrace and extend" actually meant in the broader context of Microsoft's machinations.

The DOJ's clever hook won them the case.

DiabloD3··on Qwen 3.8 27B
I would not expect Ollama to be doing the right thing fwiw.
DiabloD3··on Anthropic investors bet on $2T valuation in record IPO
What a weird way for Anthropic to announce they're going out of business.
DiabloD3··on Microsoft Plugs Nearly 400 Security Holes
I think a lot of software engineers at Microsoft need to either be retrained or just let go entirely, and maybe so does a lot (or probably most) of the management.

Why are you still using languages that make it trivial to write most kinds of security bugs? Microsoft has written more than one language, they have/had in-house talent on how to do this, and one of those languages is C#, which bootstrapped itself after the whole Sun v. Microsoft case (the one where Microsoft tried to EEE Java, and Sun yanked their technology license; Microsoft then just removed Java(tm) and made their own replacement dialect on-top of their homegrown VM)...

But more importantly, for the things that can't, or shouldn't, be written in C#... they have hired a lot of ex-Mozilla Rust developers (the ones that left during Mitchell Baker's disastrous attempt to fire "rockstar developers") to get that ball rolling on that.

Why hasn't a majority of the important code been rewritten in Rust yet? Not only that, Microsoft/Github advertises that LLMs are really super good at doing language-changing rewrites, so why aren't they dog fooding this and/or are they admitting that LLMs can't do this?

DiabloD3··on STV: A full-motion video codec for the Atari ST
H.261 is from 1988.

JPEG is from 1992.

DiabloD3··on FFmpeg 9.0
They can't, unfortunately.

Linux's drivers makes that happen, not ffmpeg; ffmpeg merely calls the API.

Intel's own first party drivers simply follow the ACPI tables, Linux ignores them.

DiabloD3··on Handbook.md shows that long policy documents do not reliably govern agents
See the end of https://news.ycombinator.com/item?id=49100144
DiabloD3··on Disrupting supply chain attacks on NPM and GitHub Actions
It very much is actionable.

You might not know this, but the world of software development existed before NPM and Github, and it continues to exist after them. You are not beholden to either of them.

DiabloD3··on Handbook.md shows that long policy documents do not reliably govern agents
Tell me where I can buy a car for $1k or $2k.
DiabloD3··on Handbook.md shows that long policy documents do not reliably govern agents
Every single analysis I've seen done by anyone has basically come to the same conclusion: assuming sane BF16 or Q8 model quant, and BF16 or Q8 KV context quant (ie, not intentionally screwing over the model), you get about 250k before it pukes.
DiabloD3··on Handbook.md shows that long policy documents do not reliably govern agents
One of the biggest fixes I've seen is just getting rid of traditional sampling. But first, let me say something about quantization, just to get this out of the way.

Like, lets say you already did the sane thing, your model[1] is already either FP16 or Q8 (and quantized by a competent practitioner of the art, ex: unsloth or bartowski), and your KV context is already FP16 or Q8... which means you now are already ahead of the major companies.

Google, OpenAI, and Anthropic heavily compress both K and V to insane levels, which might not appear to be so bad on short prompts, but especially with thinking enabled, its sort of the equivalent of JPEGing a JPEG repeatedly. Every time the model thinks, and records its thoughts into the context, and then reads it back later to think more, it becomes further and further imprecise.

All models with heavy KV context compression go off the rails somewhere between a quarter and a half of a million tokens. Every. Single. One. Every team that releases a model that has a limit of a quarter of a million did this on purpose, and it was the smart thing to do.

Now, lets say you dip your toes into samplers; this includes stuff like temp, top k, top p, min p, etc. I won't describe what they do, there are already good ELI5 articles out there to help you with that. They are, however, the original samplers, and the only ones the big companies use. None of them use the newer samplers that massively outperform them.

You know what you get with most providers? Temp as a knob, and it only goes between like 0.0 and 1.5. What if you want higher? Nope! What if you want to tune the other knobs? Usually no, too (OpenAI seems to still offer it on their higher end API plans, but Google has eliminated all knobs, and Anthropic apparently removing everything but temp in the future). What if you want other samplers? Not allowed.

Even restricting yourself to normal samplers, what if I wanted temp of 100, and min_p of 0.9 and no other samplers? That produces sane results for creative writing tasks, yet I could never do this, even I was paying for some $200/mo plan at Anthropic/OpenAI/Google/etc.

All of these samplers also cause a sort of JPEGing a JPEG repeatedly sort of error, it is the third source of it (model quant and KV cache quant are the other two). Long context insanity is probably caused more by sampling error more than it does by model and KV quant error.

What other samplers are there? Llama.cpp impls dynatemp, mirostat, top-n-sigma, and some others.

The one that I think more people need to look at is top-n-sigma. Temp and top-n-sigma alone has produced results that, even on ridiculously complex and purposefully tricky prompts, let models that have 1M native context happily go to the very limit of it without any signs of tell-tale degradation.

Want to go have fun with your new found freedom? Get Qwen 3.6 27B in Q4_K_M from a reputable dealer, set llama.cpp to do Q4 KV (yep, after I just said don't do that), and then run it with `--samplers "temperature;top_n_sigma" --top-n-sigma 1.0 --temp 1000`, and then compare it to the normal recommended values of `--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00`.

You will find that a lot of problems suddenly go away: long context degradation vanishes (which Qwen 3.6 will still do inside it's 250k context, especially when both model and KV are below Q8), forgetting what you said or it said earlier goes away, losing the plot half way through goes away, overreliance on cliches (Qwen has its own form of Claudisms, but are more subtle and less grating) also goes away, hallucinations happen far less, and lazyness also goes away (Qwen 3.6 has no real lazyness defects, but Gemma 4 does).

I have tested those top-n sigma settings on code generation as well. It is better than stock, and better than the commercial offerings by any of the American providers, but I still don't think LLMs can replace human programmers: it still can't think nor reason.

But yeah, moving to more modern samplers has done more to unfuck LLMs IMO than anything else you can do.

[1]: Assuming it isn't a native 4 bit model of some kind; quantizing them correctly to work in a local inference engine without actually quantizing anything is a bit of a PITA. See unsloth's work on the QAT Gemma 4 releases.

DiabloD3··on Handbook.md shows that long policy documents do not reliably govern agents
This is a problem with long context models. To put it as simple and as bluntly as possible: just because they claim you can use 1M tokens in your context doesn't mean its true and you should do that.

Due to extreme quantization of models and the context's KV cache, and also just really shitty samplers provided to the user (hell, most are just getting rid of sampler knobs altogether), this problem will absolutely continue.

Want it to go away, almost like magic? Local inference. When its under your control, and no longer being forced to hold it wrong, all of the common LLM defects will go away.

DiabloD3··on LLM Usage in Debian: Three Proposals
That is called a hallucination.
DiabloD3··on Intel Starts Shipping High-NA EUV Silicon
I don't understand the joke.

Intel has already stated their deal with Nvidia, this came out in December.

If you think there will be shipping Druid products, I guess you can find out sometime in 2028.

DiabloD3··on AMD Ryzen 7 7700X3D Review: 3D V-Cache Gaming Performance for Less
What GPU? Its usually Nvidia owners that complain about lackluster Wine/Proton performance.
DiabloD3··on Intel Starts Shipping High-NA EUV Silicon
Battlemage was still while Pat Gelsinger was CEO, and he was the only real champion of that product, understanding its place in Intel's overall strategy.

He doesn't work there anymore, and neither does a significant fraction of the GPU engineering team.

DiabloD3··on Intel Starts Shipping High-NA EUV Silicon
There will be no future Xe GPUs.

Druid, the 4th generation, has been shelved entirely, and all consumer DGPU products, possibly _all_ DGPU products, have been killed for Celestial.

Its likely the only products shipping with Celestial will be IGPUs for the next generation and a half, until Serpent Lake comes out with RTX graphics tiles from Nvidia, and then never another Xe product ever again.

DiabloD3··on AMD Ryzen 7 7700X3D Review: 3D V-Cache Gaming Performance for Less
A lot of games no longer run properly in Windows, and the only way to keep playing them is with Wine/Proton.

Windows performance is also often worse than Wine/Proton's.

DiabloD3··on Nobody knows what a used GPU cluster is worth
GPU clusters have, largely, no actual value.

If anything, you might have to pay to have them disposed of, they don't really have any meaningful used eBay market outside of the randos that want to do high end extreme local inference in their basement.

Also, as for RAMmageddon, the inference SBCs that all of the AI bros bought don't have DIMMs, they're not even the right chip: its all GDDR and LPDDR. The only DDR DIMMs being consumed are for regular non-inference machines that help run the business and service infrastructure behind the scenes.

DiabloD3··on Show HN: Fast NF4 dequantization Triton kernel (1.41x faster than bitsandbytes)
Neat, can't wait to see a llama.cpp PR for this.
DiabloD3··on OnePlus to Pull Out of American and European Markets
OnePlus basically already went under a few years back.

They merged with Oppo, and ever since its just been rebranded Oppo phones... they're fine, they work, but they're not the magic OnePlus was.

DiabloD3··on OnePlus to Pull Out of American and European Markets
Skip Samsung, they tend to not honor warranties and are ROM swap resistant.

Google is okay for now (I recently got a Pixel 10, its fine, it can run Graphene).

Next year Motorola flagships will also have official Graphene support.

DiabloD3··on Alternative(s) to run CUDA on non-Nvidia hardware
Ironically, this is what people claim AI can do with a snap of the fingers.

Should be real simple if the HN AI echochamber is right, right?

DiabloD3··on Alternative(s) to run CUDA on non-Nvidia hardware
I love how people say things like "extension spaghetti", as if all other non-standard APIs have the same problem: hardware gets new features that people want to use from that API, API gains extension to use that hardware feature.

CUDA is no different, in fact, often worse. Nvidia is bad at documenting which hardware does what things, and CUDA users often have to use third party tables to figure out what hardware can't do what and disappoint customers who unwisely invested into it.

DiabloD3··on Alternative(s) to run CUDA on non-Nvidia hardware
Weird, most people have the exact opposite experience.

Having to deal with closed source opaque poorly documented stacks sucks.

DiabloD3··on Alternative(s) to run CUDA on non-Nvidia hardware
Weird, since the most used open source inference engine is faster on Vulkan on platforms that offer multiple options, with the sole exception being Nvidia, due to poor Nvidia driver quality (which I am forced to assume is intentional, Nvidia wishes to maintain their moat after all).
DiabloD3··on Alternative(s) to run CUDA on non-Nvidia hardware
Its easier to just get rid of your legacy code entirely and use Vulkan for compute, or have your compiler emit SPIR-V directly.

No reason to tie yourself to Nvidia's moat.

DiabloD3··on GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]
You'd have less problems with 27B, btw.
← PreviousPage 2 of 34Next →