HNHacker News
TopNewBestAskShowJobs

DiabloD3

47,547 karma · joined September 12, 2010

Blog: http://adterrasperaspera.com/ Email: diablod3@gmail.com
submissionscomments
DiabloD3··on Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
The title of the paper is correct. The paper does not seem to actually get to the point in a generic way, but hyperfocuses on, effectively, one type of error compensation.

Highly quantized models, especially with highly quantized KV caches, will, effectively, attend to the wrong tokens and be unable to easily discern highly similar tokens. The bastardized way of explaining this is gradient descent techniques get stuck in localized minimum and global maximums, so what happens when you turn the slopes into hard stair steps?

We need to move to smaller models and smaller caches and better samplers, not new quant methods (although I'm willing to also take those too).

DiabloD3··on Pentium II at 600Mhz with Voodoo 3 Emulated on 86Box with M6 Mac Mini
Classic is the version that was ported to the Goldsrc engine (which, itself, is a descendant of the Quake 1 engine) as a stand alone game.

TF2 is the same concept as TF/TFC, but turned into a full fledged game with its own soul.

DiabloD3··on Postgres SELECT DISTINCT Does Not Scale
The article seems to have changed since I commented.
DiabloD3··on Postgres SELECT DISTINCT Does Not Scale
The article seems to have changed since I commented.
DiabloD3··on Postgres SELECT DISTINCT Does Not Scale
"Postgres SELECT DISTINCT Does Not Scale"

Correct. This is documented in depth: DISTINCT sorts the results first.

The article's use case seems to imply the author did not know about GROUP BY, nor does it imply the author knew about indexes, nor ANALYZE. Postgres 18's new skip scan indexing also could help here, so ensuring the planner chooses that could help.

DiabloD3··on Show HN: Koi.rest – watch some fish and regain your balance
You didn't make anything at all, though.

Also, please talk to a lawyer before slapping a Copyright on anything you didn't make. Machine-made products cannot be Copyrighted in the US and most Berne Convention countries, and you cannot own the output of an LLM, nor can you legally shield yourself from the consequences.

DiabloD3··on Reddit mod ordered to pay Nintendo $4.5M in Switch piracy lawsuit
Ahh, The Verge, always causing problems where none exist.

The guy was a brazen pirate and operated a ToS-violating subreddit about Switch piracy. He was sued for his actions, and lost.

This is like writing a headline "McDonalds employee ordered to pay", or "Amazon employee", or "Walmart employee", or "Guy who uses Linux", or "Good upstanding Christian", or whatever unrelated thing you want to attach here.

Like, I'm not pro-Nintendo nor pro-Reddit here, both are toxic reprehensible companies that harm users and partners alike, but the guy did what he did and had his day in court.

DiabloD3··on Disney+ changes subscriber agreement to allow ads on every plan
Thats a weird way for Disney to announce the end of Disney+.
DiabloD3··on Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
Surprised its not meaningfully slower.

Vulkan and ROCm paths are missing a few optimized versions of the quants they're using.

DiabloD3··on I had Gemini train its own replacement for $9
I agree, this is misleading.

If he wanted a Gemini replacement verbatim, its called locally inferring it's sibling, Gemma.

DiabloD3··on Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
That's the other way of doing it, which solves the context rot problem in a more complex way. The model at the top says, "hey, sub-agent, go figure out the answer to this question and give me the answer", and that sub-agent can go consume 250k+ context to return an answer that might be a couple of words, and thus not contaminate the main context with that now thrown-away context.

However, this is not something that is inherently part of models or inference engine, but part of the harness.

Harnesses are very hit and miss, and are not integrated into the stack, and I think that will have to happen eventually. Like, conceptually similar to an LLM performing a tool call that just calls itself recursively, I think this would go a long way to making LLMs more viable for being an actual product people could conceivably want.

DiabloD3··on Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
A lot of this is managed by the inference engine, and has nothing to do with the model.

Models that use, for example, sparse attention mechanisms are just trying to make the bad situation slightly less bad, such as using less RAM for context (thus requiring less context quantization) or using less bandwidth (thus running faster).

If people keep using temp, top-k, top-p, and min-p, and nothing else for samplers, we're ignoring ~3 years of sampling research that virtually eliminates the worst of context rot issues.

DiabloD3··on Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
That usually ends up being a poor use of LLMs, and is an unsolved problem with LLMs.

RAG was supposed to be the way out on that, and ended up being mostly abandoned.

DiabloD3··on Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
The article doesn't really describe the problem: if your prompt is 35kb, your prompt is confusing, unfocused, and doesn't work right on any LLM, and is needlessly bloating your context.

At this point in time, due to how most people and companies run their inference engine, regardless of the model (yes, this includes the newest from OpenAI and Anthropic and the Chinese Tigers and Dragons), you run out of useful context that the model can accurately attend to around the 250k mark no matter how much they advertise their context size is.

You need to cut your prompt up. If you believe LLMs work, have the LLM help you shape the overall plan, and then have multiple sessions run each step in the plan without being bloated with the context of previous successful steps.

I don't see LLMs being production-ready until the context rot and sampling problem is fixed forever. This has not occurred, and the big inference providers aren't even bothering to integrate any of the research on that subject.

If anything, many of the bigger companies are actively making inference quality worse just to extend their runway a tiny bit farther before they go bankrupt.

The only thing the article gets right is this: if you're serious about LLMs, abandon Big AI and infer locally only. This is the only way you have control over the quality of the output.

DiabloD3··on PlayStation cancels Kojima's PHYSINT, Xbox steps in
No, Guerrilla Games. The engine also already runs on virtually every platform, including the Xbox and PC, but also phones.

Even though Guerilla Games is currently owned by Playstation Studios, they are still effectively an independent studio.

Microsoft, otoh, owns like half a dozen engines, uses none of them, and licenses none of them. The only way I'd see a engine swap happening is if Kojima somehow swung a deal to license the newest version of the Doom engine.

DiabloD3··on Trezor's email provider has been breached
Trent Reznor on good email providers: "I just want something I can never have" (maybe)
DiabloD3··on ChatGPT Was Built on Concealed 'Mass Piracy', Authors Tell Court
Those aren't equal, however.

Aaron Swartz was murdered by a company that prints scientific journals that paywall papers paid for with US government grants; his "crime" was using the access granted to him by MIT, legally. At no point did he break any laws nor the license granted to him by the paywall service.

OpenAI, Anthropic, and Meta all _knowingly_ pirated Copyrighted works en masse, knowing they did not have a license to do so, and then distributed those works, again, knowing they did not have a license to do so.

I do not think anyone should be murdered by the state, but the difference between Aaron Swartz and Sam Altman is one of them did it to fix a great injustice, the other did it as a fly-by-night get rich scheme.

Notice the one that did things legally got murdered, while the criminal continues to walk the streets.

DiabloD3··on 1% increase in immigration leads to a 1.78%-2.97% increase in Far-Right voting (2024)
far right == 1930s and 40s right, from what I can tell.
DiabloD3··on Juul gets OK to sell vaping device with age-gating technology
Can't wait until Trump is out of office and we can just ban this stuff.

Kind of insane harmless things like THC are Schedule I, but Nicotine isn't.

DiabloD3··on [dead]
Can we just start flagging all AI written slop and have Dan moderate it?

(Sorry Dan, but this needs to change)

DiabloD3··on Nvidia agrees to acquire Hugging Face for $13B
Heh, thats a weird way to announced the death of HuggingFace.
DiabloD3··on Hot Chips 2026: CUDA Targets RISC-V – By Chester Lam
Don't really know why this is interesting, CUDA is kind of a niche legacy API that Nvidia has kept on life support a little too long.

Just use Vulkan for compute.

DiabloD3··on Everything I own, owned
They do run it when its off as well.
DiabloD3··on Everything I own, owned
Then turn off the warning: all modern OLED monitors that do not force you to run it can both change when the timer fires, but also just disable the notification altogether.

This is what I have done on mine, and mine as well also runs it when its off.

If your monitor is a Samsung OLED panel, yours is a near identical sibling of mine.

DiabloD3··on Everything I own, owned
And thats why we have the warning that everyone bitches about: "Hey, the timer is up, but I'm not gonna make you do it, its just gonna wear the display out faster".

Early OLEDs really did need to be pixel cleaned every 8 hours according to manufacturer estimates, the choice isn't have warning or not, its have a lifespan or not.

I don't want to blame early adopters for being early adopters, but they early adopted, and this is the early adoption problem.

DiabloD3··on Everything I own, owned
They can.... which everyone bitched about: they would do their cleaning cycle, no matter if you were busy or not.
DiabloD3··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
It does depending on the technique.
DiabloD3··on Hook, hold, harvest and hide: Meta's alleged strategy laid out in first week
Its the same way "embrace, extend, and extinguish" was coined. "Embrace and extend", without extinguish/exterminate, was used in Microsoft corp docs that were handed to the DOJ during the famous antitrust case, but it was the DOJ that turned it into EEE after realizing what "embrace and extend" actually meant in the broader context of Microsoft's machinations.

The DOJ's clever hook won them the case.

DiabloD3··on Qwen 3.8 27B
I would not expect Ollama to be doing the right thing fwiw.
DiabloD3··on Anthropic investors bet on $2T valuation in record IPO
What a weird way for Anthropic to announce they're going out of business.
Page 1 of 34Next →