HNHacker News
TopNewBestAskShowJobs

hugmynutus

69 karma · joined April 18, 2025

submissionscomments
hugmynutus··on New Mac Studio with M5 Max and M5 Ultra
Apple puts the NAND flash controller on package, so it cooled as part of the SoC/CPU as they share an IHS. Given it is part of the fabric it only has to maintain signal coherence on millimeter to micrometer scales, while enterprise/consumer NAND PCIe have to adhere to standards which force them to signal at 800-1200mV and keep that signal stable for centimeter scale distances.

Basically, Apple gets to cheat because they shove everyone onto the same IC & make the OS that runs on this SOC.

hugmynutus··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
Qwen3.8/Qwen3.6 has a weird self doubt/thinking too much problem. You can prompt it away. I would say it "approximates" Opus 4.X class models well enough especially for coding/linux problems.

The only reason I stopped using it as much is I was getting 25-35tok/s on Intel B70 (non-quant) which made some responses slow. For a long running/autonomous task, it would probably be sufficient.

hugmynutus··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
If they did not use my data for training, they should permit me to run the first 1-3 layers of the model locally and send them dense hidden state vectors. In my experience these compress very nicely without much effort.
hugmynutus··on DeepSeek V4 Pro 0813
Nullius in verba
hugmynutus··on Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
Because nobody knows.

Apple doesn't let you "pass" the GPU through to a VM like most other ARM/x86_64 processors (forwarding interrupts and PCIe memory regions). There are symbols defined to do this within the kernel (if you dump the binary) but they aren't used in retail macos.

Instead you end up creating a paravirtual device that emulates the GPU acting like a 'normal PCI device' which you give to clients. This is usually reserved (by other hardware vendors) for when you're doing multi-tenat time sharing of higher end GPUs (like Nvidia enterprise cards can do).

These paravirtualized GPUs then just have 'less features' and Apple (being Apple) states no reason why.

hugmynutus··on The Nixpkgs core team has disbanded
> I’m curious if you have an opinion on how you think this should be handled?

To "totally overcomplicate things" but do it correctly

- Binary states a list of constraints (namespace:name [<|>|>=|<=|!=] semver). - ldconf/ld.so either integrate into your package manager and/or are easier to update (I'm not writing conf files by hand and/or flakes). I should simply be able to recursively scan. - Give the runtime linker an SMT constraint solver (when <1000 this is nearly instant) - Cache known states to avoid solving NP hard problems every time you launch `cat`.

> At this point I’m just on team static linking or Windows DLL black box style. The Linux approach of dynamic libraries which function same as static is just the worst of every world.

Honestly same.

Having the package manager <-> elf runtime <-> shared libraries more-or-less be a blackbox is probably for the best (which is sort of what windows does with the install-shield/install-wizard stuff). But nobody in Linux Land really wants to "improve" userland, other then change to flavor-of-the-month display managers.

> I also blame C/C++ toolchains for being very very bad on Linux.

They honestly aren't, it is more your pacakge manager *is* your library manager. Because of the absolute bullshit of shared libraries.

If you pretend it is the 1970/80s everyone at your company has the architecture, unix version, etc. It is pretty nice. No cross compiling, multiple OSs, everything just sort of works pretty well. But like I said, "pretend".

hugmynutus··on The Nixpkgs core team has disbanded
Shared Dynamic Libraries. Not to say they aren't useful. Keeping applications small, keeping security updates simple, sub linear ram scaling, yada-yada-yada, big wins cross the board, not debating that.

The pervasive link detection structure of effectively everything on GNU/Linux assumes is (more-or-less) objectively incorrect behavior in any sane security minded context. Binaries should be able to declare the interface/contract they expect (args/types/abi/exceptions/cryptographic signatures/digests). Then the runtime linker "match" against the local system. The current system is basically 2 levels of string equality checking.

Nix goes into the right direction by getting your runtime linker/elf-runtime & package manager "integrated". But this sort of just feels like putting 'lipstick on a pig' and dancing around the core issue that `foo.exe` cannot ever realize `foobar-v1.8` and `foobar-v1.7` are both installed on the same computer. Nix just manages the environment/symlinks such that `foo.exe` doesn't realize this fact.

hugmynutus··on AMD acquires Taalas to boost inference performance by etching models in silicon
HN is rightly pointing out putting a model into an ASIC is kind of dumb.

HN is failing to understand that AMD knows this well.

Taalas has WO2025217724A1 pending and AMD wants that because it is immediately a function block they can sell to anyone doing FP math, since large (mostly) read only memory banks are ideally suited for that micro-code type stuff.

hugmynutus··on Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge
yellow/red tint is an extremely common problem not matter the photograph source you train on

source: work at a photograph start, even training on raw images things get tinted, it is an uphill battle

hugmynutus··on Claude Code uses Bun written in Rust now
> I don't see what should ever be wrong with a function pointer. [...]Care to explain what's the issue here?

1. You're writing code you don't have to

2. That adds runtime overhead

3. That when you screw up has non-trivial security & resource management side effects

This is objectively indefeasible in nearly any vaguely professional context.

hugmynutus··on Claude Code uses Bun written in Rust now
This reads like cope because you're re-inventing RAII from first principles.

I cannot take this seriously as tutorials on robust Zig Allocation Pools will store a deinit method for each item within the pool, so when the pool deinits, all internal objects can be deinit'd.

That is just RAII & dtors from first principles, except with extra overhead of manually storing fat pointers yourself (and the bugs that come with this). Instead of using a language with builtin guarantees & optimizations around handling this so your object pools don't need to carry around a bunch of function pointers. C++ has aggressive de-virtualization passes so at runtime a lot of the 'complex object hierarchies' can be flattened to purely static function calls.

hugmynutus··on Qwen 3.8
> It's hard to say what their motivation is.

Involution is a major problem in Chinese industries [1]. Where companies will sell their products at a loss, effectively playing fiscal chicken [2] with one another to dominate a market. It is such an issue the government has had to step in to prevent EV companies from destroying themselves by more-or-less requiring companies sell their goods at a profit [3].

The straight forward line of reasoning that AI/LLM labs are applying this logic to their profit.

I think (we) Americans are reading a bit too far into this assuming government intervention, conspiracy, etc.. Chinese markets are downright cut throat. They're using those tactics to compete with US labs.

1. https://www.reuters.com/business/autos-transportation/what-i... 2. https://en.wikipedia.org/wiki/Chicken_(game) 3. https://www.theguardian.com/business/2025/aug/05/china-warns...

hugmynutus··on AMD silently removes memory encryption from consumer Ryzen CPUs
Everyone jumping up about "enshitification". I tried to enable this feature on QEMU and it broke my VMs because the secure memory system was board-line hopelessly broken/non-functional.

Did anyone even use this feature?

Yes it is dishonest to remove features but from perspective AMD disabled a feature that never worked in the first place. The feature never should've been advertised as enabled.

hugmynutus··on Twenty One Zero-Days in FFmpeg
GStreamer is just a different front end to ffmpeg.

ffmpeg's core functionality (encode, decode, streams, pipes, channels) are all implemented in `libav` which gstreamer links against.

hugmynutus··on How JPL keeps the 13-year-old Curiosity rover doing science
Add to the list that martian dust contains a massive amount of carcinogens meaning any air/dust lock has to be an ISO-6 clean room.

Sure it is a "solved" problem but all the solutions are very heavy.

hugmynutus··on Running local models on an M4 with 24GB memory
It makes no difference at all.

Edit: Opus prior to the context nerf it worked more often than not. Current Opus 4.7 is practically unusable.

hugmynutus··on Running local models on an M4 with 24GB memory
> What it gives me in Swift, most closely resembles stuff that enthusiastic newer folks would do, and want to show off.

The same is true for rust-lang. Code that will immediately clone/re-allocate anything passed by reference and collect everything to the heap that is passed by `Iterator`/`IntoIterator`.

It is a massive performance anti-pattern and the hallmark of somebody "struggling" with the borrow checker. Naturally a lot of 1st & 2nd 'I just learned rust' projects lean on it. Which is totally fine for humans, you're learning. But with LLMs that pattern is now burned into their eigenvectors with the heat of a billion hours of H100 training time.

It has gotten to a point that all code I generate with Opus or Codex if there as iterator or reference in the argument, I start a fresh context, with a sort of `remove unnecessary clones, collections, and copies from the following code: {{code}}`

hugmynutus··on Shall I implement it? No
This is because LLMs don't actually understand language, they're just a "which word fragment comes next machine".

    Instruction: don't think about ${term}
Now `${term}` is in the LLMs context window. Then the attention system will amply the logits related to `${term}` based on how often `${term}` appeared in chat. This is just how text gets transformed into numbers for the LLM to process. Relational structure of transformers will similarly amplify tokens related to `${term}` single that is what training is about, you said `fruit`, so `apple`, `orange`, `pear`, etc. all become more likely to get spat out.

The negation of a term (do not under any circumstances do X) generally does not work unless they've received extensive training & fining tuning to ensure a specific "Do not generate X" will influence every single down stream weight (multiple times), which they often do for writing style & specific (illegal) terms. So for drafting emails or chatting, works fine.

But when you start getting into advanced technical concepts & profession specific jargon, not at all.

hugmynutus··on What does " 2>&1 " mean?
> Or is it restricted to 0/1/2 by the shell?

It is not. You can use any arbitrary numbers provided they're initialized properly. These values are just file descriptors.

For Example -> https://gist.github.com/valarauca/71b99af82ccbb156e0601c5df8...

I've used (see: example) to handle applications that just dump pointless noise into stdout/stderr, which is only useful when the binary crashes/fails. Provided the error is marked by a non-zero return code, this will then correctly display the stdout/stderr (provided there is <64KiB of it).

hugmynutus··on Anthropic officially bans using subscription auth for third party use
> Commodity hardware and software will continue to drop in price.

The software is free (citation: Cuda, nvcc, llvm, olama/llama cpp, linux, etc)

The hardware is *not* getting cheaper (unless we're talking a 5+ year time) as most manufacturers are signaling the current shortages will continue ~24 months.

hugmynutus··on Gathering Linux Syscall Numbers in a C Table
docs.kernel.org is generated from in tree readmes, docs, type/struct/function definitions. Making it a lot easier to read/browse documentation that would (previously) require grepping the source code to find.

I realize the site also hosts some fairly out-of-date articles, there is room for improvement. Those hand written articles start with an author & timestamp, so they're easy to filter.

hugmynutus··on Linux is good now
Setting up `apt` to pull from a different repo (to say install firefox.dpkg instead of snap) requires like 3-4 commands which are easily searchable.

I'd had effectively zero issues avoid snaps.

hugmynutus··on A Fast 64-Bit Date Algorithm (30–40% faster by counting dates backwards)
> i'm placing my bets that in a few thousand years we'll have changed calendar system entirely haha

Given the chronostrife will occur in around 40_000 years (give or take 2_000) I somewhat doubt that </humor>

hugmynutus··on Uv is the best thing to happen to the Python ecosystem in a decade
> As long as your

linux core utils have supported this since 2018 (coreutils 8.3), amusingly it is the same release that added `cp --reflink`. AFAIK I know you have to opt out by having `POSIX_CORRECT=1` or `POSIX_ME_HARDER=1` or `--pedantic` set in your environment. [1]

freebsd core utils have supported this since 2008

MacOS has basically always supported this.

---

1. Amusingly despite `POSIX_ME_HARDER` not being official a alrge swapt of core utils support it. https://www.gnu.org/prep/standards/html_node/Non_002dGNU-Sta...

hugmynutus··on How to save the world with ZFS and 12 USB sticks: 4th anniversary video (2011)
Buddy, I have 24Tb HDDs in my pool today.

If anything the opposite has occurred. HDD scaling has largely flattened. Going from 1986 -> 2014, HDD size increased by 10x every 5.3 years [1]. If anything we should have 100Tb+ drives if scaling kept going. I say this not as a but there have been directly implications for ZFS.

All this data stuck behind an interface who's speed is (realistically after a file system & kernel involved) hard limited to 200MiB/s-300MiB/s. Recovery times sky rocket. As you simply cannot re-build parity/copy data. The whole reason stuff like draid [2] were created is so larger pools can recover in less than a day by doing sequential parity & hot-spairs loaded 1/N of each drives data ahead of time.

---

1. Not the most reliable source, but it is a friday afternoon https://old.reddit.com/r/DataHoarder/comments/spoek4/hdd_cap...

2. https://openzfs.github.io/openzfs-docs/Basic%20Concepts/dRAI... for concept, for motivations & implementation details see -> https://www.youtube.com/watch?v=xPU3rIHyCTs

hugmynutus··on A cheat sheet for why using ChatGPT is not bad for the environment
You are correct to point out the larger questions of supply chain cost (and their environmental impact) are not addressed in the root link.
hugmynutus··on A cheat sheet for why using ChatGPT is not bad for the environment
Thanks!
hugmynutus··on A cheat sheet for why using ChatGPT is not bad for the environment
I find this unconvincing. The actual discussion of LLM generation is very lacking.

The original link [1] cites a discussion of the cost per query of GPT-4o at 0.3whr [2]. When you read the document [2] itself you see 0.3whr is a lower bound & 40whr is the upper bound. The paper [2] is actually pretty solid, I recommend it. It uses the public metrics from other LLM APIs to derive a likely distribution of the context size of the average query for GPT-4o which is a reasonable approach given that data isn't public. Then factoring in GPU power per FLOP, average utilization during, and cloud/renting overhead. It admits this likely has non-trivial error bars, concluding the average is between 1-4whr per query.

This is disappointing to me as the original link [1] attempts to bring in this source [2] to disprove the 3whr "myth" created by another paper [3], yet this 3whr figure lies directly in the error bars their new source [2] arrives at.

Links:

1. https://simonwillison.net/2025/Apr/29/chatgpt-is-not-bad-for...

2. https://epoch.ai/gradient-updates/how-much-energy-does-chatg...

3. https://www.sciencedirect.com/science/article/pii/S254243512...

Edit: whr not w/hr

hugmynutus··on The Policy Puppetry Attack: Novel bypass for major LLMs
This really just a variant of the classic, "pretend you're somebody else, reply as {{char}}" which has been around for 4+ years and despite the age, continues to be somewhat effective.

Modern skeleton key attacks are far more effective.

hugmynutus··on The Size of Packets
Luckily 400Gb/s nics are already on the market [1]

[1] https://docs.broadcom.com/doc/957608-PB1