HNHacker News
TopNewBestAskShowJobs

Farmadupe

159 karma · joined January 21, 2021

submissionscomments
Farmadupe··on Btrfs/ZFS/bcachefs under workloads classic benchmarks skip
ah, yes that's obviously fine; I'm definitely coming fromt he perspective of "bored guy on a Saturday evening who happened to stumble across someone else's wall of charts on HN"
Farmadupe··on Btrfs/ZFS/bcachefs under workloads classic benchmarks skip
I'd actually recommend some excellent books on the "philosophy" of data presentation: The first that comes to mind is "the visual display of quantitative information" by Edward Tufte seems to be freely available online, and the other one on my mind is "how charts lie" by Alberto Cairo (which doesn't seem to be freely accessible)

But if it helps, just some "initial gut feel observations" from me:

* It's definitely not possible to find issue with the the _sheer amount_ of results, but there's just far too much for a human to absorb, all presented at once

* Overall text size is quite small, and difficult to read

* The page doesn't make a strong statement of _what_ is under test: the first words are: "modern-fs-benchmark Multi-device CoW filesystems under workloads classic benchmarks skip" -- which defines the webpage in terms of what it is _not_, without stating what benchmarks are actually present.

* The first line of teh page contains run statistics that probably eithre want to b at the bottom, or just don't need to be in the webpage at all: "latest run 2026-09-18 18:50:45 UTC, kernel 7.0.0-1012-azure, 593 runs recorded · 145 trend points shown"

* A significant proportion of the free text is caveats. There's nothing wrong with being transparent about limitations, but they may be a sign that there might be alternative ways to present the data, or that the data may be flawed (depending on the caveat)

* Theres several categories that I think have been invented for the purpose of collation, but I don't think are defined on the page. I think "Overall Core" and "Core I/O" aren't explained, which means by definition it's impossible for a reader to understand the score table.

* And as we're all aware right now, current Claude models are currently struggling to write coherent English. There's several incoherent sentences on the page. It's a Claude issue.

Farmadupe··on Btrfs/ZFS/bcachefs under workloads classic benchmarks skip
Yes exactly. The issueThe epherrality of the VMs isn't an issue, it's the _shared_ part that's the concern here. Going by the fact that the kernel is listed as "kernel 7.0.0-1012-azure" I feel like it's a fair risk that there may have been noisy neighbours.
Farmadupe··on Btrfs/ZFS/bcachefs under workloads classic benchmarks skip
> compare shapes and ratios, not absolute MB/s

In this case, given that the author's own disclaimer (above) already disclaims the numeric readings, I'm not sure how it's possible to make any inference on "shapes and ratios" derived from the numeric readings.

Farmadupe··on Btrfs/ZFS/bcachefs under workloads classic benchmarks skip
> CI runs use loop devices on shared ephemeral VMs (one VM per filesystem): compare shapes and ratios, not absolute MB/s. Each job records a host-calibration anchor — see the table.

I think if you're not using baremetal for such tests, it's likely that the results are simply not comparable at all? What if another tenant is also using the disk?

Farmadupe··on Btrfs/ZFS/bcachefs under workloads classic benchmarks skip
> 2G of random garbage is written directly onto one member device (behind the filesystem's back, offset 1G — python injector; uutils dd mis-seeks on dm devices), caches dropped, then a full scrub: btrfs scrub -B, zpool scrub + wait, bcachefs scrub, md/lvm sync-action 'check' (which can only COUNT mismatches — no checksums to know which copy is right).

I'm not sure that nuking 2G of the underlying block device is a recoverable error on any filesystem that I'm aware of? Can you confirm if any ofthe filesystems really came out of the other side in a usable state after scrubbing?

-----

> Trivial-op p99, idle (ms) # A trivial operation — one 4k write + fsync every 200ms (like a shell appending history or an editor updating its swap file) — run alone for 10s. p99 of the fsync completion

In fact, if it's OK for me to ask, are any of the metrics tht you used standard industry metrics? It looks like several of the tests are bypassing the kernel's page cache? -- which I worry may fall into the trap of "I modified the system to be unrepresentative of reality and then tested it".

----

> kernel 7.0.0-1012-azure

Can you confirm if you tested on a bare metal machine? were you the only tenant?

Farmadupe··on Btrfs/ZFS/bcachefs under workloads classic benchmarks skip
@farlight assuming that you're the creator do you think you'd be able to rework the HTML/CSS? I'm sure you've got good data but speaking on behalf of my eyeballs, the results page is... hard to read!
Farmadupe··on Don't be the out of touch Kung Fu master
Thankyou for putting the last three months of my life into words <3

> and the explanations of the agent are convincing enough that surely, it knows better than you.

I feel this in my bones. I also get to watch the misalignment feedback loop close itself when the next agent sees that security rules aren't followed because of a hallucinated 20 line justification in a code comment, and then it decides that the project _is_ a demo and then confidently writes even more security holes into the codebase.

Then when you catch the issue, the agent pushes back against the fix because it would need a schema change and production DB migration.

Farmadupe··on Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(If it helps, I ask my own question of myself too -- I mostly don't write code by hand any more as I find that an LLM writes it faster and with less bugs -- Is that therefore proof that my time was never worth my paycheck? I hope not but at the same time I would actually be proud if I had got away with being an accidental charlatan/fraudster at my employer's expense during my entire career)

-----

Similarly, if what I said really is true, I would be implying that LLMs are charlatan/fraudster detectors (to some statistical level). And I refuse on principle to believe that that is actually the case.

Farmadupe··on Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
hmm, assuming that this article is part written by claude and part human-written, can anyone help me find a rule of thumb for "how to know if the article is worth reading"?

Because on the one hand, the prose and the presentation is painful (narrating irrelevant points, nonlinear X-axes, ambiguous chart labels, etc etc),

But on the other hand, the result that I'm assuming the author means to communicate ("on these evals, generation quality seems fairly good") sounds worthwhile to share?

Because I really struggle with this question at the moment. Am I allowed to draw an adverse inference that "if the writeup presents irrelevant text side by side with the data, then this may be a sign that the author does not understand the task that they are attempting to write up"?

Farmadupe··on Qantas Airbus A380 engine failure in 2010 (2023)
Don't know why you're getting faded out, because having worked on exactly this engine, software mitigations were put in place. That's exactly your `anomaly detection system`

It was a belt-and-braces fix, because the primary fix was of course to manufacture the engines correctly in the first place.

The "lots of temperature sensors" thing is actually easier said than done; jet engines already have lots of temperature sensors, and they will turn the engine off if a huge overtemperature is detected, but it's not a free win to add more of them. It increases the risk of a faulty sensor shutting down a perfectly good engine in flight, which from a safety-critical perspective is a borderline disaster condition all by itself.

(technically you usually cross-correlate temperature sensor readings with pressure sensor readings to prevent sensor faults turning engines off, but still, extra sensors isn't always a free win.)

Farmadupe··on Firefox's AI Switch Is Off. Telemetry Isn't
Thankyou! that title was painful, I spent 20 seconds term rewriting it in my head..

* Firefox's AI Switch Is Off. Telemetry Isn't

* Firefox's AI Switch Is Off. Telemetry Isn't off [complete second clause]

* Firefox's Telemetry Isn't Off [remove negative clause]

* Firefox's Telemetry Is On [remove double negative]

----

Admittedly I didn't read past the first two paragraphs, but I assume it is a several hundred word repetition of its tautologous title.

Farmadupe··on NanoGPT Speedrun Frontier
Yeah, I'm with you on this, I think this is just what fable/opus-5 slop looks like now...

- "A frozen verify.py accepts the claim" (what does it mean to freeze a python script?)

- "which trains the recipe eight times on fixed seeds it can't touch" (what does it mean to not be able to touch a seed)

- "One other detail is that we gave an estimation of the speedrun noise in program.md that was slightly too large. 62 out of ~100 runs measured it themselves instead of trusting our number" (What does it mean for a "run" to "distrust" a noise measurement)

- "One important disclaimer is that our benchmark has a lot of variance" (Actually this one makes sense, but congratulations for burying the lede that your entire article is bogus.)

- "Almost every model finds the same winning ideas. What separates the best traces is what an experiment leaves behind" (This implies the the graph would show every model finding a plateau in whatever metric the experiment is measuring, but I don't see every model scoring the same in the graph)

Farmadupe··on Claude Fable 5 Promotional Access
I'm totally feeling you here as well. Fable is definitely quicker to first output token on `claude.ai` (ie less internal reasoning tokens being generated), and given how much more expensive decode is than prefill, I'm sure that must pay for itself pretty nicely, on top of any architectural changes that they must have made since the opus 4 architecture was locked in.
Farmadupe··on GPT-5.6
Yes this is an extremely well known result for exactly the reason you guessed. It's not just abcktracking, asking an LLM to present a conclusion and then justify is also an excellent way to provoke hallucination as the model con concts "any justification that plausibly justifies the words it's already said".

This is the actual reason why openai _invented_ reasoning models, to give them time/space to work out a solution, rather than having to magic a correct solution out of thin air from token 1.

It's less important now that all models do reasoning, but it's still almost always better to make the output come out last rather than first.

Farmadupe··on Claude Fable 5 Promotional Access
Is it too cynical to read the first quote as "when openai catch up, we will restore fable 5 to subscription plans"?

In other words, they're saying "We know have a monopoloy and we will take all of your money until someone offers a cheaper competing product"?

And "We intend to do this as quickly as we can" read as a non-promise then as much as it does now?

---

Well, I am assuming that this is all based on the fact that Anthropic have enough compute capacity to serve fable 5 now and in the future, and they're "only" limiting it in the coming week to get richer quicker. Since compute hardware is presumably relatively un-fungible, I'm assuming Anthropic isn't offering a week of fable-5 for everyone on rented hardware that they're paying exporbitant fees for?

In other words, my cynical read is "if they can serve it under the terms of subscription plans for a week", then they could serve it under the same terms for a month?

Farmadupe··on GLM-5.2 is a step change for open agents
2 and 3 bit quants are often closer to gibberish than hallucination, and that'll happen regardless of context.

I shouldn't claim too much, I haven't tried GLM5.2 at 2/3 bit quantization, but if I were a betting man I'd put money on "useless even as a chatbot"

Farmadupe··on Rotary GPU: Exploring Local Execution for Large MoE Models Under Limited VRAM
It's amazingly vacuous isn't it? I think the most interesting read was the fact that they were surprised llama.cpp crashed when they used a bad set of commandline arguments.

Although in the section immediately above the observation they claimed that they ran 10 whole completions with 100% success rate. So who knows.

I have to admit I slightly miss the flood of AI-psychosis research papers that seemed to be popping up a couple of months ago. Good to know there's still one or two new ones floating around.

Farmadupe··on Local AI needs to be the norm
I'm not sure there's a one-stop shop for this at the moment. I think the process is:

* Have a box with sufficient spare (V)RAM -- probably 8G for simple categorization with qwen3.5-4b, and 24G or more for more intelligent categorization with qwen3.6-27b or gemma4-31b.

* Download or compile llama.cpp. Choose a model, then choose one of the "quantized" builds that will actually fit on your hardware. There are literally hundreds to thousands of these per model on Hugging Face.

* Spend half a day tuning command-line parameters until llama.cpp doesn't crash.

* Watch llama.cpp regularly OOM itself, then put it in a systemd service with a memory limit so it doesn't take the entire machine down when it dies.

* Download all your photos to a folder.

* Start vibing a Python script to categorize your images by repeatedly prompting the LLM with each image in turn.

* Spend days tweaking/refining the prompt to try to get the LLM to actually do what you want.

The endgame is one of:

* The local model categorizes your images. Yay.

* The local model is too slow and you give up. Boo.

* The local model is too slow, so you spend $1k-$10k on hardware. Your image categorization task becomes a cover story for buying new gear. Yay.

* The local model can't understand your categorization metric, so you give up. Boo.

* You eagerly await news of the next open model being released. Yay?

* You consider replacing your local model with a frontier model, but then you realize you'd be spending $500 to categorize your photos. Boo.

* You refuse to allow Google/Gemini/Anthropic to train on your nudes. Boo.

Farmadupe··on Running local models on an M4 with 24GB memory
Yup, especially when for a lot of us, the price of the frontier subscription has become a cost of doing business over the last 6 months.

If you're already doing big boy stuff with big boy models, then... just carry on trucking!

Only place I'd differ is for vision/OCR tasks. Small/medium open weights models are as good as SoTa, and token prices for prefill are kinda very not worth it for larger batch tasks.

Other thing that people forget is, if you want to have even a smallish LLM as a reliable personal service, you've got to carve out 16-24 of (V)RAM and leave it permanently running.

Farmadupe··on Local AI needs to be the norm
The latest rounds of open weights vision language models are incredibly good. Like, massively good. Open weights vision capabilities trade blows with frontier models. Over the last few months I'd roughly rank capabilities as Gemini -> {chatgpt and SoTa open weights models} -> Claude.

qwen3.5-2b and qwen3.5-4b are great at document parsing. They can run on CPU

qwen3.6-27b and gemma4-31b are borderline better than the human eye in some cases. Their OCR isn't perfect, but they're seriously good. They can still run on the CPU but you'll be waiting minutes per document.

You can demand JSON, YAML, MD, or freeform text just by varying the prompt. Even if you have a custom template, you can just put that in the prompt and they'll do an OK-ish job.

There's also models that aren't in the r/locallama zeitgeist. IBM released a new 4b parameter model for structured text extraction last week, and there's a sea of recent chinese OCR models too.

IMO the open wights models are so good that in a lot of cases it's not worth paying frontier labs for OCR purposes. The only barrier to entry is the effort to set up a pipeline, and havin the spare CPU/GPU capacity.

Farmadupe··on Accelerating Gemma 4: faster inference with multi-token prediction drafters
If it helps, I mean it in a really literal sense. qwen3.6 27b is currently $3.20 per million tokens on openrouter right now which is way overpriced. As good as the 27b is, kimi k2.5 $3.00 and it's just in another league in terms of capability. There's no reason to spend money on it.

And even alibaba's own qwen3.6-plus is $1.95, so it's kinda easy to come to a conclusion that alibaba (nor anyone else) is really interested in hosting that model.

And don't get me wrong, I fully agree with you, qwen3.6 27b is an amazing model. I run it on my own hardware and every day I'm constantly surprised with what it can zero shot.

Farmadupe··on Accelerating Gemma 4: faster inference with multi-token prediction drafters
I wonder if for a model that small with a permissive license it might not be worth their time to host a commercial grade inference stack?

Might be easier to chuck it over the fence and let other providers handle it as it'll run in almost any commercial grade card?

Also speculating, but I wonder if it might also create a bit of a pricing problem relative to Gemini flashlight depending on serving cost and quality of outputs?

As a comparison, despite being SotA for their size, the smallest qwen models on openrouter (27b and 35b) are not at all worth using, as there are way bigger and better models for less oricemon a per token basis

Farmadupe··on Protobuffers Are Wrong (2018)
Using protobuf is practical enough in embedded. This person isn't the first and won't be the last. Way faster than JSON, way slower than C structs.

However protobuf is ridiculously interchangeable and there are serializers for every language. So you can get your interfaces fleshed out early in a project without having to worry that someone will have a hard time ingesting it later on.

Yes it's a pain how an empty array is a valid instance of every message type, but at least the fields that you remember to send are strongly typed. And field optionality gives you a fighting chance that your software can still speak to the unit that hasn't been updated in the field for the last five years.

On the embedded side, nanopb has worked well for us. I'm not missing having to hand maintain ad-hoc command parsers on the embedded side, nor working around quirks and bugs of those parsers on the desktop side

Farmadupe··on Misconceptions about loops in C
I feel like there's a semi-philosophical question somehwere here. Recursion is clearly a core computer science concept (see: many university courses, or even the low level implementation of many data structures), but it's surprisingly rare to see it in "day to day" code (i.e I probably don't write recursive code in a typical week, but I know it's in library code I depend on...)

But why do you think we live in a world that likes to hide recursion? Why is it common for tree data structure APIs to expose visitors, rather than expecting you write your own recursive depth/breadth-first tree traversal?

Is there something innate in human nature that makes recursion less comprehensible than looping? In my career I've met many programmers who don't 'do' recursion, but none who are scared of loops.

And to me the weird thing about it is, looping is just a specialized form of recursion, so if you can wrap your head around a for loop it means you already understand tail call recursion.

Farmadupe··on So you think you know C? (2016)
The author is making a deliberate point about undefined behaviour in the article. Hence them not executing worked examples.

In fact, by not doing so they are making a subtle implicit statement that it is uninteresting to consider actually attempting to execute these snippets.

The third paragraph of the "P.S" of the article (you have to press submit to see it) is the one that really gives the game away.

Farmadupe··on Microwatt: A tiny Open POWER ISA softcore written in VHDL 2008
FWIW, 85k 4-input luts is huge by the standards of any softcore with "micro" in the name.

It comfortably surpasses the capacity of most of the Actel aerospace FPGAs that I tend to work with.

And I think it's so "micro" that the majority of all of Lattice's FPGAs wouldn't fit it either.

And the Lattice ECP5 is advertised with 85k LUTs, which would seemingly limit its use to edification as instantiating this softcore would consume the entire chip. For any other purpose if you wanted a chip that was only a CPU, you would buy a CPU.

Farmadupe··on Microwatt: A tiny Open POWER ISA softcore written in VHDL 2008
...could be named for how much space is left after that soft core has been instantiated.
Farmadupe··on ZF makes magnet-free electric motor uniquely compact and competitive
Steppers are super optimized for being super precise and easy to controle while being (relatively) cheap to build.

They have weird "trick" geometry that means you can move the rotor to many different angles (tens, hundreds) with far fewer stator windings.

The cost of this is really low power and efficincy if you just want to spin them.

Farmadupe··on ZF makes magnet-free electric motor uniquely compact and competitive
Out of interest, what could make such a design better than a traditional inductive motor? Such motors already suffer inductive losses, and do not need slip rings or brushes, and presumably do not an additional set of windings to transfer power to excite the stator?

Just could the two sets of coils be optimized for their own purposes?

Page 1 of 2Next →