HNHacker News
TopNewBestAskShowJobs

janwas

590 karma · joined March 3, 2016

submissionscomments
janwas··on Platform-independent SIMD in Go
We recently published another: https://github.com/google/highway/blob/master/g3doc/tutorial...
janwas··on Platform-independent SIMD in Go
We pioneered this in Highway and shared some advice on the API. Great to see this decision taken :D
janwas··on Fearless SIMD v1.0
:) Yeah, those are best copy-pasted from existing code.

I do think we've converged on the best that's possible in C++.

janwas··on Fearless SIMD v1.0
+1 on agents making library/DIY a lot more attractive than bespoke compilers.

In C++, one can also have portable intrinsics plus simpler runtime dispatching using our Highway library :)

janwas··on Vectorized and performance-portable Quicksort (2022)
VQSort users include numpy, XLA (for sparse tensors), ScaNN:

https://sourcegraph.com/search?q=context:global+hwy/contrib/...

janwas··on Vectorized and performance-portable Quicksort (2022)
(Co-)author here :)

I'd absolutely still use Highway, and do. My experience is that even two separate implementations diverge over time and I'd have low confidence in bringing the same updates and improvements to all, even with LLM assistance.

Our programming model is 1) an agent+human to generate the algorithmic approach, 2) a C++ library (Highway) to translate to intrinsics, while filling in gaps + allowing customization, 3) a compiler to generate the actual code with some optimizations.

Asking the compiler to do #1 is a pipe dream: compiler friends tell me they are not going to devise new shuffles/data layouts (like what VQSort does). Conflating #2 and #3 means a custom compiler/IR which has high engineering costs (ABI boundaries, hard to debug/profile/sanitize). And doing #3 at runtime (JIT), or moving fusions into #3 (MLIR), vastly complicates the compiler. We can still get runtime adaptability thanks to Highway's multi-target support. Fusion has been much easier to implement manually for LLMs than to construct a general fusion infrastructure. Templates hide most data type differences and we see 2-5x speedup vs llama.cpp for 128k prefill+batch decode on Zen5.

Instead of requiring a compiler to do heroic transforms at runtime, and get it right every time, we can do all kinds of agentic exploration, then verify the result/approach, check in the source code, then we 'just' have a C++ compiler afterwards. And if/when something breaks, it's easy to update centrally, in code we can modify directly, rather than indirectly via updating a compiler.

janwas··on Is JPEG XL too slow for the web?
As posted here: https://news.ycombinator.com/item?id=49693906

several years ago, using 6 threads and vectorized C++, we saw about 240-270 MP/s decode speed. Hence this should be possible in < 100ms. Not sure how the current implementation differs from that state :)

janwas··on The case against JPEG XL
> This tells me you don't know how codecs work

Laughable.

My conservative estimate is that a few-MP image with all but a smallish region encoded using skip blocks will spend a few KiB on that. This is very expensive compared to sending only a bounding box, hence it is unattractive for purposes of updating small regions with a whole-image layer.

I am glad to hear AVIF progressive has improved. But note that my original comment was: I find it misleading to call AVIF's "up to four passes" "very flexible". I believe that stands: contrasting the flexibility of (purpose-built) JPEG XL vs. the fairly strict limitations (inherited from video) of AVIF, I am astonished anyone would still call the latter "very flexible" by comparison.

janwas··on The case against JPEG XL
Not at all.

CPU power is proportional to frequency^2. Running on 4-6 little/efficiency cores (which are widespread on mobile) is likely faster than one big core, and uses less energy.

janwas··on The case against JPEG XL
There is indeed an issue with the JPEG format that makes parallelization difficult: the lack of a 'table of contents' with offsets to tiles.

Various workarounds (for example RST markers or self-sychronizing properties of Huffman) have been proposed, but these are not great and do not work for all images.

JPEG XL ensures this information is always available.

janwas··on The case against JPEG XL
I note you made no response to the objection about representing a temporal sequence with a single, nonrepresentative and cherry-picked, screenshot.

As to bad experience, that seems like a legit personal preference, but disturbingly un-nuanced and absolutist if intended to apply to everyone.

When on a train with spotty wifi or frequent tunnels, I would actually prefer to see whatever arrived. Someone on prepaid data might want to truncate at ~100 KB, no matter what the website owner or browser PM thought.

janwas··on The case against JPEG XL
Ah, an accusation of misquoting. I actually quoted your exact words minus "is" and "it supports".

I think we disagree on the degree of flexibility, for sure.

A cap on layers (under user/browser) absolutely makes sense, but 4 at the format level is quite limiting, especially if you want to spend some of them on salient regions.

I agree that's possible, but not that it's efficient. You'd waste a few KiB on encoding skip blocks - AVIF layers represent the whole image, whereas JPEG XL can efficiently encode and update at group level.

How flexible did Jake find AVIF progressive in 2025? [1]

"it seems pretty limited. Only particular scaling values are allowed, and 1/8 is the smallest. Supposedly, additional layers are possible[..], but whenever I tried this, the encoder would error out, or explode the file size to ~400 kB, even at lowest quality. I guess that's why it's marked 'experimental'."

> Intermediate passes in AVIF can semantically be different from the final pass

Also true of JPEG XL - scans are additive.

[1]: https://jakearchibald.com/2025/present-and-future-of-progres...

janwas··on The case against JPEG XL
We can agree on wanting a better internet :)

I participated in the design of those filters, so no, I do not deeply dislike the way they look.

This gaslighting is not convincing. No matter how many filters AVIF has, I distinctly remember tile artifacts being particularly disturbing. More so than individual blocks, whose size and border effects differ; tiles are a straight line through the entire image.

JPEG XL does not have this problem because the design and codestream enables parallel decoding (thanks to per-group offsets encoded in the 'TOC'), hence does not require separate tiles.

janwas··on The case against JPEG XL
Surprised and disappointed to hear "bad-faith reading".

I have seen AVIF tiling artifacts myself. Hand-waving them away by appealing to a metric that averages across all image pixels is not convincing.

janwas··on The case against JPEG XL
Sounds like some strong assumptions here, particularly a stable and non-metered connection.

Imagine fast scroll across an image gallery on a slow connection (including cell handovers).

Or range requests, where a service worker only downloads the header+preview portion, and when clicking on the image, no need to re-download that.

Or even a browser that truncates all images, to protect users who might visit a page with huge background images that blows through their prepaid data plan.

JPEG XL anticipated, and accommodates, these use cases.

janwas··on The case against JPEG XL
I find it misleading to call AVIF's "up to four passes" "very flexible".

It seems quite limited compared to the JPEG XL ability to truncate the bitstream anywhere, or send the progressive updates for salient regions first [1].

[1]: https://opensource.googleblog.com/2021/09/using-saliency-in-...

janwas··on The case against JPEG XL
+1, this is super unfair to show the one point in time where AVIF progressive looks better - right after it receives its 'preview' (which as you say is 4x as big as JPEG XL's).

I value integrity, especially when communicating results. This is shameful.

(Disclosure: I worked on JPEG XL)

janwas··on The case against JPEG XL
Something is fishy here.

The JPEG XL report [1] measured between 240-270 Megapixels/s on 6 cores using the C++ implementation (disclosure: I was responsible for its SIMD/threading), about twice as fast as the then-current libaom.

Measuring on a single core is deeply misleading because our code was designed to scale well. I believe AVIF requires tiling in order to parallelize, which causes artifacts at tile boundaries.

[1]: https://arxiv.org/pdf/2506.05987

janwas··on Ask HN: Is it time to run the LLM engines on the CPU?
Not sure what this comment is based on. Arm introduced a scalable SIMD whose whole point is to be expandable. AVX-512 works very well, for example on Zen 4 and 5.
janwas··on Everyone should know SIMD
That is not at all my experience :) Please expand on what "vertically-oriented scope" means.
janwas··on GigaToken: ~1000x faster Language model tokenization
If you are running on large-scale data, have you validated at that scale (comparing results)? From a quick look at the code, it looks like there is a 42-bit hash (computed via single-mul hash function) which can have collisions and thus return the wrong tokens, right?
janwas··on GigaToken: ~1000x faster Language model tokenization
hm, maybe not so trivially correct here. Do I understand correctly that incorrect results can happen as a result of a 42-bit hash collision? That could happen after less than one MB of input, given the simple one-mul hash.

BTW throughput is measured for a 12 GiB file. Would be interesting to see the throughput for something more like 32 KiB, with cold start (token cache not yet populated).

janwas··on We recommend Highway over std:SIMD
Author/Highway TL here. Happy to discuss.
janwas··on Bun team is rewriting SIMD from Rust to C++
Impressive result. Congrats!
janwas··on C++26 Shipped a SIMD Library Nobody Asked For
Oops, the final T got cut off somehow, sorry about that.

https://gcc.godbolt.org/z/KM3ben7ET

janwas··on C++26 Shipped a SIMD Library Nobody Asked For
Any suggestions for improvement? We went through >5 iterations of the dispatching and I am fairly confident this is about as good as it gets in current C++. I suppose "macro hell" is a matter of taste. Objectively, we have six dispatch related macros in the example: https://gcc.godbolt.org/z/KM3ben7E The ~two dozen lines of boilerplate are generally copied from an example. But why multi-file?
janwas··on C++26 Shipped a SIMD Library Nobody Asked For
Working on one together with fastcode.org :)
janwas··on C++26 Shipped a SIMD Library Nobody Asked For
To be clear, "better abstractions" here seems to mean macros for assembly language. To each their own.

What bothers me is advocating for this, or denigrating more generally useful alternatives, without mentioning the very narrow niche where this sits.

Video codecs only change every few years. This makes it more worthwhile/feasible to spend eng time on a few kernels.

Even then, not supporting SVE (you don't, right?) gives less incentive for the Arm CPU ecosystem to invest in it, helping keeping us stuck in the NEON local minimum. Not ideal :/

janwas··on C++26 Shipped a SIMD Library Nobody Asked For
Correction (typo): Z13 lacks fp32.
janwas··on C++26 Shipped a SIMD Library Nobody Asked For
Oh, interesting :) I meant Fastcode.org.
Page 1 of 16Next →