590 karma · joined March 3, 2016
I do think we've converged on the best that's possible in C++.
In C++, one can also have portable intrinsics plus simpler runtime dispatching using our Highway library :)
https://sourcegraph.com/search?q=context:global+hwy/contrib/...
I'd absolutely still use Highway, and do. My experience is that even two separate implementations diverge over time and I'd have low confidence in bringing the same updates and improvements to all, even with LLM assistance.
Our programming model is 1) an agent+human to generate the algorithmic approach, 2) a C++ library (Highway) to translate to intrinsics, while filling in gaps + allowing customization, 3) a compiler to generate the actual code with some optimizations.
Asking the compiler to do #1 is a pipe dream: compiler friends tell me they are not going to devise new shuffles/data layouts (like what VQSort does). Conflating #2 and #3 means a custom compiler/IR which has high engineering costs (ABI boundaries, hard to debug/profile/sanitize). And doing #3 at runtime (JIT), or moving fusions into #3 (MLIR), vastly complicates the compiler. We can still get runtime adaptability thanks to Highway's multi-target support. Fusion has been much easier to implement manually for LLMs than to construct a general fusion infrastructure. Templates hide most data type differences and we see 2-5x speedup vs llama.cpp for 128k prefill+batch decode on Zen5.
Instead of requiring a compiler to do heroic transforms at runtime, and get it right every time, we can do all kinds of agentic exploration, then verify the result/approach, check in the source code, then we 'just' have a C++ compiler afterwards. And if/when something breaks, it's easy to update centrally, in code we can modify directly, rather than indirectly via updating a compiler.
several years ago, using 6 threads and vectorized C++, we saw about 240-270 MP/s decode speed. Hence this should be possible in < 100ms. Not sure how the current implementation differs from that state :)
Laughable.
My conservative estimate is that a few-MP image with all but a smallish region encoded using skip blocks will spend a few KiB on that. This is very expensive compared to sending only a bounding box, hence it is unattractive for purposes of updating small regions with a whole-image layer.
I am glad to hear AVIF progressive has improved. But note that my original comment was: I find it misleading to call AVIF's "up to four passes" "very flexible". I believe that stands: contrasting the flexibility of (purpose-built) JPEG XL vs. the fairly strict limitations (inherited from video) of AVIF, I am astonished anyone would still call the latter "very flexible" by comparison.
CPU power is proportional to frequency^2. Running on 4-6 little/efficiency cores (which are widespread on mobile) is likely faster than one big core, and uses less energy.
Various workarounds (for example RST markers or self-sychronizing properties of Huffman) have been proposed, but these are not great and do not work for all images.
JPEG XL ensures this information is always available.
As to bad experience, that seems like a legit personal preference, but disturbingly un-nuanced and absolutist if intended to apply to everyone.
When on a train with spotty wifi or frequent tunnels, I would actually prefer to see whatever arrived. Someone on prepaid data might want to truncate at ~100 KB, no matter what the website owner or browser PM thought.
I think we disagree on the degree of flexibility, for sure.
A cap on layers (under user/browser) absolutely makes sense, but 4 at the format level is quite limiting, especially if you want to spend some of them on salient regions.
I agree that's possible, but not that it's efficient. You'd waste a few KiB on encoding skip blocks - AVIF layers represent the whole image, whereas JPEG XL can efficiently encode and update at group level.
How flexible did Jake find AVIF progressive in 2025? [1]
"it seems pretty limited. Only particular scaling values are allowed, and 1/8 is the smallest. Supposedly, additional layers are possible[..], but whenever I tried this, the encoder would error out, or explode the file size to ~400 kB, even at lowest quality. I guess that's why it's marked 'experimental'."
> Intermediate passes in AVIF can semantically be different from the final pass
Also true of JPEG XL - scans are additive.
[1]: https://jakearchibald.com/2025/present-and-future-of-progres...
I participated in the design of those filters, so no, I do not deeply dislike the way they look.
This gaslighting is not convincing. No matter how many filters AVIF has, I distinctly remember tile artifacts being particularly disturbing. More so than individual blocks, whose size and border effects differ; tiles are a straight line through the entire image.
JPEG XL does not have this problem because the design and codestream enables parallel decoding (thanks to per-group offsets encoded in the 'TOC'), hence does not require separate tiles.
I have seen AVIF tiling artifacts myself. Hand-waving them away by appealing to a metric that averages across all image pixels is not convincing.
Imagine fast scroll across an image gallery on a slow connection (including cell handovers).
Or range requests, where a service worker only downloads the header+preview portion, and when clicking on the image, no need to re-download that.
Or even a browser that truncates all images, to protect users who might visit a page with huge background images that blows through their prepaid data plan.
JPEG XL anticipated, and accommodates, these use cases.
It seems quite limited compared to the JPEG XL ability to truncate the bitstream anywhere, or send the progressive updates for salient regions first [1].
[1]: https://opensource.googleblog.com/2021/09/using-saliency-in-...
I value integrity, especially when communicating results. This is shameful.
(Disclosure: I worked on JPEG XL)
The JPEG XL report [1] measured between 240-270 Megapixels/s on 6 cores using the C++ implementation (disclosure: I was responsible for its SIMD/threading), about twice as fast as the then-current libaom.
Measuring on a single core is deeply misleading because our code was designed to scale well. I believe AVIF requires tiling in order to parallelize, which causes artifacts at tile boundaries.
BTW throughput is measured for a 12 GiB file. Would be interesting to see the throughput for something more like 32 KiB, with cold start (token cache not yet populated).
What bothers me is advocating for this, or denigrating more generally useful alternatives, without mentioning the very narrow niche where this sits.
Video codecs only change every few years. This makes it more worthwhile/feasible to spend eng time on a few kernels.
Even then, not supporting SVE (you don't, right?) gives less incentive for the Arm CPU ecosystem to invest in it, helping keeping us stuck in the NEON local minimum. Not ideal :/