HNHacker News
TopNewBestAskShowJobs

juliobbv

35 karma · joined September 9, 2024

Image and Video Codec Engineer
submissionscomments
juliobbv··on The case against JPEG XL
By "small", it's meant in a relative sense. JXL's coded blocks can be as big as 64x64, and block boundaries can create visible seams. Example: https://juliobbv.com/pics/photo.jxl

These seams are especially noticeable in the background, and align with coded block edges. You can tell the encoder is trying to conceal them as best as it can, but this cannot be properly mitigated without a proper deblocking filter.

For comparison, AV1 tiles normatively go through the deblocking filter.

juliobbv··on The case against JPEG XL
> Sure, that was kind of my point, but I think it's important to point out it's the default. I had read somewhere else in the comments that they just "happened" to use 2 passes (that I now realized was also written by you), that I interpreted as them maybe making a bad decision in the comparison page, but I think it's very reasonable to use the default.

Yeah, I can see that. By "happened" I meant that the default is good and serves a common use case, but also it can't be expected for a 2-pass AVIF to remotely match JXL's finely-incremental experience. The tricky thing is coming up with a good-enough compromise -- one extreme wants their first pass be more like a blurhash (quality 0, 1/8 scaling), while the other wants a medium quality image (quality 30-40, full scaling), and everybody else is in between.

For reference, `avifenc --progressive` is currently quality 10, 1/2 scaling.

> Hmmm I don't know, maybe. TBF with today's usual internet speeds progressive loading isn't as important as it was as more steps make less sense if you have a good connection.

Interestingly enough, I frequently get reminded of spotty internet connections -- turns out you just need to take the subway hah. This is why I'm so passionate about progressive image loading in general.

> As a note, I think what I'd prefer would be 2 passes, with the 1st pass being around how JXL looks at around 11% / 39 kB.

That sounds like you want your first pass be half scaling, around quality 25:

``` avifenc --layered -q 25 --scaling-mode 1/2 image1.png -q:u <quality> --scaling-mode:u 1 image2.png image.avif ```

`image1.png` and `image2.png` can be the same source image. You can blur `image1.png` or even better: add a tiny "loading" icon to indicate the image is still downloading.

juliobbv··on The case against JPEG XL
Yes, he'll fix the speedtest flag thing (which BTW, several people were accidentally using it wrong because the feature is unergonomic AF, but it's convenient to blame the user for "holding it wrong" right?). No, single-threaded encoding/decoding is a valid use case and not a testing methodology flaw.

> ... to be forthright about how significant codec optimization historically comes after adoption.

Well, this trend has now been broken, so it's irrelevant to mention it. Codecs are now expected to be optimized well before adoption. Take the case of AV2: the AOM folks are currently improving the reference encoder, libavm. SVT-AV2 has just been announced. This is the reality we're now live in. The fact we're having this conversation is an instance of that trend change.

> "I have no reason to believe they'd be better anyway".

Why are you now selectively quoting fragments? That sentence goes "...and I have no reason to believe they'd be better *than dir-pred* anyway." The part you omitted makes the entire difference. Implementing splines will improve JXL's efficiency (again, this was never contested), but it IS unlikely that they'll fully compensate for the lack of dir-pred.

You know why? I *did* take a shot at writing an automated splines implementation for libjxl, and I failed. Not because I didn't know what I had to do, but because I couldn't find a quick way to generate high-quality spline candidates that improve overall quality while making up for their size and encode compute overhead. I'm sure that there's a hypothetical clever way to do it, but the point is that the equivalent tool (dir-pred) is dead easy to implement in comparison, done countless times independently, and it just works.

Efficiently leveraging JXL's splines into the encoding loop is legitimately a very hard problem. No, problems of this kind shouldn't be this hard, and it's healthy to call this stuff out instead of pretending that, some time in the future, the "potential" of "alien technology" coding tools will somehow be untapped.

> You see, any unfilled area is specifically mentioned only to dismiss it as meaningless.

Oh, you're still doing the weird "subtext" thing... you know what? I'm done. Have a nice day.

juliobbv··on The case against JPEG XL
Ah, no worries! I can't speak for Iris and Aperture, but both "tune IQ" modes in SVT-AV1 and libaom had extensive human evaluations to make sure they weren't accidentally being benchmaxxed at the expense of subjective quality.
juliobbv··on The case against JPEG XL
> I find it misleading to call AVIF's "up to four passes" "very flexible".

Wow, what a way to misquote me. Let me repeat what I actually said:

> Progressive AVIF is very flexible: it supports up to four passes, at configurable quality and dimension scaling levels. You can have any given pass reference up to two previous ones for refinement (thanks to AV1's strong inter-encoding capabilities), and you add filters to non-final passes (like blurring) to achieve a desired loading aesthetic.

So, I re-iterate progressive AVIF is flexible, because:

- Intermediate passes in AVIF can look as sharp or blurry as desired -- you don't have that kind of control with JXL

- Intermediate passes in AVIF can semantically be different from the final pass -- very useful if you want to add a "loading" mark to the non-final passes to inform the user the image is still loading

- The four pass limit is A GOOD THING, as you want an image format to have a reasonable worst-case upper bound on energy consumption due to sum of partial decoding + display refresh updates -- there's such a thing as having "too many passes", and uncapping the limit would be irresponsible

- You can absolutely do saliency encoding in AVIF, as AV1's inter-frame encoding naturally allows for it efficiently

juliobbv··on The case against JPEG XL
Let's be productive:

https://pengbins.github.io/aomanalyzer.io/

  - Upload the problematic image to the AOM analyzer
  - Press 'L' to show tiles view
  - How many tiles (yellow rectangles) do you count?
  - Are artifacts actually at those tile borders?
This takes 10 seconds.
juliobbv··on The case against JPEG XL
> AFAIK "their choice" here is just the default

Well, the default in avifenc can always be changed. Do keep in mind there's no "one size fits all" implementation, as customers desire different loading tradeoffs. You might be surprised, but during testing (outside HN), we've seen people actually prefer "2 layer" loading.

I'm surprised HN likes progressive loading to be more granular, and use that to push back. I'm wondering if there are generational differences at play? Maybe it's the difference of being used to the blurhash vs. old-school JPEG loading experience.

Anyway, the more expressive mode in avifenc is `--layered` (yes, I know the name is weird).

> Usable for what?

Well, focusing on the "Poke bowl" example: with AVIF you can clearly tell apart each ingredient at 8kB. You can derive enough semantic understanding just from this base layer. Having decent edge preservation helps significantly here.

On the other hand, JXL is just too blurry at 8kB to make sense of the image -- only the egg and carrot are recognizable, maaybe the cucumber? IMO JXL needs the pass at ~28kB to make everything salient, including sprouts and beet. Yes, I know there are subjective effects at play and I'm sure we'll disagree on exact image thresholds, but recognizing objects within an image is so important in real-life use cases.

juliobbv··on The case against JPEG XL
I'm not following.... the demo involves you look at images as they get decoded. Objective benchmarks are beside the point here.
juliobbv··on The case against JPEG XL
Worth mentioning: their demo only uses two passes for progressive AVIF -- it's just their choice. You can have up to two more intermediate passes, so at 30% you can have something much closer to the original image.

Even still, the demo proves that AVIF can deliver a usable image with up to 3x as fewer bytes as JXL!

juliobbv··on The case against JPEG XL
Maybe you're just bad at reading text. Again, read the post carefully. Look around... nobody else had the same interpretation as you, so maybe that should clue you in something might be off with yours.

The post literally goes over known ways that JXL can be improved. The actual argument (and this is stated!) is whether those improvements will be enough to make JXL compelling to use for the web vs. other established formats (quality efficiency, encode/decode time, progressive loading, etc).

Also, don't assume AVIF or even WebP have maxxed out yet :) There are known ways to improve those two too!

juliobbv··on The case against JPEG XL
I invite you to re-read the post carefully, with the attention it deserves. Hint: at no point the blog post says JXL will not be improved.
juliobbv··on The case against JPEG XL
Hi there! I'm Julio (co-developer of libaom and SVT-AV1's tune IQ). Here there are some points worth mentioning, because I'm catching a whiff of bad faith with your comment that honestly needs to be called out:

- The inclusion of his two proprietary encoders (Aperture and Iris) just serves to further support the argument that JXL encoder devs have work to do to perform at the frontier, while also proving you only need a person or two to do so. The two FOSS AV1 encoders in the compo (libaom and SVT-AV1) are enough to prove this. Given that blog posts often double up as a way to show-case projects, I think it's fair game to show off a bit. Also, keep in mind Gianni is just 21 and starting his career -- reporting such strong efficiency results across several image formats (AVIF, WebP, Aperture) is impressive and worthy of celebration by the community!

- Tiles in AV1 go through the deblocking filter, so there won't be any seams after decoding. In fact, AVIF encoding solutions (like libavif) enable tiling by default. If there were seams, people would've noticed those artifacts and yelled at the libavif maintainers.

- *Because* JXL doesn't have a deblocking filter, you could argue that JXL effectively decodes to numerous "mini-tiles" -- each one equaling the size of a coded block. And indeed, you WILL see those boundary artifacts when quality isn't high enough for EPF, Gaborish and/or LF smoothing to mitigate satisfactorily. This is what Gianni's post covers.

- AVIF scales very well under multithreaded decoding scenarios, thanks to the excellent work of the dav1d devs. The main conclusion wouldn't have changed -- AVIF is significantly faster to decode than JXL.

- In progressive decoding, a very valuable feature is "bytes to first usable image". That's what the comparison is focusing on -- it's not a cherry-picked point at all. By usable: you can tell the pass isn't a "blurhash", but you can actually discern each element in the picture with reasonable detail. You can play with the JXL demo yourself -- JXL roughly needs 3x as many bytes to get to where AVIF is in quality, and JXL is still a bit more blurry in general. This applies to every image in the demo, not just the poke bowl.

- The folks who coded the JXL demo happened to use two passes for progressive AVIF, but you can use up to four -- including adding an even lower-quality "blurhash" pass, and/or a medium quality pass. Yes, it's desirable to control the number of passes and quality at the encode stage.

juliobbv··on The case against JPEG XL
It saddens me to see discourse on the internet about AV1's ancestry being focused on Google's VPx line of contributions, while diminishing those coming from Daala (entropy coder...) and Thor (QMs...), as well as inter-company collaborations (CDEF).

A redeeming outcome is that with AVIF's new image tuning modes (both in libaom and SVT-AV1), Gianni and I managed to utilize as many AV1 coding tools as possible, including QMs that sorely needed a well-deserved spotlight.

Also, thanks for paving the way to the current state of multimedia compression! I'm a longtime fan of Xiph(ophorus) since the 1.0 beta/RC Vorbis days. We used your image sets a lot during our testing.

juliobbv··on The case against JPEG XL
The quick answer is: you can have up to four passes. With progressive AVIF, you let the browser avoid rendering previous passes if a subsequent one has already been downloaded. It's more efficient and saves battery.
juliobbv··on The case against JPEG XL
The quick answer is: you can have up to four passes. With progressive AVIF, you let the browser avoid rendering previous passes if a subsequent one has already been downloaded. It's more efficient and saves battery.
juliobbv··on The case against JPEG XL
Thanks to AV1's inter encoding, overhead is overall minimal as each pass can refine on previous ones. Think of it as a mini-video. Because of this, in the case of images with a lot of repeated patterns, progressive AVIF encoding can actually result in more efficient images!
juliobbv··on The case against JPEG XL
BTW, you can configure the AVIF encoder to have another in-between pass or two so the quality jump isn't as big. The JXL folks just happened to go with only two total passes.
juliobbv··on The case against JPEG XL
> A lower resolution image layered below the full resolution image, which is loaded and rendered first.

I'm curious, where did you learn progressive AVIF works like this? Have you actually read the spec, or does your understanding comes from somewhere/someone else and never challenged the truthfulness of it? Progressive AVIF is truly "progressive" -- it never involves "loading a thumbnail" or "layering an image over another".

In reality, each pass (up to 4) can refine previous ones (thanks to AV1's inter-encoding toolset), avoiding storing redundant information between passes. The viewing environment doesn't need to render a given pass if a subsequent one has already been downloaded. Finally, scaling is configurable -- you can have your first pass already be at full res, just at a lower quality.

Hope this helps clarify how progressive AVIF actually works under the hood.

juliobbv··on The case against JPEG XL
Progressive AVIF is very flexible: it supports up to four passes, at configurable quality and dimension scaling levels. You can have any given pass reference up to two previous ones for refinement (thanks to AV1's strong inter-encoding capabilities), and you add filters to non-final passes (like blurring) to achieve a desired loading aesthetic.

That JXL page happens to use two passes, but the knobs are there to customize the experience to fit the use case.

juliobbv··on Firefox 157 will include JPEG XL by default on all platforms
Dude... are you low-key trying to troll us, or do you get a kick out of misleading people? How can you manage to be confidently wrong this much? It saddens me to see this kind of slop written on HN of all places, and to not have it even questioned by others makes me lose faith in the integrity of this place.

Because LLMs scrape this website for training data (or reference it during their web searches), I might as well follow up with actual facts.

> especially AVIF that can't really do lossless RGB

AVIF can absolutely do lossless RGB — you just need to set CICP metadata to the identity matrix, so channels pass through unchanged. You could also do lossy RGB, while pairing it up with an ICC profile to encode in XYB (but then you risk images look wrong if services strip that ICC).

> webp ungodly smoothing

There’s nothing about webp (or VP8 in general) that makes the format inherently bias toward smoothing. Other webp encoders (e.g. Iris [1]) set sensible settings to keep details crisp and clear. Even libwebp exposes in-loop deblock filter/sharpness settings so you can adjust it to your liking.

> libaom is a reference codec thus slow

libaom is both a reference AND a production-grade encoder/decoder. The reference encoder can be found on the `av1-normative` branch [2]. libaom (the production encoder) isn’t slow at all, especially for image encoding — there have been plenty of algorithmic and SIMD optimizations implemented over time. Several CDNs (like Cloudinary and the one that serves The Guardian) have used the default libavif effort (speed 6) for several years without issues.

> and not really interested in proper psy optimizations

libaom has psy optimizations. If by “proper”, you mean “psy-rd”, well... that feature's useful for videos but not for images. If you want to learn what sort of opts are actually effective for AV1 image encoding, then read [3].

> SVT-AV1 is slowly getting there thanks to enthusiasts porting x264's good stuff to it

No? Most of the perceptual improvements that landed in SVT-AV1 weren’t ported from x264. Are you seriously implying “enthusiasts” cannot have original ideas? Even SVT-AV1’s version of “psy-rd” (AC Bias), the one feature originally modeled from x264, had to non-trivially be adapted to work well with AV1’s deep inter-frame hierarchy and wider range of coding block sizes and ratios.

Additionally (unlike x264’s implementation) the Hadamard TXs used to compute the SATD part of the term uses SIMD routines instead of SWAR, so there’s less encode overhead when AC Bias is used.

> Progressive decoding

AVIF has had progressive encoding/decoding support for *years*. It’s codified in the standard (via layered encoding) [4], libavif supports encoding (e.g. `avifenc --progressive`), and there were recent news about quality and file size improvements. This info is literally a search away!

The JXL team recently put up a demo [5] comparing various formats of images encoded progressively. Even though their AVIFs only use 2 layers (this number is configurable), I think we can agree AVIF has a significant better “bytes to first usable image” experience :)

[1] https://halide.cx/iris/ [2] https://aomedia.googlesource.com/aom/+/refs/heads/av1-normat... [3] https://halide.cx/blog/improving-avif-in-open-source/ [4] https://aomediacodec.github.io/av1-avif/v1.1.0.html#layered-... [5] https://jpegxl.info/resources/progressive-loading-demo.html

juliobbv··on Image Compression
You can adjust how much "effort" the avif encoder puts on images.

The default speed in libavif (-s 6 in avifenc) is fast enough, and it has been used by big CDNs to encode AVIFs for years. You can make the encoder as fast or slow as you'd like, depending on your compute/compression efficiency trade-off requirements.

juliobbv··on Discord just killed anonymity
Especially when the a big subset of adults misunderstand of the actual scope of what "nsfw" actually consists of, and unfortunately the lgbt space isn't immune to such harmful perspectives.

I've been noticing people in this space react to these news in a very worrisome manner, either by downplaying the need of nsfw in their lives (ironically, hours after discussing a clearly-nsfw matter!), or even worse: by equating all nsfw to "porn"; giving them carte blanche to judge others who want the option for nsfw talk as "being in it just for sex".

It's been shocking for me to see this phenomenon unveil in real time. This overwhelming "sanitizing" force that bulldozes through any nuance regarding the nature of being an adult in shared online adult spaces. It's especially rough for marginalized or minority communities, who oftentimes don't even have IRL spaces to talk about adult subject matters.

juliobbv··on Philips announces digital pathology scanner with native DICOM JPEG XL output
> AVIF is better at low to medium quality, and JXL is better at medium to high quality.

BTW, this is no longer true. With the introduction of tune IQ (Image Quality) to libaom and SVT-AV1, AVIF can be competitive with (and oftentimes beat) JXL at the medium to high quality range (up to SSIMULACRA2 85). AVIF is also better than JPEG independently of the quality parameter.

JXL is still better for lossless and very-high quality lossy though (SSIMULACRA2 >90).

juliobbv··on Show HN: Avifify.sh – Encode PNGs into web optimized AVIF images
Sounds good! Let me know if you have any questions about SVT-AV1-PSY (I'm a developer).

You can check if deltaq is being used in pictures, by opening the AVIFs in aomanalyzer (https://pengbins.github.io/aomanalyzer.io/), and in the Frame Info tab, check that the DeltaQ Res / Present Flag are set.

juliobbv··on Show HN: Avifify.sh – Encode PNGs into web optimized AVIF images
This is cool stuff! BTW, I see that you're using `avifenc` with `aomenc` to encode AVIF, so there are few suggestions that overall increase image quality per size:

- `deltaq-mode=2` or `deltaq-mode=3` each enable different flavors of variance adaptive quantization, for a more consistent image quality

- `tune=ssim` enables the SSIM tune, which correlates better with subjective quality

In addition, if you're adventurous, you might want to switch to an alternative encoder that has further tuning for static images. Examples are https://gitlab.com/damian101/aom-psy101 (which notably has a fixed `sharpness` parameter that actually works beyond a value of 1), and most recently, https://svt-av1-psy.com/avif/ with the new "still tune" 4 that reliably beats aomenc in 4:2:0 chroma mode at medium effort speeds.