Does the difference matter?
> Note: AVIF supports two types of progressive decoding. Spatial scalability can be used to offer a lower resolution image for network constrained users and 'progressively' provide a higher resolution image by sending just the additional data required to fill in the high frequency details. Quality scalability offers a similar progression by steadily improving visual quality with each render.
That is totally not what I expected. Though it's not clear to me how exactly you achieve it.
If layering to <img> elements is fine for your use case, then it simply means you have no need for progressive decoding.