Might be interesting to experiment with different mappings of pixel location to audio sample number, rather than just having a row by row linear scan from corner to corner.
Right now I've mostly been exploring YUV and RGB colorspace in either packed or planar formats.
Would this be middle-out compression?