They don't really "label every pixel" in the sense that I think about it.
Instead, they have a collection of dense arrays representing the image data itself, then have metadata at the per-array, or overall level.
A typical dataset I work with is multidimensional, it starts as:
1) 2D planes of multiple channel image intensities, typically 5K x 5K pixels, each covering just part of an overall field of view. These are like patches when you do panoramas- take 20 partly overlapping shots. Each plane contains multiple channels- that could be "red green and blue" or more complicated spectral distributions.
2) 3D information- the microscope takes photos at various depths, only the "in-focus" (within the volume of view) information. These can be stacked (like depth stacking) or turned into a 3D "volume".
3) Maybe the data was collected over multiple time points, so (1) and (2) repeat every hour. Other parameters- like temperature, etc, could also represent an entire dimension.
4) Every 2D plane has its own key-value metadata, such as "what color channels were used", "what objective was used" (magnification), and lots of other high-dimensional attributes (that's what they mean by "each pixel must be labelled with metadata"- the 3D position is the same for every pixel in a 2D plane.
5)
Generally all of this is modelled as structures, arrays, and structures of arrays/arrays of structures. In the case of OME-zarr, it's modelled as an n-Dimensional array with dimensions expressed in a filesystem hierarchy (first dimension typicall the outermost directory, innermost dimension usually a flat file containing a block of scalar points using some compressed numpy storage. Then at each level of the directory you have additional .json files which contain attributes at that level of the data.
Those partly overlapping 2D planes are often assembled into panoramas, which can be a lot more convenient to work with. There are various tools for working with this- I've used navigation map javascript, but napari is a desktop app wiht full suport for sectioned viewing of high-dimensional (7d) data.
OME-zarr is nice because it sort of uses the same underlying tech that the machine learning folks use, and it's ostensibly "optimized for object storage", but I still have lots of complaints with the implementation details, but it's important for me not to distract the OME-zarr team from making the standard successful.