They are very likely using VQVAE to create a dictionary of tokens and then just converting images into them with an encoder.
If you need to be able to reason about multiple objects in the image and their relative positions, then don't you need to use a tiled approach?