I'm not entirely sure how to begin to answer this question. Forgive me if I start too early.
An SVG is a list of compositing instructions for some (graphical) artefact. These instructions are given so that their dependencies (the things that must be composed before other things) are stated directly.
If you have familiar with Adobe's Photoshop, you might imagine a SVG as a set of nested layers: An SVG "renderer" will simply compose these layers together in order to have some pixels to display.
Now, to give a clear example of what I'm referring to, I am going to show you a simplified SVG that composes four png files together in tiles:
<svg width="100" height="100">
<image x="0" y="0" width="50" height="50" href="data:image/png,xxx" />
<image x="50" y="0" width="50" height="50" href="data:image/png,xxx" />
<image x="0" y="50" width="50" height="50" href="data:image/png,xxx" />
<image x="50" y="50" width="50" height="50" href="data:image/png,xxx" />
</svg>
That data:image/png,xxx stanza is the SVG-encoding of a PNG (a slight simplification: the format actually belongs to a number of different standards, that the SVG specification leverages).
That is to say, I'm not exactly suggesting converting a PNG to an SVG: I am also suggesting breaking apart the (large!) PNG into several component PNG files (tiles) so that they can be decoded independently. Note carefully my example, how none of the tiles overlap. A decoder can (trivially) determine these instructions are independent, and so process them independently.
Being able to decode parts of the resulting image independently is what the iDOT metadata makes possible: It is essentially a different encoding of the x/y/width/height information in the above, the difference is that SVG already existed.