The clever bit is yet another camera, of a different kind, which somehow determines distance-to-camera at high resolution. Supposedly the key to this is the fact that intensity is proportional to 1/distance^2, but just how they use this (given that different objects will be different in brightness) the article doesn't say.
Fortunately, the lead researcher has a press release at http://individual.utoronto.ca/iizuka/research/OmnifocusVideo... with a bit more detail: they illuminate the scene with two IR sources in turn, one closer than the other, and look at how brightness changes between the two. (So they have C/r^2 and C/(r+k)^2, where k is the known separation between the IR sources, r is the distance to the nearer source, and C is the IR reflectivity; from that they can compute r.) I guess this doesn't work so well with anything that doesn't reflect enough IR.
At present N=2, and their demonstrations all show an image with some "near" and some "far" bits and nothing in between. Presumably they've carefully focused one camera on the "near" and one on the "far".
Using N cameras in this way gets you N times the depth of field, and 1/N as much light into each camera. If instead you make your aperture N times smaller, you again get N times the depth of field, but now 1/N^2 as much light into the camera. So it does seem possible that it might be a win, for applications where you really want as much depth of field as possible, and nothing in your scene is moving much, and everything reflects enough IR.