I don't buy this part. Sure, the human visual system does all kinds of processing that causes us to perceive our surroundings differently than we do a flat image on a screen. But the input to the visual system is still an image on your retina, which obeys the same laws as camera optics.
If you ignore depth of field (which the authors don't seem to be concerned with) then the human eye behaves essentially like a pinhole camera. No matter what happens behind the pinhole, all of the light rays entering the eye must pass through the pinhole from the environment, traveling in straight lines. For any given viewpoint, the relationships of which objects are occluded by which other objects will be exactly the same for an eye as for a camera. But the authors specifically point out that their algorithm doesn't preserve these relationships.
Of course, computer graphics is as much an art as a science. If deviating from the realistic model turns out to give aesthetically pleasing results, then by all means go for it. But the reason it's better wouldn't have anything to do with more closely mimicking what the human eye actually perceives.
In any case, I would question the aesthetic benefits. To my eye, the algorithm seems to distort relative shapes and sizes of objects in a weird way. It looks great in screenshots, but when moving through a scene[1], it creates a subtly unsettling "space-warping" effect.
I also think the comparisons in the article are a little disingenuous, because everyone knows that linear perspective projections with wide FOV look horrible. I'd like to see a comparison against something else, like a stereographic or fisheye projection, which would be both more physically realistic and more efficient to render.