Fooling Around with Foveated Rendering
peterstefek.me
peterstefek.me
I was wondering what determines the speed at which the stars in foveal vision rotate. I assumed the movement is the illusion, but it turns out the stars all rotate - there's an explicit parameter determining their period in the shader. The illusion is that the stars outside your central vision seem frozen!
As far as I understand, the effect is caused by the fact that the fovea has a much higher resolution (of bright light) due to its higher density of cones. At a certain size, the fovea is still able to make out details of the elements, but the rest of the retina can not. Thus they are detected as rotating in the fovea area, but not outside (since it's basically just a colored lump there. We still perceive the outside ones as stars (and not just color lumps) because the brain interpolates them that way due to previously perceived information.
Under low light conditions, i.e. when rods would be the main means of vision, an opposite effect can be observed: We are mostly blind in the fovea area due to the lack of rods there. This manifests as objects showing more detail at night when looking slightly past them.
Do other people get the same effect?
I thought the opposite would happen.
Both the fovea and the rest of the eye perceive more detail when you're closer to an object (the fovea is no different in this respect). When you're close enough, the fovea can perceive a larger area in detail, which we see as motion in that animation.
However, I have a vision condition where I can't see blue in my actual fovea. And that point is incredibly small. Maybe 3-4 times smaller than what you are seeing here.
You need to look right at where they want you to look - so in this example you need to look right at the center point of the image. If your gave drifts off to the side then it looks bad.
The same with 3D movies, if you don't want to look directly at what the director wanted you to look at, you end up getting seasick.
Do VR headsets have gaze tracking these days?
They do (e.g. Vive Pro Eye, HP Reverb G2 Omnicept Edition). And checking that on google is much faster than writing this sentence.
They could do statistical gaze tracking on a 2D version of the movie scene using a small test audience, and then render the 3D scene based on the results.
I happen to also struggle with human face recognition, so it's likely I'm just wired different.
Since the field of view is limited in today's VR displays anyhow, you don't lose much from the foveated rendering: if you look to the side of display, much of your own field of view is now black, which is more limiting than the fact that the few pixels that you still see are rendered in a blocky fashion.
One nice benefit of the Quest 2 is that it can run Quest 2 games without be it.
It's still a feature in the Oculus runtime but they only recommend it in specific situations now, and they try and manage the quality outside of the game because way too many developers just cranked it to Max and it made their games look terrible.
You can easily tell when games have it cranked up and the quality drop off at the edge is very severe. Perf improvements are only moderate.
Eye tracking foveated rendering is much more useful but a bit of a pipe dream right now and still only gives a max improvement factor of 2-5x or something like that.
https://venturebeat.com/2019/05/23/playstations-dominic-mall...
The difference in sensation of change was palpable.
Even though there's only 24 images on the film per second, this triple exposure results in a smoother experience for the viewer. This is due to how critical flicker fusion works (a.k.a. persistence of vision). [1] (Though another reason for doing it is simply so that the film won't be burnt, since those xenon gas projection lamps run very hot.) See also beta movement and phi phenomenon. [2]
These physiological phenomena are of course very important to consider when making VR devices and games.
In order to create the effect mechanically in the film projector, each frame has to be stationary before the shutter is opened. If not, all you'd see is a blur. Thus, when the film is pulled forward, the shutter blocks the light from projecting the image onto the screen. In fact, during some half of the movie, people are actually sitting is pitch black darkness. Think about that next time you go to the movies! ;)
Please note that peripheral vision has a higher sensitivity to flicker than foveal vision. I'm unsure how interlacing affects that, though, if applickable. This effect might also be different on video systems where various forms of interlacing may or may not be used.
[1]: https://en.wikipedia.org/wiki/Flicker_fusion_threshold
[2]: https://en.wikipedia.org/wiki/Beta_movement
(Please excuse my lack of sourcing. Most of this is off the back of my head, and from books I am no longer in posission of... But the Wikipedia links should give you a good start.)
Ironically, it sounds like a complement to saccades! Where saccades involve shutting off vision when your eye moves, this involves shutting off projection when the image moves.
MEMS kHz eye tracking enables even predicting where a saccade will land 20+ ms before: https://www.youtube.com/watch?v=JEg4l5KuQgI&t=452 (AdHawk @ AWE 2018).
We have had machines that do what you describe for decades.
Here's Daniel Dennett in Consciousness Explained (1992)
When your eyes dart about in saccades, the muscular contractions that cause the eyeballs to rotate are ballistic actions: Your fixation points are unguided missiles whose trajectories at lift-off determine where and when they will hit ground zero at a new target ...
Amazingly, a computer equipped with an automatic eye-tracker can detect and analyse the lift-off in the first few milliseconds of a saccade, calculate where ground zero will be, and before the saccade is over, erase the word on the screen at ground zero and replace it with a different word of the same length. What do you see? Just the new word, and with no sense at all of anything having been changed. As you peruse the text on the screen, it seems to you for all the world as stable as if the words were carved in marble, but to another person reading the same text over your shoulder (and saccading to a different drummer) the screen is aquiver with changes.
The effect is overpowering. When I first encountered an eye-tracker experiment, and saw how oblivious subjects were (apparently) to the changes flickering on the screen, I asked if I could be a subject. I wanted to see for myself. I was seated at the apparatus, and my head was immobilized by having me bite on a "bite bar". This makes the job easier for the eye-tracker, which bounces an unoticeable beam of light off the lens of the subject's eye, and analyzes the return to detect any motion of the eye. While I waited for the experimenters to turn on the apparatus, I read the text on the screen. I waited, and waited, eager for the trails to begin. I got impatient. "Why dont you turn it on?" I asked. "It is on," they replied.
I didn't mean "random" as in governed by an RNG, just "random" as in we can't obviously predict their pattern. The opposite of that would be being able to tell when they'll happen and where they'll end before they start, and/or make them happen on demand and land on desired target.
> We have had machines that do what you describe for decades.
I didn't knew that. Thanks for citation. This leads me to ask: so why aren't we employing this for VR?
https://mailchi.mp/b4c8b26e025d/blindsight-project-update-se...
It should be feasible to detect saccades instead of predicting them, and render new pixels more quickly than the eye can move. But it is right at the edge of what's possible, and really needs more reliable eye tracking than is generally available today, along with purpose built rendering techniques and higher resolution + higher field of view displays.
GPUs do have real branching and even loops these days, you just can't have divergent branching within a shader group.
I'm not sure how efficient it is to splatter individual pixels like in the mask used in the article since adjacent pixels will likely need to be evaluated anyway if the shader uses derivatives. Too bad that the author didn't bother to include any performance numbers.
That GIFs autoplay is unequivocally a bug. Just a well-entrenched one that would take a lot of effort to fix, and no one capable of fixing it seems to quite think it’s worth it.
There are multiple reasons I might not want autoplay: bandwidth or performance concerns, accessibility problems from motion, or even just the simple “if you autoplay, then when I reach it it won’t be at the start of the video any more”.
If we had access to the internals, we could determine the visual salience of every object[1] and move the sampling pattern closer to the most salient one. Since that object is more likely to attract the viewer's attention, it would focus the rendering on those parts of the scene that the viewer actually cares about.
[1] http://doras.dcu.ie/16232/1/A_False_Colouring_Real_Time_Visu...
Maybe "which 3/4 of the laptop screen needn't be rendered/updated fully"? Or "unclutter the desktop - only reveal the clock when it's looked at"? Or "they're looking at the corner - show the burger menu"? Though smoothed face tracking enables far higher precision pointing than ad hoc eye tracking. This[2] fast face tracking went by recently and looked interesting.
[1] https://ai.googleblog.com/2020/08/mediapipe-iris-real-time-i... [2] https://news.ycombinator.com/item?id=24332939
Seems like it would be very interesting to see how changing the sample positions per frame and some sort of technique ala temporal AA would do.
TAA implies a velocity buffer, and that velocity buffer may cause you to sample outside the fovea (of a previous frame), for which you will not have enough samples to do a decent reconstruction of the image.
I've checked a few browsers on desktop and it seems to be no problem at all there, with the added benefit that most (or all?) modern browsers let's you easily change the font size that is used with Ctrl +/- (or Cmd) on almost all sites. For instance, I have the fonts scaled to 125% here on HN since I find them a bit small normally. (You can do that on mobile as well, but I'm not sure if that works on a per-site basis like how it works on desktop.)
1000px wide browser window: https://i.imgur.com/HQHkRED.png
@media (max-width: 1000px) {
body {
font-size: 250%;
line-height: 175%;
I think this is meant to make it more readable on small devices when in portrait mode, but really the text is too big even for that.