Behind the Tech with John Carmack: 5k Immersive Video
developer.oculus.com
developer.oculus.com
It's been five years since the Oculus Rift DK1 appeared, and there's still no killer app. Game developers are pulling back from VR.[1][2] The VR virtual worlds are a disaster. Sansar has about 50 (fifty) concurrent users. SineSpace and High Fidelity have similar numbers.
VR goggles are cool for about an hour, and in the closet in a month.
The technology is coming along just fine, but that's not the problem.
[1] https://www.gq.com/story/is-vr-gaming-over-before-it-even-st... [2] https://mashable.com/2018/01/24/virtual-reality-gaming-loser...
It's not (as it stands) going to be adopted as the media-platform of choice for consuming mass amounts of netflix and marvel movies, but there's way more niches to make it a viable go-to solution for many problems.
Take "Verification of Competence" activities used to on-board and train people who use specific plant/machinery. If you want to work on a mine driving a large truck, you'll usually have to do a virtual test to show your prospective employer that you can efficiently pilot $500,000 worth of machinery.
Health and safety training in complex plants. Run simulations on an immersive virtual model, rather than sitting at a computer screen.
Train driver training. Drivers need to do a particular number of trips to learn the route (accelerating and braking for turns, stations etc) this can be in field or virtual but required to ensure safety.
B2B++ B2C--
About 14 million VR headsets have been sold worldwide. That puts it about on par with the NES in the late 80s. I choose to forgive the VR market for not immediately being as successful as the mature gaming market. A lot of things were being done in the early 90s that showed just how little we really understood digital entertainment. It'll take tech improvement and product experimentation combined to really get VR to come into its own. That'll take continued investment on every front. We won't get there by magic.
Comparing them with the full lifecycle of the NES sales instead of its first 5 years is not a proper comparison.
Lots of enthusiasts did buy the Oculus DK1 and DK2 Rifts, but CV1 was the first time it was actually available in stores and marketed as being ready for consumer use.
At a more tactical level, such as beats, storyboarding, and even UX research...there's still a lot of experimentation when it comes to what works and what doesn't for storytellers. If you tell an experienced director to make an action movie, there's a lot of well understood techniques they can lean on to tell good story. So far, our experiences show that a lot of that flies out the window with VR unless you spend all your energy trying to encourage your user to frame the shot the same way a director wants you to. However, if you do that, you aren't really using VR anymore.
The problems that constrain story length are really a problem because you can only give story so much depth in 15 minutes...and if you don't have the opportunity to figure out how to tell a story with nuance, you won't learn how to get nuance right. VR has no nuance at the moment.
The indie game scene on Steam VR is thriving and these experiments in vr experiences will yield future content and entertainment breakthroughs. The groundwork, at least, has finally begun.
VR goggles are the opposite of laziness for me.
Too much effort required to set everything up. Why do that much work for little reward if I can just sit back in my couch with a remote and play a game?
As an example, I'm not sure if playing Civilization 6 in VR would be so earth shattering that I'd go through the effort (though who knows?). On the other hand, for a game like Elite: Dangerous, where essentially it's all about immersion in the experience of flying a spaceship, VR could / can be awesome. Even a simple head-tracking device can transform the setting from "hey, I'm playing a game here" to "hey, I'm sitting in a space ship here".
> Too much effort required to set everything up.
Which is why Oculus Go is being a (moderate) success. It's quick and easy to grab and use. It has its problems, sure, lack of 6dof in headset and controller, and lack of content, but we'll get there soon I think.
The Second Life continuations aren't the best examples to measure VR's traction. I'd look at user re-engagement of apps on the level of Bigscreen, Rec Room, Beat Saber, Google Earth. Those apps get people back into the headset and capable of logging 1000+ hour playtimes. Once high-end capabilities (6DoF head / hands) meets convenient form factor (standalone, wireless), I expect to see VR gain wider traction.
What a lot of people don't see are the iconoclastic milestones that are coming up, that will transform the entire domain in terms of market acceptance and saturation.
Foveated rendering is the first of those. Hand-finger tracking (e.g. a glove interface) is the second.
These technologies are not science fiction, we know they are coming out since lab prototypes already exist. Moreover, impact-wise, these technologies are not really incremental improvements but complete game changers (thus iconoclastic). The VR landscape in 5 years will look completely different to today.
I certainly don't want to be wearing any kind of apparatus that distinguishes me from a normal individual.
If we do have brain implants I don't want advertising and tracking.
I would pay for a direct internet connection to my head only if I know it is anonymous and I control it.
VR in its current state is pretty lame if you like the real world too... imho; and it isn't going to take increased resolution to fix it.
In any case, I think it's early to say what's a dead end. As you say, there are no killer apps yet. Even VR porn is still at a demo stage. It all might be a dead end, if you want to go on existence proof alone.
Honestly, I was expecting one of the VR platforms to attract simple applications/games. Resolution doesn't matter much if you're looking at stick figures.
If big, epic content is what it will take to get VR to that "real product" stage then I think the platforms will have to produce it. I could imagine "canned" content similar to old Imax movies working well, but that is not a 2 guys and a camera type of job. Who would put up that kind of money, for such a small audience, when all the first mover benefits go to the platforms.
Some are, others are going all in and it's just getting good. Compare Lone Echo to launch experiences, it's a completely different experience. Not to mention Fallout VR and Skyrim VR launched this year.
> The VR virtual worlds are a disaster
Go on YouTube and search for VR Chat, check out the number of videos and view counts. VR Might not be mainstream yet and I don't believe it will be until stand alone headsets drop but it's naive to write it off just yet.
I want something like that, I bought the old Xiaomi headset (basically a Google Cardboard) and I can't run Netflix on it (which was what I wanted them for :( )
The Standalone Xiaomi one does seem to run Netflix though :)
Why do you think that is an indicator of success? VR has plenty of hype; marketing is not the issue. There are plenty of people who are interested in watching a video on YouTube that won't buy a headset.
Because this is what people in the video game industry are literally looking at as a gauge of success.
Just look how much they court youtubers and streamers. Games are literally built around them in 2018.
(Note, I don’t like that this is true but the sad fact is that it is true)
I just don't want to wear VR stuff I'd rather get something like from lawnmower man or something without all this stuff strapped to me. Is anyone doing anything like that?
How are the videos stored? Is each gop a separate file or do you seek around in larger files? Is it feasible to stream over HTTP or is it only possible to play local videos for now?
I was originally going to put it into an mp4 file with the base stored first, so normal video players could at least play the low res version, but the Android MediaExtractor fails when presented with more than 10 tracks, so I just rolled my own trivial container file.
Peak bitrate for Henry is around 40 Mbps, so it wouldn't stream for most people. With some rearrangement of the file so each strip has a full gop continuous, instead of time interleaving all 11, the bitrate wold be cut in half, but it would still be a lot of fairly small requests, so it would call for pipelined HTTP2.
I had each gop in a separate file like HLS or DASH (except for the background which was a single file also containing the audio track). It's unwieldy but makes HTTP streaming a little simpler because you don't need range requests or an index.
Also, instead of bitstream hacking to stitch three strips into one, I encoded multiple strips into "pre-stitched" views. This means that every strip is encoded redundantly in multiple views, bloating the on-disk video size. But for streaming that only affects the server, not the client, and it's nice for the client to only download/decode one view at a time (plus the background) instead of three strips. Bitstream hacking to join the strips would definitely be better if it can work, though.
How does increasing the resolution cause aliasing? Shouldn't it be the opposite?
More resolution is going to be fine if the pixels are filtered well. If they aren't, that could certainly invite aliasing. Pixel filtering, especially in a scenario like this, is going to have a wide spectrum of techniques that can work to varying degrees, with trade offs of quality and speed.
> Image scaling can be interpreted as a form of image resampling or image reconstruction from the view of the Nyquist sampling theorem. According to the theorem, downsampling to a smaller image from a higher-resolution original can only be carried out after applying a suitable 2D anti-aliasing filter to prevent aliasing artifacts. The image is reduced to the information that can be carried by the smaller image.
- https://en.wikipedia.org/wiki/Image_scaling
In an ideal situation, your source resolution would match your target resolution. If it doesn't then you must expend resources scaling. VR can't leverage the 2D scalers found in most hardware since the source data is being mapped into 3D space. If you have to downscale your source then on top of the processing resources being used to scale in 3d space, you're also wasting bandwidth delivering those resources at a higher resolution.
High end mobile GPUs can handle 5k video at the expense of battery life but most GPUs can't.
Advocating for higher than necessary resolutions today based on future prospects is a bad gamble because consumer adoption is still in it's infancy. What if you make a poor user experience today in anticipation of a better experience tomorrow but instead kill the market so there is no tomorrow?
Video game developers and video content providers are both well versed in dynamically scaling to match the capabilities of the consumer.
I already read articles about the pannini projection, which is quite cool, but I guess occulus is facing the same issues and has similar solutions...
You can render a single quad and calculate UV in the pixel shader. Geometrically, this is the most accurate method.
You can do without any shaders, write a CPU code that builds a reasonably dense (about 8x8 pixels/triangle) static geosphere mesh with position and texcoord channels.
You can also build a reasonably dense grid mesh occupying just the screen with just the positions in clip space, and calculate UV channel in the VS based on the rotation.
GS will work too but I’m not sure it makes sense performance-wise. People normally use GS-es to implement small dynamic objects. For static geometry I won’t be surprised an immutable buffer will be faster, just because that’s what hardware & drivers are optimized for. See e.g. this http://www.joshbarczak.com/blog/?p=667
1. Cubemap (source) - 360° videos are often sourced from six different images stitched together.
2. Equirect (video) - the cubemap is transformed to "equirect" trivially with a fragment shader—which maps every output pixel to some input pixel.
3. Perspective (output) - Mapping equirect video to a perspective projection is also done with a fragment shader, just using a different transformation depending on the focal point of the user.
4. Pannini (alternative output) - You mentioned this projection, and it's an alternative to the "perspective" projection that allows a wider output FOV that minimizes distortion in the periphery.
I researched this a bit here: https://github.com/shaunlebron/blinky
Mandatory XKCD: [1]
Another (admittedly crazy) idea, for a setup with a lower-res version and a higher-res overlay, is trying to store the difference only, affording a (significant?) bitrate reduction for the high res "patches". This is very tricky to do in practice, though (needs larger range or losing the lsb; the codecs aren't really designed for this). I don't think it has ever been done.
There are also other interesting spherical tessellations that could be used - one is implemented in Google's S2 geometry library, the other, HEALPix, is used for astronomical datasets:
http://s2geometry.io/resources/earthcube http://iopscience.iop.org/article/10.1086/427976/pdf
Another idea is to take advantage of the fact head motions are really just a translation vector of the camera. There’s no need to send pixels that have just transposed locations unless they have changed in time.
If I was designing such a system I’d try to take advantage of the fact there isn’t a lot changing fundamentally in the scene when you move your head, and maintain some sort of state and only request chunks of pixels that are actually needed. You wouldn’t even have to use a traditional video codec as the preservation of state would be far more efficient than thinking about things in terms of flat pixels and video.
By "disparity map" are you thinking something like a heightmap applied to the scene facing the viewer and then you use that to skew things for each eye?
If so, how would that handle parts of the scene that are occluded/revealed to one eye but not the other?
My personal opinion is that true stereoscopic images feel better when there's enough detail; those occlusions do matter. For some imagery it doesn't matter as much though.
A three inch difference between two cameras producing simultaneous frames is similar to a three inch sideways step of one camera in time between two frames.
Ultimately what you want for VR is some kind of light field video compression. You can get a taste for what that would be like here, although it's mostly still images for now: https://store.steampowered.com/app/771310/Welcome_to_Light_F...
Secondly, the limiting factor described in the post is not space efficiency, it's decoding performance. It doesn't do much good to halve the amount of data required to represent a frame if it takes twice as long to reconstruct the raw pixels for display.
> partially-transparent or glossy surfaces.
It's all just RGB values; there is no gloss or transparency in an image. (Image layers can have transparency for compositing, but that's obviously something else.)
If audio encoding can have "joint stereo", why not visual coding.
Many areas of a stereo image are nearly identical, like the distant background.
The "disparity map" you suggest seems to exist in 3D Blu-rays (Multiview Video Coding), but there may be some technical limitations that make it unsuited for the Oculus Go.
I wonder if the problem isn't a misaligned paradigm for just what a codec does. Right now, codecs exist to take bytestreams and fill frames at a certain rate. They're a bitstream-to-framerate device.
What if instead of delivering to a frame during a certain time period, the codec had to instantly deliver whatever it could, but only geared to the actual acuity of the users retina. Seems like there would be less information, you would never have lag, and all optimization would be around the detail of data delivered, not framerate or screensize.
I also birthday parties and Bar Mitzvahs, folks. I'm here all week. (I figure it couldn't hurt to throw crazy ideas out there. Every now and then a crazy idea actually amounts to something. Coding is cool because it not only let you solve problems, sometimes it lets you change the universe the problem lives in. Good luck, John!)
I want to restate what I'm hearing so you can correct me. You are saying something like "Gee, if we could show any retinally-matched conic the user's pointing their eyeball at? To do that we'd have to have the entire scene rendered anyway"
I'm saying you can't do it now. It hurts when I do that. So stop doing that. Instead of trying to render stereo scenes quicker, try to render a moving conic real-time quicker. When the codec runs, it's optimizing areas of a frame changing over time. If instead it optimized possible vision movement paths, you'd end up solving the framerate and resolution problems. Then you could concentrate on optimizing the codec in a different way than people are currently trying to optimize it. It couldn't hurt, and there may be opportunities for consolidation if you look at the problem as visual-movement-path-rendering instead of frame rendering. I don't know. I just know it's a problem now. Set your constraints differently and optimize along a different line. Sometimes that works.
Does that explain the differences here? Or do you want me to start spec'ing out what I mean by retinal-path codec? At some point this will a bit over the top.
ADD: I'll say it a little differently. Our constraint now is "how fast can you render this frame". Codecs are built for rendering frames on screens where people watch movies and play games. But that's not the world we're trying to solve a problem in. They may be configured the wrong way. Instead, require that anything the eyeball is looking at should render consistently in, say 5ms.
This sounds like the same thing, until you realize that the eyeball can't look at everything at the same time. Different parts of the image are temporally separated. So if I update the image in back of your head once every 200ms? You're not going to know. As far you're concerned, it's all instant.
It becomes a different kind of problem.
I want to predict the temporal cost of focal path movement, then optimize the stream based on those temporal "funnels" I guess you'd call them. This is in opposition to looking at segments of the screen all equally and optimizing across frame changes. I don't care about frame changes. All I care about is how fast the eyeball can get from one spot to another -- and that's finite. If I can create the image I want instantly wherever the eyeball is looking, framerate and resolution issues are a non-sequiter. (This would also probably scale to more dense displays easier, but I'm just guessing)
This is assuming the videos are stitched pre-encode, of course... From the post, it almost sounds as if the idea would be to stitch independent H.264 streams into a new unified one using slice mangling + slices... which would be pretty crazy stuff.
(As a side note, it's a shame flexible macroblock ordering is only in the baseline and extended profiles... I still don't understand that decision at all.)
EDIT: Dawned on me that the hard part is on the client/decode side. D'oh.
If yes, would it be possible to render, for every second (half second or whatever resolution you like) a snippet which only contains the seek time until the next keyframe and than tell the webserver 'if someone seeks to time x, take this snippet and when done, jump to the original video'?
You could do that transparent with some fuse fs.
You mentioned the overhead for handling many streams being a problem. Can you go into more detail on the type of overhead you saw and do you feel this is something can can be addressed with better drivers or some application changes? Or do you think the only solution is packing multiple stripes into a single stream using the encoding bit manipulation you mentioned? It seems a shame for hardware that can support more streams leaves them inaccessible to real world applications.
If a 5120 pixel Radius is fine... then Radius of sphere = 2 pi * r, and Surface of sphere is 4 pi * r^2. So 5120 / 2pi =r substitute for r > 4 * (5120/2pi)^2 simplify > (5120)^2 /(pi) ~= we need ~1/3 aka 1/pi of 5280^2.
However, I suspect you actually want more than 5120 pixel radius.
Starting where? Turning where?
> However, if you tile like that with 5120 x 5120 you get a lot of wasted pixels at the poles.
You didn't mention tiling.
> If a 5120 pixel Radius is fine... then Radius of sphere = 2 pi * r, and Surface of sphere is 4/3 pi * r^2. so 5120 / 2pi =r => 4/3 * (5120/2pi)^2 => (5120)^2 /(3 pi) ~= we need ~10% of 5280^2.
More details on what you are trying to say at each step please.
As to tiling, I am saying if cut a sphere into 5120 slices, and the middle slice is fine with 5120 pixels. Then the poles are going to have ~1 pixel on them. How you tile them is really a question of tradeoffs.
> As to tiling, I am saying if cut a sphere into 5120 slices, and the middle slice is fine with 5120 pixels.
What kind of slices? Like sections of an orange? Like longitude?
You are off by a factor of 2 in your pixel calculation, because 5k x 5k is for a stereo pair of spheres. Equirect projections waste a fair amount, but compared to the 300% miss to get to 60 fps stereo, it isn't dominant.
Anything you can do about someone tilting their head? That feels like the biggest remaining immersion breaker with pre-rendered stereoscopic videos.
One would assume there would be a lot in common between stereo cameras' views, so presumably there are compression efficiencies to be found.
Isn't this back-to-front? Generally filtering is considered to work better in a linear colour space. Using a sRGB texture will convert each pixel to a linear colour space, before the reconstruction filtering is done, AFAIK.
VLC on the Note 8 (and the S8) can do 8K decoding at 48fps, as is demoed here: https://vimeo.com/254723180
We've spent quite a bit of time to get that right, though, and those are rare devices (and does not work with the Qcom version).
Maybe that's one video, with low bitrate (5Mbps), but it worked.
Edit: Thanks for the feedback, I guess it is perceptible. Nevertheless, I think my argument becomes valid at some N. Sure, N !== 60, but N = 144 or 120 may be more reasonable. I'm not too concerned with what N is, more so with the fact that "doubling the refresh rate" eventually becomes an act of futility.
But really, when you achieve 120Hz, it’s beautiful, it reminds me of when retina displays came out. We are a bit closer to realistic rendering.
Yes, what I was trying to get at is that just because the hardware is capable of 60 frames in a second, that doesn't mean the software was delivering 60 frames a second. The iPad Pro has a different processor than the iPad (A10X Fusion vs A10 Fusion), and in a lot of tests it's significantly faster.[1]
The iPad pro does have more pixels to push around, but that doesn't exactly negate the CPU difference, it just makes it more complicated to draw an actual comparison. That that's before we even get to the actual graphics processor, which itself could do a better job of offloading some processed to hardware (better OpenGL/Metal/whatever support). For all we know, you were seeing a average of 35 updated frames a second on the iPad, and you're now seeing an average of 55 updated frames on the iPad pro. In that case, the doubling of the screen refresh might help a little (in reducing noticeable laggy frames a bitas it can update between what would be frames at 60Hz), but it wouldn't be earth shattering. I doubt it's that bad, but as an example, this should show how a Hz rating on what a screen is capable of doesn't mean much.
The real benefit of higher screen refresh rates is to better support different lower native refresh rates. Much video content is at 24 FPS. A 30Hz or 60Hz screen can't represent that faithfully, and will need to double some frames. A 120Hz screen can perfectly represent 24 FPS content[2], and that's the real reason screens (and TVs) ship with that refresh rate. Different media (television, internet video, DVDs, Blu-Rays, video game systems, etc) all have different refresh rates they want to deliver.
1: https://www.notebookcheck.net/A10-Fusion-vs-A10X-Fusion_8178...
2: I'm ignoring that it's often actually 23.976 FPS or something.
Its painful to go back.
[0] https://www.pcgamer.com/how-many-frames-per-second-can-the-h...
[1] https://www.reddit.com/r/Competitiveoverwatch/comments/5mhqh...
It's an interpolation thing; it's much easier for your brain to do the tracking of an object when it smoothly moves around, instead of it having to do interp / extrapolation, and this should demonstrate why
Can I ask why you think it sounds like a placebo? I don't really see why there's any logic behind 144hz being the ceiling of how well your eyes can see, and 144hz -> 240hz is a big jump
144-200 was not very noticeable. 144-240 was absolutely noticeable. A couple of games that dual 980 ti's cant reach 240 in currently, hopefully Volta changes that. The monitor was also on sale, so I got it for a crisp $300 with no tax.
Using my iPhone immediately after makes the phone display feel cheap for a bit.
Even on your current 60Hz monitor, you can see how the image is blurry in the 60Hz band. On 144Hz I could see the pixel-level details (on 1080p) almost as clear as a motionless image.
In reality though, all LCD displays have persistence issues. (You end up seeing bits of the previous frame suspended over the current frame).
Higher quality monitors designed for higher display frequencies tend to also be tuned to have less persistence per frame. There are also tools/tricks some brands employ to remove / reduce the persistence (and blurring), which can be effective (or annoying, depending on how they do it) even at 60hz. There are also "120hz" and higher displays with such bad persistence issues that they are way worse off on the ufo test than a good low persistence 60hz display.
As an easy example. IPS panels tend to have a lot more trouble switching between frames quickly than VA or TN panels, so they also tend to exhibit more persistence issues per frame. This is then directly noticeable on the UFO test between these kinds of panels. IPS panels have of course been getting a lot better in this regard in recent years!
There has been a bunch of research involving VR induced nausia indicating that having a framerate of 90 or higher reduces the incident of nausia over 30 and 60.
For example: https://www.researchgate.net/publication/320742796_Measureme...
There is also a misconception of treating visible motion details and temporal artifacts as the same thing ("what human can see?"). There are diminishing returns around 100 FPS for motion details, but it's still far too low to eliminate artifacts like blurring or judder (in VR usually connected to vestibulo-ocular reflex). This is why current VR headsets already have effectively 300+Hz-like persistence. We may need something like 1000 FPS if not more to achieve clean vision at full persistence, so obviously strobing tricks are necessary to get around it. And you don't need a fighter jet pilot pilot to see it. Everyone can notice these problems.
The most important quality issue is for the image update to match the display strobe. That's why FreeSync/Gsync is such a big deal.
Doubtless there is some refresh rate where it ceases to matter, but the point here is that we're still struggling to push an acceptable frame rate in VR (and 4k, and other high resolution "formats") without making significant sacrifices in other aspects of the video quality. Assuming that image quality and available processing power both continue to increase, it will continue to be important to include framerate in the balance, i.e. we want to draw enough but not be wasteful.
I played breath of the wild and it was almost bad enough to keep me from playing. I think 60fps is fine for this, but in the future hope to see higher refresh rates as an option.
My main use case from VR has been escaping reality. Having multiple monitors floating in space allows me to prevent distractions. I fell asleep in VR once. I woke up looking at the stars, it was very confusing but interesting.
Almost everyone I've met or talked with that can notice the difference between those kind of framerates seems to either be diagnosed with it, or at the very least clearly show the associated traits.
I think most people I associate with are diagnosed ADD, but they're also high skill gamers.
That's not true for everyone. For me and a few others, there's even a clear difference between 120 and 144 Hz, and to 60Hz it a gulf. A Refresh rate of 60Hz in a high paced, close up, fight becomes a bit of a where's Waldo situation. It's as if my brain stop doing motion estimation because the frames doesn't really fit together, so I kind of stitch together the individual frames and try to guess what next 'slide' is going to show up.
I know there are a few other with the same issue, and anecdotally it might be a trait that occurs with some forms of ADHD/ADHD-PI. But I don't have a good reference on that, only experience and people I've met.
And how old you are (ducks) But seriously it's 2018! It's the distant future where Doom is run in emulators in browser windows.
It was influential enough that for many years, the genre we now call “first person shooters” was known as “Doom clones”.
Namely, that when he created Wolfenstein 3D (Doom’s predecessor), he basically only rendered the walls + sprites and not the floor or ceiling, because computers + graphics engines weren’t as capable at the time.
I would prefix that with "in its current state".