Disney rendered its new animated film on a 55,000-core supercomputer
engadget.com
engadget.com
That means almost half a year (~172 days) was spent rendering the film on this supercomputer. The listed running time of 108 minutes * 60 seconds * 24 frames per second = 155,520 frames in the film, giving us an average render time of 1,221 compute hours per frame, or a rendering speed of 2.27e^-10 FPS. Which means that, if Moore's Law continues to hold, in 26.7 years or so we'll have a super computer that could render this film in realtime at 24FPS, and that's just neat.
Congrats to everyone who worked on this; really impressive technical achievement.
So, he may be wrong about what makes this happen, in essence he's probably correct. In the time frame he outlined, I expect both Ray Tracing and Radiosity solutions on both hardware and software to match and possibly exceed what he has outlined in terms of capability.
You can massively parallelize rendering a movie in advance because you can do each frame on its own CPU. Rendering in real-time is much less easy to extract this kind of parallelism from, particularly if you have hard real time constraints.
If we can make higher pressure capable small pipes we could run then all faster. I bet there's billions being invested in solving this problem, but that doesn't mean it's solvable.
Of course, that's probably a bit of movie magic; in the real world there's surely some iteration.
So while it's great for placing objects in the scene and getting the camera position correct, that's about it. There has been a move to using progressive raytracing more and more over the past year or so, but it's still very early days with that.
On top of that, as previs is almost always done early (before most asset generation - modelling, texturing, maybe animation too) is done, it doesn't always give you all the info you need - i.e. it's often the case with mechanical hero objects in the scene (i.e. aircraft, robots, weapons, etc) that the basic geometry used for previs isn't good enough, and later on in real lookdev/lighting (or sometimes even at animation stage) serious issues are found - i.e. the previs has a robot moving through a street, but when animation actually start trying to rig the robot to get it moving realistically, they find for the arms to swing, it can't actually fit in the street.
So there's a huge amount of iteration that works up-and-down the pipeline, sometimes causing lots of different parts of it to re-do work.
Generally it goes: Previs -> Modelling / Texturing (linked, as you need UVs) -> Layout -> Animation - > Lookdev / Lighting -> Compositing.
Any non-trivial change before Lookdev / Lighting will trigger new renders having to be done for some layers / scenes.
So basically the impression from the DVDs is a simplistic fantasy. While the narration says 'it's such a wonderful place' the people working there until 9 every night are thinking 'why is everything broken'.
TV work, on the other hand, can't deal with that kind of time burden, so it's shifting to more novel solution which lately involves near real-time rendering (see element3d for example).
Blinns law is completely true. The limiting factor on render times is the impatience and work schedule of the people submitting the jobs.
In feature animation things are changing and being adjusted constantly. Each shot is worked on by many people across multiple disciplines, all of whom do many, many iterations to hone in in the final shot. So the entire production pipeline is designed to make handling changes as efficient and easy as possible. Look at the amount of processing resources disney used for this movie.. I can guarantee you they are cutting all the corners they can and maximizing efficiency wherever possible. Doing a single monolithic render every time something changed would be extremely inefficient.
That said, there are a few studios that do produce complete final frames in-render. Usually places that write their own renderers in-house and therefore have a sort of macho academic attachment to showing off how much their renderer can achieve out of the box. Blue Sky is this way because their whole pipeline is designed around their cg studio renderer. Anecdotally I have heard that Pixar did everything in renderman for a long time but more recently they have started using more of a compositing workflow since it's so much more efficient. So it's possible that disney may be this way now since they have this fancy new hyperion renderer they wrote, but I certainly wouldn't bet on final frames all being done in-render.
So for example if your scene has objects A, B and C, you might break it up into a layer for each. Then to render the A layer, your primary rays would intersect with object A only. But any secondary/indirect rays would intersect all of A, B and C, so you'd still get the correct global illumination on object A. Breaking it up this way just makes things easier to adjust in compositing. Also the indirect for each layer often will be rendered as a separate pass entirely so it can be dialed in comp independently, or regenerated if the scene geometry changes.
It's also important to understand that almost all the lighting in feature rendering is incredibly faked and not physically-correct at all. This is particularly true in animated movies. It's just important that the end product looks plausible, not that it's academically correct.
Global illumination helps add a bit of realism and nuance, but its contribution to the final image is pretty subtle, especially between objects that are far apart. Nobody would ever notice if the indirect lighting wasn't perfectly correct in all but the most close-together objects. So a lot of the time it won't be re-rendered if the scene changes slightly, as long as it still looks decent.
Without global illumination, to get the illusion of light bouncing between objects lighters would have to place spot lights for every bounce they want to fake. Like a ground light pointing upwards below a character to fake light bouncing off the ground etc. People really used to do this, and it's a real pain. This is the "time consuming manual lightning" they talk about. Global illumination tools make that process automatic and are able to get much richer light interactions than anyone could set up by hand. But they're still just tools. And like any tool, artists will break them apart and use them in whatever hacky way they need to get a shot to look right. Getting the exactly correct solution to the rendering equation is just not important.
But without a source, I'm not going to stand by that...
The latency between the sites would kill its linpack performance. It's a "supercomputer" in the same sense that seti@home is a supercomputer -- it's really just a few thousand machines working on an embarrassingly parallel problem.
I would imagine that right now studios are looking at 10gb ethernet, especially since SSDs should be making their way into the enterprise level gradually. Many times reading a scene goes at 50MBs.
It helps to realize that these studios are not flush with cash. They also might not always spend money in the best places (although usually they know what they are doing or they go out of business). Articles paint them to be super high tech, but really everything is built out of commodity hardware except for the backbone disks and routers.
Also upgrading to 10Gb means upgrading thousands of boxes, not just a couple. It could also mean running new cable but you would probably know more about that.
Take the infrastructure of an Amazon or a Google and its not really a material chunk of their resources. In fact, Amazon could no doubt offer an 'Elastic Render' service ala EC2 and really help a lot of CGI companies become profitable. The thing that kills those companies is keeping all the hardware after they don't need it any more. If you don't store it properly it becomes worthless, if you store it too long it becomes worthless, if you leave powered up and running it sucks money out of your account long after the checks from the projects come in.
http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using_clu...
Running Renderman [1], Blender [2], or your own in-house/custom rendering engine in AWS should be trivial at this point (using S3 as a cache for staging before processing).
Now, if you meant that should wrap a nice API around it like Elastic Transcoding (which is still terrible in my opinion compared to services built for transcoding like encoding.com), my hunch would be that the market isn't big enough for that sort of effort.
Dozens of companies exist that have massive render farms and they rent them out to shops working on movies and commercials, but most of the major players have large dedicated in-house infrastructure.
I don't understand how people make assumptions like this, unless they work at AWS at see the customer account info and real world workload data.
http://www.sidefx.com/index.php?option=com_cloud&Itemid=213&...
AFAIK they won't even let you spin up anything other than one or two instance sizes they kinda got working. The pricing was ridiculous, $6+/hr if I remember correctly.
[0] http://techreport.com/news/25615/amazon-web-services-adopts-... [1] http://www.pcworld.com/article/2599080/google-buys-zync-for-...
tl;dr: I doubt the engineers building the cluster didn't think of that.
Also, it's not really practical to use amazon-like service for production previews if you don't have a fat pipe running to it, so there will always be a need either for a small local farm or high end gpus or both.
(I have absolutely no idea, and the advert..cougharticle is light on technical content, but the fact that they didn't use them makes me suspect they couldn't use them.)
GPU's still don't cut it for high-end VFX/Animation rendering, as you're always memory / IO bound, e.g. reading in +20GB geometry and >200GB textures for an average large scene. That doesn't fit on a GPU and copying stuff on-and-off over PCIE is really slow.
Generally you get the geometry in memory (so 20GB in memory, + stuff like acceleration structures) and page the textures with 10/20 GB texture cache sizes. A dual high-end Xeon can almost compete with a K6000 at raytracing, and copying stuff in and out of memory is much faster. Then you're just limited by network bandwidth / latency.
Which makes the renderprocess especially slow, because even though you can reuse some parts of the process, some other parts can’t be reused.
1. Films are 24fps. 48fps is exotic and extremely rare.
2. Some stereo is planned to be done with conversion, some films are a hybrid between conversion and actual renders for each eye. If this movie is stereo it was very likely done by rendering each eye, as that has been the approach for their stereo films after Meet The Robinsons.
3. Films are still rendered at 2k. 4k is still extremely rare.
4. Typical frame times are probably between 2 to 10 hours on 8 cores. Many passes go into a typical frame, but usually only a couple are heavy (many hours).
5. Amazon and other clusters actually are being used and investigated by CG companies to be used during peak times (the last 6-8 weeks of a show). Profitability is not something that comes down to rendering however. It is an expense, but the vast majority of the money goes to pay the hundreds of people who work at the studio.
6. Hardware is not stored. It is on the farm and switched on, or decommissioned.
The factoid I like to use is: "Toy Story 3 took on average 7 hours to render 1 frame (24 fps), and at most 39 hours." (source [1]) That means if we wanted to render the next frame dynamically based on user input, the user would have to wait on average 7 hours to see the next frame, vs 16.666 ms that we shoot for in real time rendering. If one frame takes on average 7 hours, and there's 24 frames per second, then it take 168 hours or 1 week to render 1 second of video. Better get it right the first time (or have a super computer cluster to do all of that math)!
Many corners are cut to achieve something of lesser quality than can be achieved via pre-rendering; we call them "approximations."
[0] http://nickdesaulniers.github.io/RawWebGL/#/28 [1] http://www.wired.com/2010/05/process_pixar/all/
Here's a technical writeup that was on HN earlier this week - revolutionary stuff.
http://www.fxguide.com/featured/disneys-new-production-rende...
In fact, they are far from being pioneers in the field (right now raytracing is widely used in production rendering), but it's a big step for Disney and they have built their raytracer from ground, allowing them to implement it with some clever tricks.
Per frame.
Parallelize per frame and you can only have a few computers. They actually parallelize per pixel (see my other reply).
Rendering a single frame across multiple machines sounds wasteful. They would have to load the exact same textures and models for a single frame across all of them. When batch rendering, it would be more efficient to do that work just once per frame.
By the way, if you haven't heard of Blinn's Law [0], you may find it interesting.
[0] http://blog.boxxtech.com/2013/07/15/blinns-law-and-the-parad...
On an individual box frames are split up into tiles and the tiles are rendered on individual cpu cores.
The unit of parallelization for renders is always the frame. On a machine with multiple cores each core will render tiles in parallel, but across machines jobs are split up by frame. This is both because it's the simplest type of distribution for people to understand (machine goes bad? You lose one frame, not arbitrary pixels in an image). But also because it's most efficient for a single machine to read all the data for a single frame instead of multiple machines requesting the same data repeatedly across the network. I think people tend to underestimate just how massive the geometry and scene description files are for a typical feature and how much of the work involves managing the storage and network efficiency.
What if a frame, or several frames or small sequence of seconds is requested to be previewed?
Possibly it's more dynamic, and not just set up one way.
It's possible to parallelise this across several processes on different machines, but this is obviously less efficient because you'll have to do all the render startup tasks (reading resource files, building acceleration structures) multiple times.
In raytracing you render in reverse starting from the "eye", not forward from the lights, so each pixel of each frame can be parallelized. (With perhaps a small amount of mixing at the end to handle aliasing/quantization errors.)