Nvidia Tesla Supercomputer (960 cores)
nvidia.co.uk
nvidia.co.uk
Visual Effects production often needs lots of computing resources for rendering. This is a task that can be heavily parallelized. Usually this is done on a frame/per core basis, with memory being the limiting factor. Now I've started using EC2 recently which made me realize a different approach to rendering tasks.
Normally a facility has a fixed number of CPUs available to run jobs. So for example 100 cpus. Let's say the average number of frames to render is usually around 200. So if each frame takes an hour to render and your's is the only job on the queue then it will take 2 hours to see the complete shot. And it's much worse in the real world with many artists working on many shots. Whereas using a resource like EC2 you can always render your entire shot in 1 hour regardless of the number of frames. You're only limited by the time it takes 1 frame to render, and the cost is the same if you use 1 cpu or 200.
In other words one can trade depth for breadth. Now this may seem obvious to anyone familiar with EC2, but for me it means that tools like Tesla which where once in high demand for this type of work are now much less valuable. I would expect that it's price/performance is much better then EC2, but where is the cutoff? How many hours do you have to run that Tesla to come out ahead? And if you're running 1 Tesla for that many hours, might it not be worth a premium to get your answer sooner by running more massively parallel (but for less wall clock time) on an EC2 like service?
I suspect these types of tools becoming only relevant to real time applications (due to reduced latency vs. EC2), and or nonstop computing (assuming there's a big price/performance win).
Put simply is that ray-tracing is a per-screen pixel ray which is quite good at glossy surfaces with predictable lighting and plastic appearance. Radiosity, on the other hand, simulates actual light photons/waves. It is slower and more computational intensive, but it produces far more realistic results. In particular, it is good at lighting/shadowing, and is more directly applicable to rendering non-plastic surfaces, including sub-surface scattering (like flesh or hair).
All that said, REAL modern renderers, are wacky hybrid of every technique :-P
The primary trick used to parallelize radiosity is to partition the scene spatially. Treat the virtual polygons which separate the rooms as light absorbers. When the room is finished with all of the available light, send the light absorber across the bus as light emitters to the other processes rendering the other rooms. Iterate back and forth until the light absorbers absorb under a threshold of light.
Those 960 ALUs are just the main FP32 and INT32 hardware, too. There are others (FP32 MUL and special function units, and an FP64 ALU too, per core).
It's pretty nice. As nice as C can get, that is... The biggest hurdle for me was thinking data-parallel as opposed to sequentially or even task-parallel.
When precision is needed CUDA is much less useful, say you're running 10^10 simulations then with single FP precision you will only have a result accurate to 5 significant figures.
What I'm interested in is seeing how OpenCL plays out, as Khronos seems to have the entire industry behind them (with the exception of Microsoft).
Optical Character Recognition could be done, by for instance, recognising each letter on a different core.
(just kidding)
I'm gonna love one of these babies for my AI work!