RTX3080 TensorFlow and NAMD Performance on Linux
pugetsystems.com
pugetsystems.com
RTX 3080 - ~30 TFLOPS, 760 GB/s
So for most things compute you'd expect anywhere from 70 % to 200 % performance increase. Note the significant increase in operational intensity in this generation, to about 160 FLOps per Float Load/Store (up from ~90). The fact that we're not seeing 200 % increase more widely can point at things like this being a problem for existing applications (so RTX 3080 is _even more_ memory constrained as previous cards [1]), and perhaps also some applications struggling feeding enough work items / scheduling issues in general.
[1] Alternative view: Even more operations you can do for free when you have to process a given buffer anyway!
Jesus.
I think I know what Jesus In A Leather Jacket has planned for next year though.
It's fun to compare it with the historical top500 list: https://www.top500.org/statistics/perfdevel/
If it's effectively a paper launch because they aren't available in commercial quantities. Much like Intel's overly binned high end chips.
Some people are calling this rather scummy behavior on nvidias part.
"Nvidia has the cheapest card that this performance level ever!" is what the headlines will say. But if only a few hundred/thousand of these cards show up at this price over the next 6 months or so, and the AIB models go for $100+ more is really more of a way of manipulating headlines with no desire to follow up on the savings.
Did they buy? At what price?
It's still not cheap but who says moore's law is dead (I know its apples and oranges, but 60% gen-to-gen performance is great). The first games I played were already 3D and looked ok, but the idea that we'll probably be playing movie-quality raytracing within a decade is really something to look forward to.
Not to mention that scaling is going to be a pain.
I think it's more that if you can render 8K at 60fps, you can probably render 4K at 180fps or more which gives you a lot of wiggle room on a 60hz or even 120hz display.
Until the angular resolution of a pixel is good enough an inch from your eye there will be increasingly higher resolutions in the HMDs
I tried playing a few games at 8k (just for fun) and found that it simply wasn't worth the frame rate hit, or really even that noticeable an improvement with the assets in the games I tried.
One nice property of 8k resolutions is that they have an integer scaling factor from both 4k (2x) and 1440p (3x), so if you have an 8k monitor you can play games at either of those resolutions with high quality scaling.
I had a 4k 32" monitor at work and found that it simply didn't give sharp, high resolution text, driven by either a Macbook or by a Linux box. And you wouldn't really expect that, either: a 4k 32" monitor is only ~140dpi, which is only marginally higher resolution than the ~100dpi screens we had for many years.
I think the best point of comparison for the Dell UP3218k monitor which I have is a Retina Macbook Pro screen: subjectively, it's a similar experience in terms of text sharpness and legibility (>200 DPI, glossy), just in a 32" form factor.
I suspect 8k at 32" is actually a bit higher resolution than necessary (~280dpi), but there's nothing else on the >30" high resolution monitor market other than Apple's 6k display, which is significantly more expensive.
Be aware that there's no Mac OS support for 8k displays, but Linux and Windows on a desktop with a reasonably modern NVidia GPU work great.
NVIDIA is using Samsung's 8nm process, which has about 60 million transistors per square millimetre (MTr/mm^2).
That's not cutting edge! The crown is currently held by TSCM's 5nm process, at 173 MTr/mm^2.
Some time next year, TSMC is starting "risk production" of their 3 nm process, which is expected to hit about 300 MTr/mm^2. That's a solid FIVE TIMES higher density than the process used for the RTX 30xx series.
Unlike general-purpose CPUs, where transistor density does not linearly translate to real-world performance, GPUs are designed for embarrassingly parallel problems and have nearly linear scaling. More transistors equals more "CUDA cores" equals more performance.
The only thing holding back GPU performance is memory bandwidth. Current-gen consumer cards are just shy of 1 TB/s of memory bandwidth, but to get 5x performance, they would need 5 TB/s memory throughput to match. That's... difficult. Even with HBM2E, you'd need to stack a bunch of them to get near that.
But yeah. 8K gaming is crazy. Real time raytracing was an utter fantasy just a few years ago, and I just played through Control at 60fps and it was a visual feast.
I grew up in an era where wire frame 3D graphics took seconds to redraw the screen. I used keyboard macros to control a CAD program because it had no hope of keeping up with mouse movements.
My unborn son is going to grow up to play in a world of 8K raytracing as standard, with visuals better than Pixar movies of just a few years ago. That blows my mind.
We still don't have real time ray-tracing. Even the demos that only use ray-tracing are throwing the strict minimum of rays and they apply a series of complex filters (using machine learning) to remove the artifacts and the noise.
The RTX cards dedicate something like 30% of the silicon and power "budget" to neural-nets.
This isn't something designed to appease only the ML crowd, it's used for gaming. The graphics you see is 30% ML noise reduction and upscaling.
I used to think of "AI accelerators" as some sort of gimmick, one-trick ponies like Apple's face recognition. Useful for a handful of apps, a few seconds at a time.
But no, in the RTX series of cards, the "AI stuff" is drawing nearly 100 watts and doing real work, making ray tracing viable and making 1080p look better than 4K.
I assume by the end of the century we will end up with some kind of physically based rendering at a near photonic level of detail.
Both Renderman and Hyperion have denoisers that Pixar and Disney use on feature films.
> And because we could not afford to render images to convergence, we needed to develop a robust denoising solution
> The denoiser is used on most production shots, and is run automatically unless disabled
https://www.yiningkarlli.com/projects/hyperiondesign/hyperio...
(The fact that Hyperion uses a denoiser does not mean that it isn't rendering via path tracing. Similarly, the fact that real-time rendering uses a denoiser does not mean it isn't rendering via path tracing. dealing with noise is the name of the game, and machine learning isn't somehow "out of bounds")
Exactly. This is why I'm not too worried about Intel not being on the smallest current-gen process, even though this AMD fans are jumping up and down about this.
The reason lower nm is better, beyond cost savings, is heat output and power consumption. If heat becomes a problem, the chip has to limit itself, as we're seeing with laptops. Nvidia's new cooling solution seems to fit the bill just fine, so 8nm is no problem.
I suspect the next generation is going to be on 7nm, is going to be a bit faster and will consume a fair bit less power, which will be nice, especially if you plan on training neural nets all day.
Not sure how you'd fit 20 stacks on an interposer. I can't solve all your problems.
Cooling and getting enough power to the chips is going to very challenging. The 3080 already needed a redesign of its cooling solution and power delivery mechanism. At 5nm or 3nm, things are going to be a lot more difficult.
The problem is you can't use that law to increase speed of CPUs. You don't need more transistors, you need faster transistors and Moore's law does not help with that. We had 4GHz 15 years ago, most of CPUs still work under 4 GHz nowadays. But with more transistors you can implement some common operations in hardware (like crypto, vector operations). Also you can just increase core count. Or you can put energy-efficient core along with energy-hungry core. And those things happen with CPUs. Unfortunately many workloads are still single-thread capped.
GPU on the other side is inherently multi-threaded. You need 8k? You just need 4x transistors compared to 4k resolution. So GPUs will evolve even further and there's no limit, at least until we hit transistor wall.
Its funny to me that some people complain about nvidia price gouging etc, but if the sales of 3080 are any indication then based on basic supply&demand idea if anything they are massively underpriced. Lowering the prices would seem to just mean that more people would be disappointed from not actually being able to get them because the supply is so limited. The little I know about chip manufacturing also suggests that its likely that fabs have been producing ga102s at full capacity and its not like you can just get another fab to produce more of them in blink of an eye, so availability is probably really just constrained by production.
As a complete layperson when it comes to economics, I’m curious if someone has extensively analyzed ‘market value’ vs MSRP over the lifecycle of a product rollout and come up with a workable formula for calculating the optimal price.
Another thing to consider is that the number of pixels pushed to your screen is one thing. Another thing is that you will need higher quality assets with more detail too.
> Ampere's benefit is that it can deal with dense and sparse matrices differently. Its cores are twice as fast as Turing's for dense matrix and four times as quick for sparse matrix that have all the needless weights removed. The upshot, per SM, is dense processing at the same speed - it has half the cores, remember - and twice the overall throughput for sparse processing.
https://hexus.net/tech/reviews/graphics/145342-nvidia-geforc...
"Sparsity is possible in deep learning because the importance of individual weights evolves during the learning process, and by the end of network training, only a subset of weights have acquired a meaningful purpose in determining the learned output. The remaining weights are no longer needed.
Fine grained structured sparsity imposes a constraint on the allowed sparsity pattern, making it more efficient for hardware to do the necessary alignment of input operands. Because deep learning networks are able to adapt weights during the training process based on training feedback, NVIDIA engineers have found in general that the structure constraint does not impact the accuracy of the trained network for inferencing. This enables inferencing acceleration with sparsity."
So the idea seems to be that at the end of training, there's fine tuning that can be done to figure out which weights can be zeroed out without significantly impacting prediction accuracy, and then you can accelerate inferences with sparse matrix multiplication. They consider training acceleration with sparse matrices an "active research area."
I could see it being nice for the sake of running large language models on consumer, or really cool for the few edge computing applications that can actually demand and power conventional GPUs (e.g. self-driving cars.) It's probably not a great boon to the researcher who wants to reduce their iteration timeline though.
[0] https://developer.nvidia.com/blog/nvidia-ampere-architecture...
https://gist.github.com/Chick3nman/bb22b28ec4ddec0cb5f59df97...
Heres a list for 2080ti for comparison.
https://gist.github.com/binary1985/c8153c8ec44595fdabbf03157...
(The older numbers looks more inline with what I remember and here is an alternative benchmark for RTX Titan showed similar: https://lambdalabs.com/blog/titan-rtx-tensorflow-benchmarks/)
"The SOFTWARE is not licensed for datacenter deployment, except that blockchain processing in a datacenter is permitted."
https://www.nvidia.com/en-us/drivers/geforce-license/
This doesn't preclude use in a workstation, but they don't want you building a DGX A100 competitor using RTX 3080 or 3090 cards.