Nvidia Pascal GPU to Feature 17B Transistors and 32GB HBM2 VRAM
wccftech.com
wccftech.com
> The Pascal GPU would also introduce NVLINK which is the next generation Unified Virtual Memory link with Gen 2.0 Cache coherency features and 5 – 12 times the bandwidth of a regular PCIe connection. This will solve many of the bandwidth issues that high performance GPUs currently face.
To point out the obvious, this sounds like it could be fantastic for deep learning. Not just is the RAM big enough to hold a lot of current datasets, it'll help alleviate the latency bottlenecks in updating stuff.
IT departments aren't even buying Nvidia at all, they are buying Intel integrated.
The deep learning applications and those that want massive amount of RAM buy Teslas, which have exactly that. Titan X is a gaming card and it's marketed as such.
Soon headless, socketed solutions will be the preferred form factor for HPC. I image the desktop and server product lines will diverge at that point. It'll be curious to see what will happening to PC gaming at that point.
For reference: Titan Z DP FP: 2707 GFLOPS ATI 5870 (released sept 2009): 544 GFLOPS Titan X (current generation - NVidia recognized that gamers don't value DP performance): 192 GFLOPS
I sometimes wonder how it came to be that the Titan Z preceded the Titan X...
E: to clarify, review sites are aware that STEM people need double precision and 6/12GiB GPU memory for something
Maybe some music videos, or the intro of an HBO show..
Until then, enjoy your eyes and dog faces on family pictures.
Also I think there's a bit more than just ZOMGtrippy. The whole project points out that "if you're looking for it, you'll find it." The system amplifies noise until it sees something familiar. Understanding that humans do this too (to a point), is a critical point in comprehension and discussion of the world, opinions and people. But yeah, it will seem rather "iconic" and dated 10 years from now.
I'm thinking that using CNNs to generate or change images could be a huge game changing technology. It could potentially be bigger for hollywood than 3d graphics. In 10 years, we could see this stuff everywhere, albeit in a more developed state. It would be used to actually transform or enhance images in a more directed way.
That 32GB seems too high for cost constrained consumer market though - may be they will have a leaner variant for desktops/gaming.
But the Internet effect on the video card market is huge. There are hundreds of online benchmarks, huge fanboy communities, et cetera. The benchmarks magnify small differences causing a winner-take-all effect. The communities create a bandwagon effect. It's a wonder that AMD maintains the share that it does.
Yes, AMD is obviously #2, but it should be a close #2. In most markets close #2 is not a bad position to be in.
Zen definitely looks interesting, let's see if it's as big as a success as the K7 architecture
> AMD's continued existence as a serious CPU maker depends on it
I agree
How isn't it?
Okay, but they didn't used to be. They sold all their fabs. And when their reliance on third-party fabs meant they couldn't keep up with Intel, they did... nothing. Now, you could argue that they don't have the resources to do anything. But that doesn't mean it's not their fault that they didn't do something, it just means they couldn't.
If the fab sitting idle or production is bellow running cost then they are in trouble. Better to spend that money on R&D.
They were later compensated, but the damage done was incredibly destructive.
That's not quite true -- Intel aims for a big IPC (instructions/clock) improvement for each "tock" generation (Nehalem -> Sandy Bridge -> Haswell -> Skylake), and IIRC has pretty much delivered. Some benchmarks are really hard to push because they're memory/cache-miss bound (so it's really just about throwing in more memory channels and clocking them up), but a lot of things have gotten seriously better for tricky integer code, especially in e.g. branch prediction/uop cache in the frontend and available execution ports in the backend.
While the headline is hyperbole, the fact is that Intel has passively admitted that newer process sizes are taking more time to achieve.
It's some what asinine to reduce the complexity down to a single number but I've found that they are usually pretty reflective of what you find in practice.
You've overlooked that Maxwell / 980 Ti has a ton of overclocking headroom. OC to OC, you're looking at 20-30% performance difference at more commonly used resolutions like 1440p and 1080p. The gap narrows only at 4K.
Voltage is locked on Fury right now and AMD isn't saying why. The best OC I've seen on one is 10% higher clocks, with Maxwell 30% is common. Huge difference considering both cards are in the same price bracket.
Lastly pricing is also an issue with the Fury cards. Usually AMD undercut Nvidia, but this time the flagship matches the 980 Ti in cost. As the underdog, AMD is going to have a tougher time winning people over at the same price point. This happens in any market.
I like to support the underdog, but I couldn't this time given the large disparity in OC performance.
I didn't care either way until NVidia pulled the hairworks stunt, which was a pretty controversial move.
This thing is definitely not for the consumer market, much less the cost constrained one. This is a small dedicated number crunching machine that, for reasons unfathomable to me, can spit out rendered 3d environments with admirable speed.
I liked to joke that no serious computer has keyboard/mouse/video ports because no serious computer would be used like that. That assumption held well until the late 80's. ;-)
But no. For gaming, this is the superlative of overkill.
This article you just read is typical Nvidia "we are the best .. in 2 years, you just wait!!1". They produced same charts, slides and test results before releasing Tegra, Tegra2 and Tegra3. Every time it was supposed to revolutionize industry and beat competition, every time it released late and benched slower than products on the market.
My ZX81 only had 8192 bits of memory (1KB).
For the goal of full VR, computing tech has a long way to go
Seems like Moore's law is alive and well in the graphics/attached processor space.
Note, that both NVidia and AMD rely on TSMC to manufacture their chips, so they're completely constrained by TSMC's ability to implement new process nodes.
2012: GTX 680, 3.0 TFLOPS (~2.0 attainable)
2013: GTX Titan, 4.4 TFLOPS (~3.2 attainable)
2014: GTX 980, 4.6 TFLOPS
2015: GTX Titan X, 6.7 TFLOPS
Looks to me like they're doubling perf roughly every 2 years.
Meanwhle, my Core i7-5930k's SOL is <1/2 of 2011's GTX 580 at 672 GFLOPS and it still doesn't have fast approximate transcendentals. Skylake begins to fix this, but c'mon, GPUs have had these for almost a decade now...
GTX 680: 3.5B transistors
GTX Titan: 7.1B transistors
GTX 980: 5.2B transistors
GTX Titan X: 8B transistors
Core i7-5930k: 2.6B transistors
What the data above suggests to me is that relying solely on Moore's Law to predict performance is a fool's errand. Going forward, process transitions are obviously slowing down and IMO victory will go to those who make the best use of the available transistors. Just like programmers who make the best use of the caches and registers in these processors get dramatically better performance than those who can't be bothered to even think about such things.
Intel's business strategy of backwards-compatibility is a giant albatross for them here in that they spend a lot of transistors on this, but clearly otherwise profitable. In contrast, while GPUs are mostly backwards-compatible, they usually oops I meant nearly always oops I meant always need some refactoring to hit close to peak performance. But that usually leads to ~2x performance improvements per generation so far.
Whenever someone complains about having to do this I ask them if they prefer this over hand-coded assembler inner loops for maximally exploiting SSE/SSE2/SSE3/SSE4/AVX2/AVX512? Usually, I get some dismissive remark about leaving that to the compiler. Good luck with that plan IMO.
There are obvious downsides to the architecture, but the need to be backwards compatibility shouldn't hurt it too much.
GPU workloads are very different in that generally you don't have to look particularly hard to find a bunch of parallelism that you can exploit (if you did, your code would run terribly); so you can generally gain a load of performance by just scaling up your design.
CPUs are super restricted by the single threaded, branching nature of the code you run on them, and this is what makes CPU performance a little more nuanced, and not directly comparable.
A paper (http://www.ic.unicamp.br/~ra045840/cardoso2013wivosca.pdf) states that a mostly-microcode solution would still require 20% of the die area to be dedicated solely to microcode ROM.
I can't remember where I read it but something like 30+% of an Intel CPU die area/power consumption is due to the x86 ISA. Apparently the original Pentium CPU was 40% instruction decoding by die area. And the ISA has grown enormously since then.
Ironically, to really hit peak performance of a modern AVX2 or later CPU, you have to embrace many of the design principles that lead to efficient GPU code:
1. Multiple threads per core to make use of the dual vector units introduced in Haswell
2. SIMD-like thinking to remap tasks into the 8-way and soon to be 16-way vector units
3. Running multiple threads across multiple cores
4. Micromanaging the L1 cache and treating the AVX/SSE registers as L0 cache
Where the CPU prevails is for fundamentally serial algorithms that cannot be mapped into a SIMD implementation. Mike Acton's Data-Oriented Design covers this case nicely IMO.
Give that GPU a highly serialized workload and watch the actual performance take a nosedive. There's not much of a reason to compare a sniper rifle to a carpet bomb.
Sarcasm aside, this is probably purely business driven decision. The tech is there and it's nearly zero difference for the manufacturer to put either 8GB or 16GB dies on the board. It's the same as with SSD's - companies have to milk existing capacity tier to offer users next ones, otherwise they will have lower profits.
IMO, AMD pulled the trigger a tad too soon on their HBM cards, but to be fair their last line of R9 and R7 cards weren't interesting at least to me (and I have a Radeon HD 7870). So, they had to go first to get their customer base excited for the future.