Nvidia CEO Reveals New TITAN X at Stanford Deep Learning Meetup
blogs.nvidia.com
blogs.nvidia.com
Nvidia says memory bandwidth is "460GB/s," which will probably have the most impact on deep learning applications (lots of giant matrices must be fetched from memory, repeatedly, for extended periods of time). For comparison, the GTX 1080's memory bandwidth is quoted as "320GB/s." The new Titan X also has 3,584 CUDA cores (1.3x more than GTX 1080) and 12GB of RAM (1.5x more than GTX 1080).
We'll have to wait for benchmarks, but based on specs, this new Titan X looks like the best single GPU card you can buy for deep learning today. For certain deep learning applications, if properly configured, two GTX 1080's might outperform the Titan X and cost about the same, but that's not an apples-to-apples comparison.
A beefy desktop computer with four of these Titan X's will have 44 Teraflops of raw computing power, about "one fifth" of the raw computing power of the world's current 500th most powerful supercomputer.[1] While those 44 Teraflops are usable only for certain kinds of applications (involving 32-bit floating point linear algebra operations), the figure is still kind of incredible.
There have been calls that Moore's Law is dead for at least 10 years, and yet desktops keep catching up to supercomputing.
I know this increase is not in single-threaded general-purpose computing power in the fashion of the old gigahertz race... But on the other hand, the scope of what's considered "general purpose" keeps expanding too. Machine learning may be part of mainstream consumer applications in 5 years.
It doesn't seem likely since the Titan X pulls 250 watts. Maybe if we get asics for deep learning which consume less power; and a lot of research is done on deforesting trained networks so users can compute classifications more cost effectively, then sure.
Moore's Law was never about single core CPU performance. It is about the number of transistors in a single chip; this has always been its definition.
That's the definition I've always known. It has nothing to do with money or speed or performance, but usually these things are correlated.
In a particular area (eg, single core cpu performance) it eventually follows a sigmoid curve (where all hockey sticks must go) as lower hanging fruit exposes itself (eg, higher parrallelism and GPU computing).
Kurzweil (say what you will) has written a lot on the topic.
The experience and returns to scale are still plodding along but Dennard Scaling is gone and with it, its magic, vanished by physics.
With four or six cards, as many systems built around this component are likely to be spec'd, you'd be able to build something that could fall in the top 5 in the last ten years.
Calling this great news is exaggerating greatly.
Give it up.
This is important for image stuff, true, but it becomes really important for things like Memory Networks for NLP tasks and Neural Turing Machine-like architectures. Bigger networks mean they can "remember" further "backwards" in their dependencies, and so do things they can't do before.
That's GREAT news.
Also, the speed up is pretty significant. People train for weeks at the moment - a 10% speed up means often means savings days. This might not seem much, but it lets you iterate much, much quicker.
If people really cared about 20% speedups that much, they could train for 20% fewer iterations at a slightly higher learning rate. Or they could use very slightly less deep networks. Or slightly less wide netwoks. Or slightly less connected networks. Or. most realistically, some combination of many very small changes.
Of course it's nice to have faster machines, but this is hardly a big enough change to make a very noticeable difference, let alone a dramatic difference.
There's nothing wrong with iterative improvement, but let's not pretend it's a revolution.
He is probably one of the guys who decides where the direction of company research is heading and the one who supervises those projects.
Throwing NVidia cards at it. Like everyone else.
https://github.com/baidu-research/persistent-rnn is a pretty interesting way to use all those NVidia cards though.
From Anandtech's GTX1080 review page2: As a result while GP100 has some notable feature/design elements for HPC – things such faster FP64 & FP16 performance, ECC, and significantly greater amounts of shared memory and register file capacity per CUDA core – these elements aren’t present in GP104 (and presumably, future Pascal consumer-focused GPUs).
This requires confirmation though, it depends on whether it uses the consumer chip or the HPC chip.
Edit: AT has an article up http://www.anandtech.com/show/10510/nvidia-announces-nvidia-...
They think it's likely a consumer card, so lower FP16 and FP64 perf. Should be a gaming monster though
We know almost certainly it has lower FP64 perf than GP100 because it has 22% fewer transistors and the FP64 units take a lot of transistors count. However it is less clear about what the FP16 performance is. Nvidia could have decided to match GP100 on that regard.
When they quote "CUDA cores", they've been counting float32 fma functional units; e.g., Tesla K40 has 192 float32 fma units per SM x 15 SMs => 2880 "CUDA cores".
http://docs.nvidia.com/cuda/cuda-c-programming-guide/index.h...
fp16 and fp64 are likely different functional units with different issue rates, as is the case with old hardware; unless they've managed to share the same hardware (since for P100 the quoted fp64 rate is exactly half the fp32 rate, and the fp16 is exactly double the fp32 rate).
Very curious about this as well. I got two GeForce 1080s as soon as they came out, and was very sad to discover that the stated advantage of Pascal architecture (speedup on FP16) is completely lacking on those.
http://www.anandtech.com/show/10325/the-nvidia-geforce-gtx-1...
http://www.anandtech.com/show/9059/the-nvidia-geforce-gtx-ti...
http://www.geforce.co.uk/hardware/10series/titan-x/
The Tesla cards are the ones with no display outputs.
I'm specifically wondering about FP16 handling though. Single precision FLOPS are never mentioned in the blog, nor on the NVidia page. It would be a shame if the FP16 units on this card were gimped in the same way as the GTX 1080...
The announcement skipped energy usage as well.
From Anandtech's 1080 review (page 2): "As a result while GP100 has some notable feature/design elements for HPC – things such faster FP64 & FP16 performance, ECC, and significantly greater amounts of shared memory and register file capacity per CUDA core – these elements aren’t present in GP104 (and presumably, future Pascal consumer-focused GPUs)."
HPC obviously has different requirement but Nvidia can work with interrogators and customers with less of a backlash when fixing issues in this segment.
NVidia actually care about research, researchers and the scientific computing market.
Next time someone complains about the lack of OpenCL support, again, in another framework remember how much work NVidia puts into supporting people who use their cards for scientific computing, and how they listen to them.
While ATI was always about gaming, gaming, gaming, NVidia always worried about the Pro market (Quadro, Linux support, even with proprietary modules, etc) and now with Deep Learning
And now they can sell their deep learning processor for embedded applications (equals $$$)
There are some numbers for 2015 at [1]. In the Enterprise, HPC and Auto markets they had a bit more than $1B in revenue. I believe the Auto market includes some Tegra numbers, but that is "only" $180M.
Gaming is a little more than $2B, and "OEM & IP" is $1B.
[1] http://www.nextplatform.com/2015/05/08/tesla-gpu-accelerator...
Something like OpenCL does not face the same conflict nVidia would if porting core APIs across a wide set of competing technologies. With CUDA, nVidia prioritizes themselves above AMD, intel, FPGAs and whatever parallel compute technology the future holds.
But the truth is that without the hardware vendors putting significant resources into OpenCL it just isn't competitive and won't be until that happens.
The truth is that most of the work in Deep Learning is developing new NN architectures and other algorithmic optimisations. If you are working in the field there is no reason to put up with second class support from non NVidia vendors - just build in TensorFlow, Torch or a couple of other frameworks and wait for the day (one day, we are promised!) when OpenCL is competitive. Then the framework backends get ported, your code keeps running the same, and it can run on all those other architectures.
Everyone has been waiting for that day since One Weird Trick[1]. There isn't really anything to indicate it is getting closer, and AMD's dismissal of the NVidia "doing something in the car industry"[2] doesn't give me a lot of confidence.
Anyway, I hope I'm wrong. Maybe Intel will step-up.
[1] https://arxiv.org/abs/1404.5997
[2] http://arstechnica.co.uk/gadgets/2016/04/amd-focusing-on-vr-...
They seem to be trying hard to create a new market, as Intel's integrated graphics are now good enough for most of the laptop and desktop market.
Also, it needs a Xeon processor. You can plug the nVidia card wherever you want
What? The first KNL parts will be host CPUs rather than PCI add-on cards.
To really make use of this kind of hardware, the compiler is just the start. You need debuggers, profilers, MPI integration, and, last but not least, a sensible programming model in the first place. Nvidia has been pushing that stack for close to 10 years now.
Even in Python. Imagine firing up a 288-hardware-core Multiprocessing pool using a AVX-512 enabled Numpy...without changing your code. Personally that's numerical porn for me and I want Intel to shup up and take my money, now.
Talking about Numpy, isn't that what Numba Pro is already doing with GPU?
Some background: I occasionally find myself doing FPGA design for my doctoral work and am realizing that the job market for when I get done may be better for me if I was fluent in GPGPU programming as it is easier to build, manage, and deploy a cluster of such machines than the same for FPGAs.
My current problem has huge numbers of XOR operations on large vectors and if OpenCL or CUDA could be learned and spun up quickly (I have a CS background) I may be inclined to jump aboard this train vs buying another FPGA for my problem.
Throughput of integer operations ranges between 25% and 100% of floating point FMA performance. 32-bit bitwise AND, OR, XOR throughput is equal to 32-bit FMA throughput.
NVidia has really invested into their developer resources. Of course, if your time to write code and debug driver issues isn't that important, then an AMD card using OpenCL might be the right choice.
(I'll try to be honest about my bias against NVidia, so you can more accurately interpret my suggestions. I think along the lines of Linus Torvalds with regard to NVidia... http://www.wired.com/2012/06/nvidia-linus-torvald/ )
A problem just calculating, say, hamming distance or 1-2 integer bit ops per integer word loaded will probably be memory bandwidth bound rather than integer op throughput limited. More complicated operations (e.g., cryptographic hashing) that have a higher iop / byte loaded will be limited by the reduced throughput of the integer op functional units rather than memory bandwidth.
For "deep learning", convolution is one of the few operations that tends to be compute rather than memory b/w bound. It's my understanding that Sgemm (float32 matrix multiplication) has been memory b/w limited for a while on Nvidia GPUs. Though, if you muck around with the architecture (as with Pascal), the ratio of compute to memory b/w to compute resources (smem, register file memory) may change the ratios up.
http://www.psc.edu/publicinfo/news/2001/terascale-10-01-01.h...
It cost $45M, and peaked at 6 Teraflops. (I think that was on 32 bit floats, but I can't find the specs. It might have been 64 bit floats)
"Total TCS floor space is roughly that of a basketball court. It uses 14 miles of high-bandwidth interconnect cable to maintain communication among its 3,000 processors. Another seven miles of serial, copper cable and a mile of fiber-optic cable provide for data handling.
The TCS requires 664 kilowatts of power, enough to power 500 homes. It produces heat equivalent to burning 169 pounds of coal an hour, much of which is used in heating the Westinghouse Energy Center. To cool the computer room, more than 600 feet of eight-inch cooling pipe, weighing 12 tons, circulate up to 900 gallons of water per minute, and twelve 30-ton air-handling units provide cooling capacity equivalent to 375 room air conditioners."
NVidia is still taking the one-size-fits-all approach to AI and graphics, maybe it is time to develop AI-specific hardware.
I'd guess NVidia sells far more cards than Google will produce TPUs.
I think it's a related to storing float arrays as arrays of 8-bit integers in memory and converting them into floats just before using. It's 2x more space efficient than fp16
http://www.kanadas.com/program-e/2015/11/method_for_packing_...
Does anyone knows the performance of the half-precision units (16-bit floating point)? It's probably 1/64 the FP32 rate, but Nvidia may have been generous and uncapped it at 2× FP32 like GP100, which would be a big difference (128× factor!)
X2 would have been easier all round.
Looks like its still crap :(
Long live the titan black!
Well I'm glad I got my 1080 now.
780ti > 980 = ~15% increase 780 > 780ti = ~12% increase 680 (potentially should've used the 690 as a reference even tho it was a dGPU card but it was the) > 780 = ~27 increase.
For the most part in generation that there wasn't a near ~2 year gap, and in generations where there wasn't a huge change in GPU memory type or a major architecture change (like dropping the hardware scheduler in Fermi for Kepler in order to save silicon space for things that actually matter(ed) for game at the time) there isn't a 50% increase gen to gen. 15-20% increase gen to gen with comparable price point cards (nvidia has been making it harder now by charging 150-200$ more per card and effectively bumping a price point lately) is what you should be expecting.
The inline results of this generation do diverge from the over the top marketing from Nvidia in the past year about the 10x increase in performance that would be delivered, largely due to the upgrade to HBM. It is certainly probably that had they delivered on HBM we would be seeing a rare jump in performance above the trend, but clearly this has not happened. It seems unlike to arrive on the 1080-ti after not making an appearance on the Titan and so we will need to wait another generation to see what difference that actually makes.
We can't clock these things faster like we used to in the days of 6Mhz CPUs. We can't shrink feature size since 10nm is proving to be a difficult node to crack. We can't jam more features onto the die since we're already producing some of the largest possible dies.
The easy gains are gone. Now we're stuck dealing with the hard stuff. Gains will be slower.
These are not the same thing. As I said: I am well aware of the engineering challenges at this scale. You are trying to argue that it is hard: so what? Should Nvidia get a badge for "trying hard". It still doesn't change the fact that what they have achieved in this generation is the same (relative) improvement as previous generations. This is not what they were selling in the run-up to the 1080 launch when they were still claiming 10x improvements due to the new memory subsystem.
They did not do what they said that they would. This is quite a simple fact - why do you try to argue that what they have done is hard. The difficulty is irrelevant. Do you think that an evaluation of the performance that they achieved should include how hard they worked on it?
If they were falling behind in transistor counts you might have a case, but they're not.
It doesn't really matter if it's above or below the trend line. Maybe the trend line is bullshit now because all the things driving it were the easy gains we've exhausted.