Caffe2 adds 16 bit floating point training support on the NVIDIA Volta platform
caffe2.ai
caffe2.ai
I'm interested more in understanding why Caffe2 would outperform Theano, Tensorflow, MXNet, etc. once Volta chipsets are generally available, beyond early pre-release optimization, particularly when most of the front-runners are already leveraging / taking into account NCCL, CuDNN, NVLink, etc. When the burden of adding support for new NVIDIA primitives is so low, what gives Caffe2 an advantage beyond an ephemeral "we were partners with NVIDIA first" one-up that would last for a couple of months at most?
(Apologies in advance if this post sounds overly negative, but I am constantly evaluating the current crop of frameworks for the trade-offs they enforce on the problem space, and a definitive answer would be very helpful.)
(In fact there are even some experiments with a very low number of bits - can't find the link now though)
There is also some work on pure binary networks, xnornet and binarynet. Low precision doesn't seem to effect these networks as long as they are trained with binary weights in the first place.
A vendor-made benchmark where the vendor is outperforming other vendors.
Let's not jump to conclusion on what's really fastest.
https://dennisforbes.ca/index.php/2017/04/11/floating-point-...
Most desktop chips offer little advantage for FP16 (at best that the source memory footprint is smaller, maybe better cache hit rates), but newer GPUs can actually blaze. If it has the magnitude/precision needs for a task, it's a big win.
The reason you don't typically have FP16 on CPUs is that in the CPU world, cache dominates die area and power usage. The floating point ALU component of a CPU is relatively tiny. Still, even in the CPU world, Intel recently added instructions to convert to/from FP16 before operating on it, because they know there is a memory bandwidth advantage in using a smaller floating-point format.
It's interesting the speed up isn't more pronounced between Volta and Pascal considering the Tensor cores on paper give you about 6x the MFlops. The price differential looks large.
From AnandTech: "By the numbers, Tesla V100 is slated to provide 15 TFLOPS of FP32 performance, 30 TFLOPS FP16, 7.5 TFLOPS FP64, and a whopping 120 TFLOPS of dedicated Tensor operations. With a peak clockspeed of 1455MHz, this marks a 42% increase in theoretical FLOPS for the CUDA cores at all size. Whereas coming from Pascal, for Tensor operations the gains will be closer to 6-12x, depending on the operation precision."
EDIT: Nvidia's advertised TFLOPS are:
FP16 FP32 FP64
V100 30 15 8.5
P100 21.2 10.6 5.3
K40 4.29 4.29 1.43That's a core that does 4x4 FP16 matrix multiplication + 4x4 FP32 accumulation in one go.
That's where V100 gets its boost, up to 120 TFLOPS.
I think it's great that what started as a specialist gaming device is now being used in industry for Big Things. The development cost that Nvidia et al. have invested in new designs has undoubtedly been financed in (large) part by the gaming community. Now income and advancements for both sectors feed into the other and gamers like me are reaping the benefits with reduced price:performance across the range.
Still waiting for them to produce anything worthwhile buying instead of AMD and NVidia GPUs.
There are some very interesting talks by Intels experts on how to use these guys with AVX and AVX512 (or however its called now) and the multiplicative improvement both multicore and AVX give when used together (the sales line goes along like this "10x for multicore, and on top of that 4x for AVX, voilá 40x speedup!").
I don't work in an HPC environment, so I don't really know if they stand up in reality to Intels claims.
While the "tensor" units may land in consumer stuff (though I don't think it'll happen this year) I expect there to be little "trickle-down" planned or wanted at this stage.
The new tesla volta super flip flop at 1.21 gigawatts blah blah blah.
Just saying.
Also Kudos to Nvidia for the buzzword/made up word creation for their products.
How are things in the red camp ? There was some HIP thing where Fiji was as good as Pascal.
Wait for ryzen.
Wait for vega.
Wait for ryzen again.
Vega is coming out this year. Volta isn't consumer based. Go check the enterprise prices.
Consumer Volta is mid 2018, so this, from a consumer perspective, is a paper launch.
Also, recent Linux patches hint that there will be dual-chip water-cooled Vega boards too which is likely not even 1080 Ti regime: https://www.techpowerup.com/233208/linux-drivers-point-to-up...
As for scientific computing I suppose there are workloads where having your GPU and CPU on one die might give you latency savings that are worth it but mostly NVidia is going to keep owning that market for the next few years.
The mid-range has a lot of potential as it's based on Fiji which NVIDIA can't match in terms of memory BW with any of their existing ML products (nor near future ones). This is important IMO because ML inference is memory-bound and GDDR5 on NVIDIA's cards is no match for even 1st gen HBM.
The big cards is announced to be 12.5/25 Tflops SP/HP, so as long as it's not "tensoring" it should be in between P100 and V100. If they get the price right (and given that even P100 is currently still $5.5-6 w/o tax), it will hopefully gain traction.
A lot hinges on their fully OSS (!) software stack, but they are getting things done, just had a new release: https://rocm.github.io
You don't call your PC a specific purpose computer just because all you do is use google chrome.
Potential customer: "What's an MPU?"
N: "Oh, it's the same thing we used to call a GPU"
P: "Why did you change the name?"
N: "Because it's a little bit weird to keep using the word 'graphics' for something that is being made more and more specifically for other things"
P (puzzled, and marginally less likely to make a purchase): "Oh right"