Intel Prepares to Graft Google’s Bfloat16 onto Processors
nextplatform.com
nextplatform.com
Numerical details: https://software.intel.com/sites/default/files/managed/40/8b...
Support for bfloat16 is already present in MKL-DNN (https://github.com/intel/mkl-dnn)
Disclaimer: I work for Intel
Dropping denormals is a huge mistake. This is easy to see if you draw out a number line for a very tiny floating-point format, for example with a 2-bit exponent and a 2-bit fraction. (do this on a sheet of graph paper) Without denormals, there is a huge gap surrounding zero.
Strangely, the infinities were kept. Treating these as NaN is far less harmful than dropping denormals. Treating -0.0 as 0.0 and never producing -0.0 would be less harmful. (the PDF didn't say what happens) Even treating NaN values as normal numbers is probably less harmful than screwing up the denormals.
IEEE floating point has lots of crazy stuff to annoy hardware vendors. Most of it isn't all that important, but denormals matter.
The Bfloat16 format does allow for denormals. Intel's implementation mangles them, changing them to 0.0 on both input and output.
Nervana is discontinued, isn't it? Compatibility doesn't matter. It's pretty compatible anyway, as long as you aren't demanding bit-identical output.
You mean the CPUID bit? That's free. Toggling denormals isn't.
> Nervana is discontinued, isn't it? Compatibility doesn't matter. It's pretty compatible anyway, as long as you aren't demanding bit-identical output.
Nervana isn't discontinued according to their website[1], and bitwise compatibility does matter, certainly more than denormals do.
Denormals are far more important than bitwise compatibility. To be clear, you would still be able to load a Nervana-produced number into a processor that supports denormals, and the other way would work too. You'd just avoid mangling numbers that are near zero.
If you still think denormals don't matter, seriously do what I suggested: draw it out on graph paper. They matter.
bfloat16 doesn't handle denormals, why is there a bit to toggle it? What's it called (so I can CTRL-f for it)?
> If you still think denormals don't matter, seriously do what I suggested: draw it out on graph paper. They matter.
No, I get how denormals work, I know what you're pointing at. But ML genuinely doesn't care, neural nets don't give a damn about mathematical purity[1]. In contrast compatibility matters because ML doesn't give you any guarantee that it's not depending on the behaviour at small values, and minor differences in rounding does cause issues. For example, Leela Chess Zero had difficulties with reproducibility because different GPUs round floats differently.
[1] Fun but relevant aside: https://openai.com/blog/nonlinear-computation-in-linear-netw...
Bfloat16 obviously can handle denormals. The encoding is possible. There would be no need to handle the issue if the encoding did not exist.
As hex, these would be denormal: 0x0001 to 0x0080, and 0x8001 to 0x8080. It's the same as plain old 32-bit IEEE, with half the bits lopped off.
ML is all about difficulties with reproducibility. I don't see a reason to get upset about denormals when a 3D-printed turtle can be confused with a rifle.
You mean the standard ones for normal floats? You certainly wouldn't want to reuse that for bfloat16s.
> ML is all about difficulties with reproducibility. I don't see a reason to get upset about denormals when a 3D-printed turtle can be confused with a rifle.
These are different things, despite the similarity in terminology.
Having Nervana and friends on a Xeon chip could be a huge positive change for software. Not only could we toss out the issue of GPU memory transfer, but Nvidia GPUs aren’t so great with concurrency, and here with the linux kernel we might have a chance to beat Nvidia. Naveen sure would like that... Nervana once had a Maxwell compiler that was better than Nvidia’s.
Nowadays it's really a moot point with Nvidia's Cutlass being open source.
One other thought about the Nervana-Xeon convergence is that the support for more memory (thru DDR, Optane, or even just mmap'ed NVME) will be a big win for modeling and large-minibatch SGD. For example, the minibatch fetching could be pushed to the hardware / OS instead of Tensorflow (or the crazy guy behind Tensorpack) using a threadpool. A lot of training is still I/O bound at some level, and processors only support so many PCI-e lanes...
bfloat helps with this because the data is half as large. But of course if you're doing Nvidia you're probably already doing (IEEE) fp16.
https://en.wikichip.org/wiki/intel/microarchitectures/cooper...
says "Higher bandwidth (174.84 GiB/s, up from 119.209 GiB/s)"
I don't know if memory bandwidth matters for this type of job, though.
I've found I get about 75-80% of the advertised bandwidth both from my real app (TLS crypto) and a toy memory copy benchmark using AVX256 instructions. The toy memory copy benchmark is how I realized that my bottleneck was actually memory bandwidth and not CPU horsepower on Broadwell based servers.
To make a stab, I suppose it might depend on whether all requests are coming from a single memory bank or spread evenly across all memory banks, assuming fully populated (again from the link "Octa-channel (up from hexa-channel)")
When GPUs kickstarted deep learning research in 2012, people had already studied shallow models on mapreduce for a decade or so. Once NVME / Optane and modern CPUs get 1-10TB of useful “memory” in the hands of grad students, there should be another wave of new research. To date, my experience has been that 1TB of “memory” is only commonly available in industry.
Why implement bfloat if you get just slightly less performance emulating it with AVX512, which already exists? Maybe it’s an “us too” claim?
AVX512 is expensive. I believe if you have an AVX512-heavy workload it can cause the processor to throttle.
Running programs doing the same thing (eg, Hamiltonian Monte Carlo where the likelihood function has or has not been vectorized), the avx512 version is far faster than scalar, and routinely 50%+ faster than avx2.
The avx512 instruction set itself also provides conveniences that make it easier to explicitly vectorize, even if most compilers don't take advantage of them on their own. Masking load and store operations in particular (they're better about masking to handle branches).
On why avx512 vs a graphics card: I need double precision, and my code routinely has maximum widths smaller than the 32 or 64 a graphics card would want to computer in parallel.
So is 105C just a very conservative number, compensating for the probe location, or are there process specific things which brings it down to 105C?
When running Stan (NUTS/HMC) on Xeon Phi, telling Eigen to use avx512 provided a noticeable speed up but I didn't look at the assembly to be sure.
In the example I give there, the logdensity and gradient evaluation was about 25x faster than Stan, and sampling was about 20x faster. A simulation fitting many data sets for my dissertation took about 9 hours. 20x is the difference between running overnight, and taking a week.
If I understand correctly, one problem Stan has is that it uses a var datatype for its arrays, which interleaves the values (Scalar) with pointers (vi_). https://github.com/stan-dev/math/blob/master/stan/math/rev/c...
This interleaving is going to cause problems to an autovectorizer. To get a SIMD vector of the scalars, you'd probably have to load two vectors, and then blend them.
Even with arrays of doubles, I found Eigen's fixed size arrays to get about 3-8x worse performance than my Julia library (3-8x worse than my Julia library for Mx32 * 32xN, for combinations of M and N = (3,...,32) ): https://bayeswatch.org/2019/06/06/small-matrix-multiplicatio...
I compiled the Eigen benchmarks with: g++ -O3 -fno-signed-zeros -fno-trapping-math -fassociative-math -march=native -mprefer-vector-width=512 -shared -fPIC -I/usr/include/eigen3 eigen_mul.cpp -o libeigenmul.so
How did you tell Eigen to use avx512? At the time, I was getting errors when specifying -DEIGEN_ENABLE_AVX512. http://eigen.tuxfamily.org/bz/show_bug.cgi?id=1705
that looks pretty cool, though I don't yet know enough Julia to understand all of it. The speedups make sense given that Stan's compiler/math lib doesn't do much in the way of smart data layout. I would still keep in mind that the metric worth using for benchmarking is the number of effective samples per second, and this also depends on the HMC variant you use.
> Eigen's fixed size arrays to get about 3-8x worse performance than my Julia library
seems unsurprising that Julia can specialize a lot better than verbose C++ templating, no? (still, good job, very worth checking out)
> I was getting errors when specifying -DEIGEN_ENABLE_AVX512
I used this flag with Eigen 3.3.1, I think, on GCC 6 or 7. This was for Xeon Phi, so I tried to use icc but despite supporting C++11 it doesn't handle Stan or Eigen's template metaprogramming.
This is all the more reason to use Julia, but my graduate student days are long past..
I was getting similar effective sample sizes/sample size in both after switching to a diagonal mass matrix, like Stan uses, from the dense mass matrix DynamicHMC.jl uses by default (the HMC backed library I'm using).
Given how common it is for folks to run Stan over night or for a week to study prior sensitivity, internal coverage, type I and II errors, etc, via Monte Carlo, I think a focus on speed is worth while.
> seems unsurprising that Julia can specialize a lot better than verbose C++ templating, no? (still, good job, very worth checking out)
The C++ library Blaze did a lot better than Eigen, but still not as well. But yes, Julia'a meta-programming is much easier to work with. Julia expressions are Julia objects that you can manipulate like anything else, so I can write all the functions I want describing how to generate matmul kernels as a function of matrix size and CPU Info, and how to loop over them.
That approach feels much more straightforward. I haven't looked at the code bases of Eigen or Blaze, nor am I that familiar with template meta-programming. But I'd guess they define matmul recursively for arbitrary fixed sizes, and then have some templates defined for specific sizes (the kernels) -- or ideally have some clever way of generating the kernels from there.
Regardless, I agree that this is much easier in Julia. Aggressive specialization is also better aligned with Julia's compilation model in general, because methods get compiled just before they're used. Defining a million possible specializations doesn't have the cost of compiling a million specializations.
> This is all the more reason to use Julia, but my graduate student days are long past..
I'm defending this week, and next Monday will be my first day in an industry job. They expressed openness to Julia, but my biggest fear is that they'll renege so that I'll only be able to work on or use Julia in my spare time at home.
That's an interesting comment. We've always used Stan's default of a diagonal, but I think we'd benefit from mixed metrics, which doesn't seem possible in Stan, but looks somewhat doable in some of the HMC libs in Julia.
> Given how common it is for folks to run Stan over night or for a week to study prior sensitivity
Yes, we changed the walltime on our Slurm cluster to support Stan Jobs running up to a week long, have used multiple million core hours on this. Stan still isn't so shabby but it's a hard problem.
> I'm defending this week, and next Monday will be my first day in an industry job.
Good luck and congrats on the job. You'll probably have to bite your tongue and look for opportunities where Julia's advanced compilation model (as you described well above) is going to more than pay for the cost of deployment/extra language etc.
=> A vector multiply using bfloat16 may not be much faster than one using float32, but it will do more multiplications.
I'm unclear on the advantage you are trying to explain. If both AVX and bfloat are SIMD instructions that cannot be the reason implementing bfloat is better. I'm expecting something like "bfloat16 is more specialised so it can have larger registers" or something?
[Edit]
Sorry re-reading your comment, I think you are trying to say that. The key part being:
> fit more numbers vector registers
(vs AVX i assume), so adding bfloat16 would provide more registers vs AVX with similar gate usage due to greater specialization?
For CPU-bound algorithms, one would expect that bfloat16 in 512 bit vector registers would be about equal in speed to float32 in (hypothetical) 1024 bit vector registers.
Also, for algorithms that are memory-bandwidth bound, halving the size of your numbers will (about) halve memory pressure.
Sorry i'm pretty ignorant of AVX so trying to understand... is this because the smallest word size in AVX is 32bit? compared to bfloat16 is using twice the register space? Or rather with the same register space in bfloat16 you can have twice the numbers? (with no negative effects on convergence in NN due to exponent size.)
Tensorflow OpenCL support bug [1] has been open for FOUR years now (with the discussion devolving into an Intel PlaidML flame war).
AMD OpenCL is now ROCm ?
At the end of the day, I cant run ANY accelerated workloads using Intel graphics or AMD ....because there's simply no software support anywhere.
OTOH, if you have a nVidia stack... boom. you get accelerated python https://developer.nvidia.com/how-to-cuda-python
Are you running containerized workloads on kubernetes ? it has BAKED-IN support for nvidia-docker (https://kubernetes.io/docs/tasks/manage-gpus/scheduling-gpus...)
Is anything going to change ?
You can, there’s an Intel framework: OpenVINO that is targeting Intel hardware, processors and HD video cards, you can convert TF graph to Vino and use their inference server, that mimics TF serving.
There’s TF ROCm port by AMD as well.
As it stands, CUDA is the only API with a decent implementation available on the three major desktop OSes. It's ridiculous.
I don't see how GPU with AVX512 can compete with TPU's or GPU's, BF16 or not.
You don't really anyway. A GPU instance in Azure will beat whatever extension accel intel does on CPUs by a mile either way.
Where there is lock in is one level higher...for ML all the other players are going FPGA which I reckon is a bad move.
FB's approach on the other side requires entirely redesigned and separate execution units. That's harder to justify, that silicon will remain dark for non-DL usage.
I imagine it still keeps most of the performance benefits since it's eliminating around 2/3rds of the longest binary component.
Ditto for them calling it "Brain Floating Point" [1]. I mean it appears to be a typical floating-point numeric data type with some of the precision truncated to reduce its cost.
I guess they might be trying to make it sound super-fancy to preempt people seeing it as cheap or/and low-tech?
[1]: https://en.wikipedia.org/wiki/Bfloat16_floating-point_format
Also, converting BF16 to FP32 and back is just a vector shuffle that sticks/drops 16 extra mantissa bits at the end of each FP16 value, so it's cheaper than other floating point conversions. This means that even if you occasionally have to escape to FP32, the overhead is low and you keep all of the memory bandwidth benefits.
It may well be a stupid q but I really don't know, and always assumed they would be [-1..+1] and that fixed point would suffice. Clearly not.
So not quite as good as a 6-inch slide rule.
Edit: voters in this thread seem pretty uptight.
Speaking of boiling the ocean . . .
It's possible that some other custom format is better in absolute terms for models trained specifically for the custom format. But for the current ecosystem, where models are trained primarily using float32, bfloat16 is a very good choice.
It seems like a huge mistake if they missed key optimizations, but I'm happy to take your word for it.
Are there write-ups that go in detail about these mistakes? It's the kind of somebody-is-wrong-on-the-Internet topic that would result in flaming blog posts. :-)
I mean, I helped design the posit spec and the twos complements treatment is something not even John Gustafson understands... The key insight is that the hidden bit is -2 for negative numbers (instead of 1 as it is for positive numbers). It's kind of nonobvious and I happened upon it by accident one night while fooling around with circuit diagrams. If people really get serious about it I'm sure though that it will get rediscovered by EDA folks smarter than I.
But I'd be rather glad if they implemented unums already.
[1] https://github.com/kstenerud/compact-float/blob/master/compa...
For weights during training, 7 bits of mantissa also seems a bit low - it's common for weights to adjust much less than 1% during a single batch of training, which this couldn't represent.
I think this is more a "we want something which is faster but is compatible with existing code written for fp32's".
You're right that the mantissa is small. The trick is that you always accumulate into fp32 and then truncate down to 16 bits at the end. You'd do this for any 16-bit floating type.
Source: I work on this at Google.
In particular the latter describes a generic framework that can be used to generate a lot of different number systems. Could hardware implement this, allowing us to compose and choose the number system by just setting some simple flags?
Intel (for better or worse) takes a very experiment-results-driven approach to choosing which features to implement in hardware. So this result -- that software emulation of a feature works almost as well as a hardware implementation would -- probably makes Intel less likely to implement the feature in hardware.
ARM, AMD, RISCV etc, will probably come to similar conclusions.
https://www.nextplatform.com/2018/12/16/intel-unfolds-roadma...
Just in case, since I can see someone else being serious about this, I think the gist is that neural-networks tend to be fairly approximate things such that we're not particularly concerned with having a lot of precision in many cases. This use-case wouldn't seem to demand a 32-bit variant too often.
But.. if you want it anyway...
Higher-bit extensions would seem to be floating-point values that favor range-over-precision more than typical floating-point numerics with the same bit-count.
If we take that to an extreme, we can talk about ranges over infinities and infinitesimals -- this is, much like the hyperreal-number system [1].
And ya know what's funny?
Some guy's been pushing for such a primitive numeric data type [2] since the early-2000's [3]!
[1]: https://en.wikipedia.org/wiki/Hyperreal_number
Graft, as understood in American English, is a form of political corruption, being the unscrupulous use of a politician's authority for personal gain.
Edit: by the way, I really couldn't fit the term with the article. And realized I was probably looking at the wrong definition. This one might be much more apt:
a shoot or twig inserted into a slit on the trunk or stem of a living plant, from which it receives sap.
This is the definition they most likely meant.
Does that really deserve downvoting? If so, explain it to me please, because I just can't see why.
On the other hand, many of the readers of this blog are native English speakers, and to most of us the metaphorical meaning of the title was clear. A very useful function of voting is to re-order the comments on the page. This is a useful comment, but only for a small subset of readers. As such, it's perfectly reasonable that it would appear after the other more technical comments.
Which is to say that while you don't deserve to be downvoted, the comment arguably does. I personally upvoted it, because of my personal experience with intercontinental miscommunication involving the word "graft", but I can see why others would want to prioritize other comments. Not because it's a bad comment --- I'm sure a few people found it really helpful --- but because it's a meta-comment on a technical site that helps only a small portion of the audience.
So you should proudly keep making helpful comments like this, and not take it at all personally when they are downvoted to the bottom of the page. This is a case where you should feel confident that you did the right thing, despite the apparent feedback.
In this situation I would personally not downvote, but rather upvote the other more technical comments, which provides a similar result, without penalizing the commenter.
Your interpretation might be right, but a downvote will often be interpreted as "I didn't contribute to the conversation". The fact that "graft" is confusing, and that I am trying to shed light on it, is a way for me to try contributing, and therefore, in my view, shouldn't be penalized.