Float-Parsing Benchmark: Regular Visual Studio, ClangCL and Linux GCC
lemire.me
lemire.me
To make fast binaries with Visual C++, use Visual Studio to create and configure the project. The most important settings are following:
• C/C++ / Code Generation / Enable Enhanced Instruction Set
• C/C++ / Optimization / Whole Program Optimization
• Linker / Optimization / Link Time Code Generation
While it’s technically possible to set up the correct VC++ options with CMakeLists.txt, it’s relatively hard: need to detect VC++ compiler, and pass these weird command-line switches. Most cross-platform C++ libraries aren’t doing that.
would be interesting to see the benchmark result using cl commands directly.
I told it to generate a VS solution to see what build settings it had applied.. and it is set up very poorly, not surprising performance is ass.
This is release settings:intrinsics = off
SIMD = off
whole program optimizations = off
link time code gen = off
You'd want to compare lto vs lto and not vs not.
This is not vs not. It is a fair comparison in that sense
GCC and clang are both conservative with instruction set withou march option.
I'm often disappointed by MSVC codegen when examining it in godbolt, I'm not surprised by the results here.
Surely some kind of fixed point format would be better for a lot of use cases?
I suppose at this point we're kind of locked in but it seems to me you could have instructions on the processor to do maths where a lower register is below the decimal point so an add would be
ADD {R1 R2 }{R3 R4}
You wouldn't need all new silicon for a FP unit. And these could be used for doing double width math aswell.
"Some DSP architectures offer native support for specific fixed-point formats, for example signed n-bit numbers with n−1 fraction bits (whose values may range between −1 and almost +1). The support may include a multiply instruction that includes renormalization—the scaling conversion of the product from 2n−2 to n−1 fraction bits.[citation needed] If the CPU does not provide that feature, the programmer must save the product in a large enough register or temporary variable, and code the renormalization explicitly."
From: https://en.wikipedia.org/wiki/Fixed-point_arithmetic#Hardwar...
Take the Analog Devices Blackfin range, if you want fixed point, which you're welcome to, they want you to use 1.15, so that's a 16-bit two's complement value between -32768/32768 and +32767/32768. Additions and subtractions are the same as integer maths, we're adding fractions but the adder doesn't care.
For the multiply, we're doing the same operation, 16-bit x 16-bit = 32-bit, but we can shift away the left bit and then snap off the bottom 16-bits to get another 1.15 fixed point value.
If you want this sort of fixed point type, lack of "hardware support" is not what's holding you back, nothing remotely modern would struggle to perform that shift operation on general purpose hardware.
You'd have to do a imul instruction followed by a right shift. On, say, Ice Lake, the imul has a reciprocal throughput of 1 and a latency of 3. The shift has RT of 0.5 and a latency of 1.
So compared to regular integer multiply, a fixed-point implementation on Ice Lake would be 0.67x the speed if you were throughput bound and 0.75x the speed if you were latency bound.
But most (all?) other operations wouldn't be any slower than regular integer operations, so a real program would slow down by less than the figures calculated above.
And thinking back to 8bit 16bit 32 bit processors, it seems to me that hardware support for a fixed point format would gain you more than spending that same silicon supporting floating point. But fixed point never seems to have been pursued.
Because fixed point leads to catastrophic loss of precision way more often.
For example, if you compute the product of two 16:16 fixed point values that both are around 1000, you like want to store the result in a 20:12 or 21:13 fixed point value.
If, on the other hand, both are about 0,001, the result best fits a 12:20 or 13:21 fixed point value.
You find that kind of issue with every operation. Want to sum a hundred numbers? Depending on their values, you may want to have up to 7 or 8 more digits before the decimal point in the result.
Want to compute sin or cos? A 16:16 fixed point results wastes half the storage. You probably want to return a 1:31 fixed point value.
So, what to do if you want to write a general-purpose library? Provide multiple functions for multiplying fixed width numbers?
Returning a pair (int, place of radix point) probably is the best you can do, and that is what IEEE floating point is.
* 9.192631770e9
* 2.99792458e8
* 6.62607015e−34
* 1.602176634e−19
* 1.380649e−23
* 6.02214076e23
* 6.83e2
A floating-point system is capable of representing all of these numbers with the same relative accuracy, independent of the fact that the scale is over 50 decimal (and 150 binary) digits apart from smallest to lowest. A fixed point number cannot do the same; it only works well when the numbers you are operating on are approximately all the same scale.
However fixed point is, as already noted, fundamentally just integer operations so there's nothing holding you back.
Language support can make it somewhat easier though, like templates in C++ can make arbitrary fixed points ergonomic, or in Delphi which has a "native" currency type which is a 64bit integer with a multiplier of 1000 (four decimal digits).
For arbitrary real-number operations that you find in more general physics and statistics, you need something more though as the scaling factor needed can depend on the input values.
The primary thing floats have that fixed point numbers don’t is the same relative precision regardless of magnitude. Combined with the ability to represent an enormous magnitude, and reasonably good precision near 0 and 1, this makes floats easier to use than fixed point numbers. With fixed point, you always have to think about the range and precision required, and with floats you don’t. Add in standards for representing Inf and NaN, which are pretty useful, and it’s easy to see why floats are quite popular.
another question to ask is why did we land on real numbers as the standard?
because they are a good approximation of a lot of physical processes, like sine and cosine which are endemic in graphics, physics, electronics, signal processing, radio, etc. one could even make the argument that video games have driven floating point processors in a lot of the CPU / GPU industry.
but of course they are not the only way to do things. some geometry packages use rational number types, some financial systems use a base-10 Decimal number type, mathematicians studying number theory need the Big Integer type (unlimited digits), but these are small use cases for a small number of people relatively.
so floats are very convenient way for a huge number of people. and the number of people who need to use other number types is rather small which is why they are not implemented in a CPU. which makes them slower , which means they are even less desirable for a lot of applications.
as we see new software, we see new number formats, like for neural networks "float 8" , an 8 bit floating point number, are becoming a thing. which would be incomprehensible why you would want such a thing to most people 20 years ago and yet now it is in mainstream GPU products.
https://stackoverflow.com/questions/38975770/python-numpy-fl...
I'm unclear on what he means by the backend. Is ClangCL essentially the same as using LLVM for Windows with the Microsoft linker, or is there a difference in code generation?
It uses the llvm backend.
Unless Daniel is using something special, he is incorrect. It is not the vs backend.
Microsoft did have a clang frontend + c2 backend project, but I thought they deprecated it.