Someone's been messing with Python's floating point subnormals
moyix.blogspot.com
moyix.blogspot.com
2004 - GCC 3.4: - Refinements and Extensions: Additional optimizations added, including better handling of subnormal numbers and more aggressive operation reordering.
2007 - GCC 4.2: - Introduction of Related Flags: `-funsafe-math-optimizations` flag introduced, offering more granular control over specific optimizations.
2010 - GCC 4.5: - Improvements in Vectorization: Enhanced vectorization capabilities, particularly for SIMD hardware using SSE and AVX instruction sets.
2013 - GCC 4.8: - More Granular Control: Introduction of flags like `-fno-math-errno`, improving efficiency by assuming mathematical functions do not set `errno`.
2017 - GCC 7.0: - Enhanced Complex Number Optimizations: Improved performance for complex number arithmetic, benefiting scientific and engineering applications.
2021 - GCC 11.0: - Better Support for Modern Hardware: Optimizations leveraging modern CPU architectures and instruction sets like AVX-512.
2024 - GCC 13.0 (Experimental): - Experimental Features: Additional optimizations focused on new CPU features and better handling of edge cases.
Sources: - GCC documentation archives - Release notes from various GCC versions - [GCC Wiki](https://gcc.gnu.org/wiki/) - [Krister Walfridsson's blog](https://kristerw.github.io)
2012: I file the obvious bug:
https://gcc.gnu.org/bugzilla/show_bug.cgi?id=55522
Early 2023: bug is fixed.
Feb 2024: clang follows suit:
https://github.com/llvm/llvm-project/pull/80475
And maybe later this year this problem will finally be gone in common Linux distros.
Like it's a quirk that floating point math can be non-deterministic but how bad is it, and is it actually a bug?
It's not too hard to imagine how an algorithm might depend upon that — there could be a branch for the case where `x == y` and then a branch that relies upon dividing by `(x-y)` and assumes that it's not a division by zero.
Look, -ffast-math isn't always fine, but specifically I'm looking at DAZ/FTZ enabled, possibly non-deterministically process-wide (which is bad, don't get me wrong!). But part of why this doesn't phase people is because in practice, programs don't care and the observable effects don't lead to bugs.
Did you intend to say "without FTZ/DAZ enabled"? If so, that's completely and provenly false.
https://en.wikipedia.org/wiki/Sterbenz_lemma
Perhaps a better example than `z/(x-y)` is `log(x-y)`. Unlike division by 0.0, `log(0.0)` often throws an immediate error whereas `log(5e-324)` is a finite — and meaningful! — result.
It would also make it impossible to have a decent non-zero check for
if ((x-y)!=0) progress((x-t)/(x-y));
Having an option that says "unsafe" creates the expectation that unsafe options are labeled like that, and then enabling said unsafe option from another option that doesn't have "unsafe" in the name is a bit surprising.
I remember this one because of this part:
> Unbeknownst to me, even with --dry-run pip will execute arbitrary code found in the package's setup.py. In fact, merely asking pip to download a package can execute arbitrary code (see pip issues 7325 and 1884 for more details)! So when I tried to dry-run install almost 400K Python packages, hilarity ensued. I spent a long time cleaning up the mess, and discovered some pretty poor setup.py practices along the way. But hey, at least I got two free pictures of anime catgirls, deposited directly into my home directory. Convenient!
Like dry-run not even working as you'd expect? I wonder if that got fixed after this article?
It’s because academia is a sort of anchor that preserves poor practices.
It largely isn't, but python is a big ship and it takes time to turn. There's been a lot of movement in python packaging semi-recently and as far as I can tell using setup.py has been considered legacy for a while now.
Biggest problem is that the new way to do packaging is not documented very well. It's split over three or four different projects. It mostly amounts to "you need to make a pyproject.toml file" but it's somewhat tricky to find what such a file is even supposed to contain.
It has been done with distutils: it has been around for more than a decade. The result is the ecosystem that you are complaining about.
tar xzf our_custom_package.tar.gz -C "$(python3 -c 'import sysconfig; print(sysconfig.get_path("platlib"))')"
I'm only half-joking. # Initially copied from
# https://github.com/actions/starter-workflows/blob/main/ci/python-package.yml
# And later based on the version jamadden updated at
# gevent/gevent, and then at zodb/relstorage and zodb/perfmetrics
Although the linked config was modified 2 years ago to removed the unsafe math options[1],
the copy-paste propagated before then.Naturally anyone asking CoPilot, ChatGPT, or other modern LLM-based interface for config code will likely get something based on this, with the problematic -Ofast option included.
[1] https://github.com/zopefoundation/meta/commit/9c07520e90d9a1...
It's less common for other languages because they just don't do it as much.
That will cause some installations to fail if wheels are not available. However, given that wheels are increasingly common (even for pure-python "source" packages), this can be used as a sort of bisect to enumerate and isolate/audit/file issues on/sandbox/etc. the remaining setup.py-based distributions.
os.system('cp -r ./pics/ ~/.local/share/koneko/') pip install nukemydrive #no, this package doesn't exist (yet)
rather than than the harder to remember rm -rf --no-preserve-root /Even when the errors happen to be discovered soon enough, people will be puzzled about their origin and they may waste a lot of time analyzing their algorithms, because there are a lot of things that can cause numeric errors.
Because at their origin the underflows will neither signal any exception nor generate any unusual value, the errors will be normally caught much later, perhaps after thousands of other operations, when their cause will not be obvious.
This fight between the people who want to get only correct results from computers and the people who do not care whether the results are erroneous as long as the results are obtained after a delay shorter by a few percents than for obtaining correct results has continued for decades, almost since electronic computers have been invented.
Errors are acceptable in games or for some other graphics or audio applications that generate ephemeral images or sounds, but they are not acceptable for any engineering purposes.
The option "-Ofast" is like smoking. There is no doubt that it is a bad habit, but whoever wants to smoke should be free to do it. On the other hand, exactly like a smoker should not be permitted to smoke in a closed room with non-smokers, the option "-Ofast" and the related options should not be permitted to alter the behavior of any other programs that are linked with an object file compiled with it.
The behavior of "-Ofast" where it affects globally the content of MXCSR is unacceptable.
The right behavior would have been for any function compiled with "-Ofast" to save MXCSR, put in it any desired value for the duration of the function execution and restore it at function exit. Moreover, when invoking any other function it should restore the original MXCSR before the call and put again in it the desired local value after the function returns back.
It is much more frequent for -ffast-math to generate bigger errors than to benefit from an accidental cancellation of the errors.
The only case when the result of underflows is predictable is when the values whose computation may result in an underflow are immediately added to other values that are known to be big enough, and they are not used for anything else.
It is not frequent to have so much information about which expressions will underflow and which will not, and about the range of possible values of the values that would be added. When there really exists so much information about what must be computed, then there are chances that the computation can be rearranged in a way that would avoid the underflows, in which case there would be no need to change MXCSR to ignore the underflows, because they would not happen.
In my experience underflows are very uncommon and indicate that you're doing something wrong, whereas fp contractions are extremely common. So I disagree with you.
a*b - a*bNevertheless, I consider that a compiler option that gives you good FMA, but only together with bad behavior at underflows, is something exceedingly stupid.
FMA is a standard operation, while ignoring underflows is rightly prohibited by the floating-point standard. There should never have been a compiler option that mixes standard operations with non-standard operations.
In C or C++, you can use the "fma" function to get FMA where desired.
This is more awkward, but it is safe.
Any better solution requires a change in the programming language. For instance, it could be required that adjacent * and + must always be contracted into an FMA, unless separated by parentheses. This would allow the programmer to select between FMA and FMUL+FADD, as desired.
FMA has appeared in 1990 (in IBM POWER) and it has been the greatest advance in floating-point computation after the introduction of Intel 8087 in 1980 (which has resulted in the IEEE standard a few years later). It is weird that 34 years later, most, if not all, programming languages do not handle well the generation of FMA operations.
> The right behavior would have been for any function compiled with "-Ofast" to save MXCSR, put in it any desired value for the duration of the function execution and restore it at function exit. Moreover, when invoking any other function it should restore the original MXCSR before the call and put again in it the desired local value after the function returns back.
I completely agree here. The problem is that introducing floating-point environment into the programming model requires disabling some of the transformations you want to do with fast-math in the first place, so doing it this way kind of doesn't work to enable what you wanted to do. Furthermore, one of the SPEC benchmarks gets like a 30% speedup if you turn on denormal flushing, and that's the kind of change that gives compiler engineer management heartburn if you want to tell them to forgo that speedup because the optimization is dumb.
I am grateful for blog posts like this one because this does help build a case for getting the compiler to stop doing stupid stuff like globally enabling DAZ/FTZ bits.
On some CPUs the penalties for exceptional cases are negligible (e.g. on AMD Zen), so it is impossible for any SPEC benchmark to get a 30% speedup when ignoring underflows.
There are also bad CPU models, which have a microprogrammed handling of the exceptional cases, which can add e.g. a penalty of somewhat more than one hundred CPU cycles for each underflow, which is very similar to the penalty for a load from the main memory (which misses the caches).
If one of the SPEC benchmarks has so many underflows that on a bad CPU model it can be sped up by 30% when ignoring the underflows, then it is pretty much guaranteed that when the underflows are ignored the results computed by that benchmark are erroneous.
Cheating at the SPEC benchmark by ignoring the underflows could be easily avoided by changing the benchmarking rules, by providing a file with the valid results and by not accepting any benchmark claims where the results do not match those known to be good.
You can cheat at any SPEC benchmark with bogus compiler options, if there is no constraint on the compiled program to be correct. For instance you could make a compilation option that deletes 90% of the generated machine code, resulting in a 10 times faster benchmark program.
For some reason, the compiler writers that handle the generation of code for floating-point operations feel that they are licensed to generate incorrect code, even when similarly incorrect code would be rejected in other contexts. The fact that the floating-point operations are inexact, so they are accompanied by some inherent errors, does not give the right to a compiler writer or library writer to introduce arbitrary errors in the computations. The reason why a standard exists is precisely for guaranteeing that the errors are bounded and you can predict their effects.
In other cases, `-ffast-math` just introduces arbitrary and strange behaviors. Sometimes you end up with higher precision than you expected. Other times you end up with less. Other times it'll helpfully just re-arrange things such that it's a zero. For example, the classical Kahan summation does the following:
t = sum + y
c = (t - sum) - y
https://en.wikipedia.org/wiki/Kahan_summation_algorithmA -ffast-math compiler will see that — algebraically — you can just substitute `sum + y` into the equation for `c` and get 0. It's `sum + y - sum - y`. And that's true for real maths. But it's not true for floating point numbers.
It explicitly destroys any attempt at _working with_ floating point numbers.
This made me chuckle
Ahh! Yeah, we need better value functions.