Re: What is acceptable for -ffast-math? (2001)
gcc.gnu.org
gcc.gnu.org
If the selected floating-point hardware includes the NEON
extension (e.g. -mfpu=‘neon’), note that floating-point
operations are not generated by GCC's auto-vectorization
pass unless -funsafe-math-optimizations is also specified.
This is because NEON hardware does not fully implement the
IEEE 754 standard for floating-point arithmetic (in
particular denormal values are treated as zero), so the use
of NEON instructions may lead to a loss of precision.This is a pretty interesting email chain and a fairly good case study for those of us who aim for zealous adherence to standards and rules vs. providing a practical and useful output for users. I know I tend to go zealous by default and often have to remind myself that I'm not just writing software to satisfy my own academic ego.
Compilers and CPUs have made huge strides since then. I'm pretty sure that -ffast-math makes a much bigger difference now.
Now have a look at the codegen [2] without and with -ffast-math.
With -ffast-math the compiler rewrites the Kahan summation algorithm to a simple sum, completely screwing the code's semantic (since we told the compiler to do so!). Keep this in mind.
1 https://en.wikipedia.org/wiki/Kahan_summation_algorithm 2 https://godbolt.org/g/uWZgW2
What's going on here? Quake 3 port to linux in 1996? Can this sentence be read otherwise? Did he possibly mean Quake 1?
The mail is from July 2001. Q3 was released in early 1999, that's already 2 years and 3 months.
Now, Q2 was released in 1997. Q3 could have started in development immediately afterwards, so that pushes it to 4 years.
Not that much of a stretch to call them 'five years', especially casually speaking. And if you're a little older (over 30) it's even easier to have a fuzzier feeling for time (for a 20 year old, 5 years is 25% of their life -- for a 30 year old just 16%, and even whole decades start to have less defined edges, especially combined work/family routine).
Or it could be a typo.
>> The main FP work I did on that thing on the alpha improved the framerate by about 50% on alpha - FP was _that_ critical for it.
I can't seem to find any evidence of one, so I am thinking it must be Q2, which was ported to Linux and NT Alpha. Jeez, just seeing that gave me a pang of nostalgia; I never got to work on an Alpha machine (spent plenty of time on SPARC though) but I always wanted to. I remember people being really excited about Alpha, thinking it was The Future™.
That was back in the days where it was hard to imagine that in 20 years the vast majority of computers would be either running on 8086 or Acorn descendants...
Edit: we must agree on a max deviation/error for this type.
Then again, GPUs now do strictly follow IEEE and certainly not for games sake. That didnt use to be the case.
In fact they even support denormal numbers at full speed (not even Intel does that).
They often lag behind or are cut down (fewer pipelines, lower clocks) for lower power consumption, but they aren't entirely different designs as far as I can tell.
And not all FPU loads are amenable to GPU computation.
[1] https://gcc.gnu.org/ml/gcc/2001-07/msg01877.html
[2] and many agree with him, Java for example moved away from IEEE around the same time this was written.
Linus is very pragmatic in this exchange, in fact, hard numerical code, as the other person is implying in the discussion, is normally not following IEEE. Especially around zero. If your code is losing precision when you divide by something very big resulting in a zero value, you do not want to rise an exception, you just want to have zero.
Numerical code is normally not following the full IEEE and ignoring the underflow and inexact exceptions.
IEEE is a lot about tracking precision loss but in numerical code, precision loss is acceptable and needed (for speed) in many situations. But knowledge of the precision loss is for some people very important, this is the raison d'être of the IEEE standard.