Fixed Point Arithmetic
vha3.github.io
vha3.github.io
You can actually use _Accum without including stdfix.h, at least for compilers that conform the embedded extension to C (ISO/IEC TR 18037 [1]). Stdfix.h just gives you a macro named `accum` among others; this approach has been used for any new C keyword since C99 (e.g. _Bool vs. stdbool.h, _Alignas vs. stdalign.h). The size of _Accum type is also not exactly defined (it can well be 4.27).
[1] http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1169.pdf
I wrote a multiplayer strategy game back in mid 90's issuing floating point and it was deterministic. No problems. Maybe chip optimizations have affected this?
We looked at fixed point, but with careful scheduling we'd get zero fpu (x87) stalls for float operations, so it wasn't a real win to go fixed. And it gave us the benefit of having more registers to use without needing to use the main stack as much, which also made the asm easier to read.
Edit: typos
Some instructions don't guarantee the maximum possible precision, and implementations can differ in the least significant bits of their result. Addition and multiplication give you the correctly rounded result. However, there is typically some leeway in fma fusion. That is, if you multiply then add, it is in some cases implementation defined whether you get a result that was rounded after the multiplication and before the addition, or you get the correctly rounded result of doing the multiply and add as one operation without intermediate rounding.
All this is with the most conservative compiler settings. Be very, very careful about your optimization settings.
And let's not get into nans.
All this adds up to a situation where writing floating point code that is guaranteed to produce the same bit exact result on all implementations is basically a superhuman task. Better to just use fixed point and be done with it.
tldr: it is hard. I was wrong by saying it is "not possible". But I wonder if it is possible if your target platform includes various consoles?
This is true in the abstract, but not necessarily true of a specific commodity chip. Processor vendors spend a lot of silicon on offering low latency and high throughput floating point support. It's a fairly recent trend of processor vendors adding fast int8 or bfloat16 vectors after the ML craze demonstrated that there was demand for vector support for more bandwidth-friendly datatypes.
While floating point numbers can be cleanly split into a mantissa and exponent, adding floats requires shifting the exponent, which can be an expensive operation. Each portion of a floating point arithmetic operation is also implemented with integer arithmetic, which limits the performance spread. Many floating point operations require significantly fewer bits though, which can lead to major speedups.
On the integer side, the serial carry dependence mean ripple-carry adders are usually a bummer, but there a ton of carry-lookahead variants that can lessen the performance impact of that dependency. If you have several integer additions at once, you can use a carry-save variant to only pay that serial cost once all of the additions are done. Finally, if you are willing to significantly change your integer encoding, there are tools like redundant number systems[0] that allow you to move around the traditional trade-offs for arithmetic circuits including completely removing the serial dependency on carries. That said, if you require rescaling your fixed-point numbers during the computation, the fixed-point implementation will require more integer operations than floating-point operations.
All of this is also quite dependent on the overall architecture of the chip too, since the number of integer units vs float units and how data is moved around can have a way bigger impact on performance than how each arithmetic operation is implemented.
ISTR the reason you couldn't use floats in the kernel had more to do with context switching and the kernel not spilling the FP registers than speed per se, but I might be misremembering.
Edit: Looks like I remembered it basically right:
https://stackoverflow.com/questions/13886338/use-of-floating...
We should all be grateful that most banks still run Cobol, which has native fixpoint support for good reasons.
That's a misconception, before widespread dedicated hardware (eg on-chip 80x87) there was software-emulated floating point arithmetic (eg Microsoft Binary Format) for general use. In contrast, fixed point solutions are very narrow purposed, that is, you choose a specific binary representation format so your specific range of values would fit with acceptable precision.
So you don't care if your bank account is out by a percent or two?
Well, others do.
There are cases where fixed-point is preferred over floating-point, such as implementing an iDFT (or in general in algorithms when the dynamic of the numbers you manipulate is fixed and you need very precise control over the rounding errors).
But the article state: "Problem: I want to do arithmetic with fractional resolution but I can't afford the CPU cycles to use floating point." which isn't true for modern hardware. The rest of the article is still very interesting, but this is just a bad start.
Dead reckoning uses a series of time delta and velocity vector pairs to continuously determine the latest position, right? So in a multiplayer setting, the vector is received over the network and due to unpredictable latency, the time value has to be as well. At that point, why not just send the new position in absolute coordinates? This would avoid any FP arithmetic inaccuracies between devices, especially the kind that accumulate over time as they would by continuously adding vectors on top of each other.
If you can run the same simulation on two different machines using the same sequence of input states from all players and get different results, that's a bug. One way to get this kind of bug is for some parts of the internal state of the game to rely on floating point computations.
https://forums.swift.org/t/pitch-clock-instant-date-and-dura...
Ignoring the reduced range, you can perfectly store integers in FP variables. Even banks could store everything in FP instead of integer variables in the lowest unit they support (for example cents). Integers at overflow fail catastrophically, so that also has to be controlled for. It is not the case that the loss of precision for large FP values is fundamentally worse, though it does happen earlier for the same total size of variable.
Integers and fixed point save space if you need a large range. The exponent of a fixed point is stored in the code, instead of in the variable. Integer instructions are encoded into shorter binary form, because they were part of the introduced set from the beginning (if you disregard the extra shifts required for fixed point). These are the advantages as far as I can tell.
For example if you have 10 buildings that gives 0.1 crystal everyday, you would gain exactly 1 crystal per day. In the user interface you wouldn't see something like "0.999 crystal per day". And it is not just for pretty tooltips, if there is a game logic that checks ">= 1.0 crystal per day" then that condition is guaranteed to be satisfied after 10 crystal buildings
Sometimes. Rarely. Specifically, it's a performance win when you don't need 32 bit of precision, when 8 or 16 bits is enough in these integers. When it happens, you can pack twice as many lanes in the SIMD vectors, and use instructions which view 16-byte vectors as 8 int16_t lanes, or 16 uint8_t lanes.
FP32 is much faster to multiply and especially to divide, compared to Int32.
> Do games and similar software use them today?
Image and video processing code which for some weird reason runs on CPU definitely does.
Even better times. :)
Can't find a more official source than Reddit, but here goes:
> It's because the position of each polygon's vertex (corner) is only calculated at a very low precision. Once the polygon moves (or the camera) the vertexes will stay at the same point, until they're closer to the next and suddenly "jump" to the new position without transition. Newer graphics hardware could interpolate smoothly here with more in-between states (that's where all the talk about "floating point precision" came from in graphics).
0. https://www.reddit.com/r/gaming/comments/bkedc/heres_a_quest...
Incidentally, the Nintendo 64 also used fixed point numbers in RDP graphics instructions, but did not exhibit the same visual artifacts as the PlayStation.
Using the CMSIS-DSP Q31 functions, in particular. On Rust `i32`s.
Or alternatively (and often more simply) by setting and clearing a GPIO pin before/after the operation being timed, and then using the oscilloscope to measure execution time
Mid-range smartphones will generally have acceptable floating point units, and you should stick to floating point math. If you're targeting low end smartphones, you never know what you're going to get.
Fixed point is generally only still useful on embedded. If you're building software for fixed point, it's generally because you know exactly what CPU your software is going to run on, and you know it doesn't have a FPU.
Also multimedia. Modern CPUs compute FP32 floats equally fast as integers. However, when you only need 8 or 16 bits of precision, RAM bandwidth often dominates computations. Most image and video codecs are still using 8 bits per channel.
Profit from fixed point can be quite large for these use cases. That’s assuming the implementation is good, with SSE2, AVX2 or NEON SIMD.