I wonder how true this still is. I've been having fun simply assuming it is, and these teardown articles and the other HN discussions around them have been very helpful.
I wonder how true this still is. I've been having fun simply assuming it is, and these teardown articles and the other HN discussions around them have been very helpful.
Yes, on modern chips float math will generally be faster than fixed point. This is not so much because the integer units get clogged, as that there's a huge amount of chip area and optimization that goes into the float units (often SIMD, and a lot of FM synthesis can benefit from this, though feedback creates data dependencies). For example, multiply-and-add is usually one cycle in float, but would always be two separate instructions in integers.
My recollection is that older ARM chips have a special issue with latency of data dependencies originating from the float unit (NEON, which is optimized for SIMD vector operations) to the integer unit. I suspect this is no longer the case, or is less of an issue.
Unless the multiply is by a constant 2, 4, or 8, in which case you can use `lea`.
Interesting to test out on the ARM Mac, and see if different dependency chains show significant latency penalties / in with reorder buffer.
The best current info I could find for the latency advice is [2]. Quoting, "Moving data from NEON to ARM registers is Cortex-A8 is expensive." Looking at [3] partially reveals the reason why: the NEON pipeline is entirely after the integer pipeline, so moves from integer to NEON are cheap, but the reverse direction is potentially a large pipeline stall. This is an unusual design decision that as far as I know is not true for any other CPUs. Edit: I found [4], which is a more authoritative source.
[1]: https://github.com/google/music-synthesizer-for-android/blob...
[2]: https://community.arm.com/support-forums/f/armds-forum/757/n...
[3]: https://www.design-reuse.com/articles/11580/architecture-and...
[4]: https://developer.arm.com/documentation/den0018/a/Optimizing...
For Cortex-A8 from [4] and the others you have linked, It makes sense to me now regarding the instruction passing data between registers, filling out the pipeline and then stalling.
Will have a peek at ARMv8/ARMv9 arch's and see what they did there regarding SVE/SVE2.
I'm specifically targeting an Intel Atom, which is obviously powerful enough for the task, but may not fit all definitions of "modern" now?
In a phase accumulator (or any numerical representation of time) it is generally desirable to have uniform precision across the full phase range. Floating point arithmetic does not have this property. On the other hand, floating point arithmetic can sometimes be more convenient than fixed-point if the hardware to hand can execute it fast. But if you're designing hardware from scratch, and especially for FM synthesis, you're probably going to find some other tricks that work even better.
bwahahaha.
I'm a DAW author. Have been for 24 years or so. My friends write plugins and DSP modules for mixing consoles.
Received wisdom from KVR Audio (and in fact, most online forums) is worth less than the distance than a flea could throw it.
Digital audio software users that know almost nothing about the subject seem highly inclined to spend their time blathering on in these forums, while the people who actually do know about it appear to have better things to do.
Depends on the forum, I guess. Lately I've been following the developer subforum of KVR, where folks talk about filter and oscillator algorithms using math I'll never understand. Discussions there seem well-reasoned and entirely civil. I haven't really seen them talk about performance optimizations, though. And I can't say anything about the rest of KVR.