Own Constant Folder in C/C++
neilhenning.dev
neilhenning.dev
The problem is that you end up promoting y from an integer to an f64 and get a much slower operation. I ended up writing my own `rem_i64(self: &f64, divisor: i64) -> f64` routine that was some ~35x faster (a huge win when crunching massive arrays), but as there are range limitations (since f64::MAX > i64::MAX) you can't naively replace all call sites based on the type signatures. However, with some support from the compiler it would be completely doable anytime the compiler is able to infer an upper/lower bound on the f64 dividend, when the result of the operation is coerced to an integer afterwards, or when the dividend is a constant value that doesn't exceed that range.
So now I copy that function around from ML project to ML project, because what else can I do?
(A “workaround” was to use a slower-but-still-faster `rem_i128(self: &f64, divisor: i128) -> f64` to raise the functional limits of the operation, but you're never going to match the range of a 64-bit floating point value until you use 512-bit integral math!)
[0]: https://github.com/rust-lang/rust/issues/83973
Godbolt link: https://godbolt.org/z/EqrEqExnc
Aren't you supposed to make a crate? Which is copying with more steps, but might make your life easier (or harder).
Yes, I suppose that would be the canonical way to go. But it's just one function; four lines! I'm getting isEven() vibes!
If that crate becomes successful, it will basically greenlight a lot of functionality that was just there because it happened to be made by the same author. And all that extra functionality will only reduce the chances of the really useful function to get into the spotlight. Both are bad for the ecosystem.
If the GP has related useful little functions, yes, pack them together. Otherwise, I'd say a crate for a small little function isn't a problem at all.
My point is that I keep my own little collection of useful code. I don’t publish it as a package or crate. It’s what I suggested to the other poster.
With "own constant folder" I expected a GCC/clang plugin that does tree/control flow analysis and some fancy logic in order to determine when, where and how to constant fold…
How is a non-expert in the language supposed to learn tricks/... things like this? I'm asking as a C++ developer of 6+ years in high performance settings, most of this article is esoteric to me.
Whatever the language, at some point in performance tweaking you will end up having to look at the assembly produced by your compiler, and discovering all kinds of surprises.
This is nonsense, but it's really common, distressingly common, for C and C++ programmers to use this sort of inappropriate global modelling. It's something which cannot scale, it works OK for one man projects, "Oh, I use the Special Goose Mode to make routine A better, so even though normal Elephants can't Teleport I need to remember that in Special Goose Mode the Elephants in routine B might Teleport". In practice you'll screw this up, but it feels like you'll get it right often enough to be valuable.
In a large project where we're doing software engineering this is complete nonsense, now Jenny, the newest member of the team working on routine A, will see that obviously Special Goose Mode is a great idea, and turn it on, whereupon the entirely different team handling routine B find that their fucking Elephants can now Teleport. WTF.
The need to never do this is why I was glad to see Rust stabilize (e.g) u32::unchecked_add fairly recently. This (unsafe obviously) method says no, I don't want checked arithmetic, or wrapping, or saturating, I want you to assume this cannot overflow. I am formally promising that this addition is never going to overflow, in order to squeeze out the last drops of performance.
Notice that's not a global flag. I can write let a = unsafe { b.unchecked_add(c) }; in just one place in a 50MLOC system, and for just that one place the compiler can go absolutely wild optimising for the promise that overflows never happen - and yet right next door, even on the next line, I can write let x = y + z; and that still gets the kid gloves, if it overflows nothing catches on fire. That's how granular this needs to be to be useful, unlike C++ -ffast-math.
From a practical point of view it is fine.
It had also been fixed recent GCC versions.
Fast math is simply saying "I care even less than IEEE"
This is perfectly appropriate in many settings, but _especially_ video games where such deterministic results are completely irrelevant.
I'm not sure I'd agree. Off the top of my head, a potentially significant benefit to deterministic floating-point is allowing you to send updates/inputs instead of world state in multiplayer games, which could be a substantial improvement in network traffic demand. It would also allow for smaller/simpler cross-machine/platform replays, though I don't know how much that feature is desired in comparison.
(Some of those are single-platform or lack cross-play support, and thus only need consistency between different machines running the same build. That makes compiler optimizations less of an issue. However, some do support cross-play, and thus need consistency between different builds – using different compilers – of the same source code.)
The perception of nondeterminism came specifically from x87, which had 80-bit native floating-point registers, which were different from every other platform's 64-bit default, and forcing values to 64-bit all the time cost performance, so compilers secretly turned data types to different ones when compiling for x87, therefore giving different results. It would be like if the compiler for ARM secretly changed every use of 'float' into 'double'.
The post can be boiled down to "Clang doesn't compile this intrinsic nicely, so just use inline asm directly. But remember that you need to have a non-asm special case to optimize constants too, and you can achieve this with __builtin_constant_p".
[1]: https://cdrdv2.intel.com/v1/dl/getContent/814198?fileName=24...
You can/would just use it in the translation units where you want it; usually for numerical code where you want certain optimizations or behaviors and know that the tradeoffs are irrelevant.
It's mostly harmless for everday application math anyway, and so enabling it for your whole application isn't a catastrophe, but it's not what people who know what they're doing would usually do. It's usually used for a specific file or perhaps a specific support library.
My understanding is that if you don't specify -ffast-math when linking then you shouldn't get crtfastmath.o linked in.
Most other language ecosystems most likely suffer from similar problem if you look under the hood.
At least compilers like Clang also give you the tools to workaround such issues, as demonstrated by the article.
just like everything else in life that's complex: slowly and diligently.
i hate to break it to you - C++ is complex not for fun but because it has to both be modern (support lots of modern syntax/sugar/primitives) and compile to target code for an enormous range of architectures and modern achitectures are horribly complex and varied (x86/arm isn't the only game in town). forgo one of those and it could be a much simpler language but forgo one of those and you're not nearly as compelling.
By learning C and inline asm. For a C developer, this is nothing out of the ordinary. C++ focuses too much on new abstractions and hiding everything in the stdlib++, where the actual implementations of course use all of this and look more like C, which includes using (OMG!) raw pointers.
but yes, given the premise of the article is that a friend wants to use a specific CPU instruction, yeah, at least minimum knowledge of one's stack is required (and usually the path leads through Assembly, some C/Rust and an FFI interface - like JNI for Java, cffi/Cython for Python, and so on)
https://www.youtube.com/watch?v=5ZOuCuGrw48 (and here's the 2009 version https://www.infoq.com/presentations/click-crash-course-moder... .. might be interesting for comparison )
My understanding is that C is a great language, but I also get that its not for everyone. Its really powerful, and yet you can easily make mistakes.
For me, I'm just learning how to use C, I'm not trying to understand the compiler or make files yet. From what I get, the compiler is how you can achieve even better performance, but you need to understand how it is doing its black magic....otherwise you just might make your code slower or more inefficient.
As for the compiler’s role in C, it’s equivalent to javac - it’s taking your source and creating machine code, except the machine code isn’t an abstract bytecode but the exact machine instructions intended to run on the CPU.
The issues with C and C++ are around memory safety. Practice has repeatedly shown that the defect rate with these languages is high enough that it results in lots of easily exploitable vulnerabilities. That’s a bit more serious than a personal preference. That’s why there’s pushes to shift the professional industry itself to stop using C and C++ in favor of Rust or even Go.
Sure you can mess up your performance by picking bad compiler options, but most of the time you are fine with just default optimizations enabled and let it do it's thing. No need to understand the black magic behind it.
This is only really necessary if you want to squeeze the last bit of performance out of a piece of code. And honestly, how often dies this occur in day to day coding unless you write a video or audio codec?
* mtune/march - specifying a value of native optimizes for the current machine, x86-64-v1/v2/v3/v4 for generations or you can specify a specific CPU (ARM has different naming conventions). Recommendation: use the generation if distributing binaries, native if building and running locally unless you can get much much more specific
* -O2 / -O3 - turn on most optimizations for speed. Alternatively Os/Oz for smaller binaries (sometimes faster, particularly on ARM)
* -flto=thin - get most of the benefits of LTO with minimal compile time overhead
* pgo - if you have a representative workload you can use this to replace compiler heuristics with real world measurements. AutoFDO is the next evolution of this to make it easier to connect data from production environments to compile time.
* math: -fno-math-errno and -fno-trapping-math are “safe” subsets of ffast-math (i.e. don’t alter the numerical accuracy). -fno-signed-zeros can also probably be considered if valuable.
In my case it was by accident as I picked up assembly and machine language before I touched C in the late 1980s.
The author is talking about a way to get a particular (version of) C/C++ compiler to emit the desired instruction. So I'd call this clang-18.1.0-specific but not C/C++-specific since this has nothing to do with the language.
Also such solutions are not portable nor stable since optimization behavior does change between compiler versions. As far as I can tell, they also would have to implement a compiler-level unit test that ensures that the desired machine code is emitted as toolchain versions change.
If you had an application where this sort of thing made a difference in JavaScript, the problem would likely still the there, you’d just have a lot less visibility on it.
I guess you’re still right - at the end of the day you see discussions like this far more often in C, so it impacts the feel of programming in C more.
I wish there was some sensible way for code that’s purely about optimization to live entirely separated from the code that’s about solving the problem at hand…
Clang transforms sqrtps(x) to x * rsqrtps(x) when -ffast-math is set because it's often faster (See [1] section 15.12). It isn't faster for some architectures, but if you tell clang what architecture you're targeting (with -mtune), it appears to make the right choice for the architecture[2].
[1]: https://cdrdv2.intel.com/v1/dl/getContent/814198?fileName=24...
[2]: https://godbolt.org/#g:!((g:!((g:!((h:codeEditor,i:(filename...
Back in the days the answer was `__builtin_constant_p`.
But with C++20 it is possible to use std::is_constant_evaluated, or `if consteval` with C++23.
But this is for scenario when you still want to keep high quality of code (maybe when you write multiprecision math library; not when you hack around compiler flag), which inline assembly violates for many reasons:
1) major issue: instead of dealing with `-ffast-math` with inline asm, just remove `-ffast-math`
2) randomly slapped inline asm inside normal fp32 computations breaks autovectorization
3) randomly slapped inline asm in example uses non-vex encoding, you will likely to forget to call vzeroupper on transition. Or in general, this limits code to x86 (forget about x86-64/arm)
4) provided example (best_always_inline) does not work in GCC as expected
The square root and div instructions used to be a lot slower than they are now.
If you really want to relax the correctness of your code to get some potential speedup, either do the optimizations yourself instead of letting the compiler do them, or locally enable fast-math equivalents.
As for the is constant expression GCC extensions, that stuff is natively available in standard C++ nowadays.
Compile-time execution emulates the target architecture, which has pros and cons. The most notable con is that inline assembly isn't supported (yet?). The complete solution to this particular problem then additionally requires `return if (isComptime()) normal_code() else assembly_magic();` (which has no runtime cost because that branch is constant-folded). Given that inline assembly usually also branches on target architecture, that winds up not adding much complexity -- especially given that you probably had exactly that same code as a fallback for architectures you didn't explicitly handle.
Most of the code doesn't need nearly that level of optimization, so you write it in a higher level language.
Making your own compiler is a lot of work compared to just doing what I've outlined above. So you don't. I's not worth it, and you'd still end up back where you are.
https://www.intel.com/content/www/us/en/docs/cpp-compiler/de...
That being said, the naming of Intel's vector intrinsics _is_ pretty bad, such as the pi64 vs. epi64 difference. ARM's vector intrinsic naming, despite looking like having come from a cat sitting on the home row keys, is more consistent and has less gotchas.
I never got that impression in my perusal of LLVM bug reports and patches. I wonder if there is an open issue for this specific case.
https://godbolt.org/z/1543TYszP
(intel c++ compiler)
> In Intel microarchitectures prior to Skylake, the SSE divide and square root instructions DIVPS and SQRTPS have a latency of 14 cycles (or the neighborhood) and they are not pipelined. This means that the throughput of these instructions is one in every fourteen cycles. The 256-bit Intel AVX instructions VDIVPS and VSQRTPS execute with 128-bit data path and have a latency of twenty-eight cycles and they are not pipelined as well.
> In microarchitectures that provide DIVPS/SQRTPS with high latency and low throughput, it is possible to speed up single-precision divide and square root calculations using the (V)RSQRTPS and (V)RCPPS instructions. For example, with 128-bit RCPPS/RSQRTPS at five-cycle latency and one-cycle throughput or with 256-bit implementation of these instructions at seven-cycle latency and two-cycle throughput, a single Newton-Raphson iteration or Taylor approximation can achieve almost the same precision as the (V)DIVPS and (V)SQRTPS instructions
(emphasis mine)
> it is possible to speed up single-precision divide and square root calculations using the (V)RSQRTPS and (V)RCPPS instructions
That option is terribly named. It should be -fincorrect_math_that_might_sometimes_be_faster.
julia> f(x) = @fastmath sqrt.(x)
julia > @code_native debuginfo=:none f.(Tuple(rand(Float32, 8)))
vsqrtps ymm0, ymmword ptr [rsi]
mov rbp, rsp
mov rax, rdi
vmovups ymmword ptr [rdi], ymm0
pop rbp
vzeroupper
ret
julia> g() = f((1f0, 2f0, 3f0, 4f0))
julia> @code_native debuginfo=:none g()
.LCPI0_0:
.long 0x3f800000 # float 1
.long 0x3fb504f3 # float 1.41421354
.long 0x3fddb3d7 # float 1.73205078
.long 0x40000000 # float 2
.text
.globl julia_g_600
.p2align 4, 0x90
.type julia_g_600,@function
julia_g_600: # @julia_g_600
# %bb.0: # %top
push rbp
movabs rcx, offset .LCPI0_0
mov rbp, rsp
mov rax, rdi
vmovaps xmm0, xmmword ptr [rcx]
vmovups xmmword ptr [rdi], xmm0
pop rbp
ret
This is almost universally true across all languages. The behaviors you see are usually the result of countless heuristics stacked one atop the other, and a tiny change here can end up with huge changes there.