Let the Compiler Do the Work
cliffle.com
cliffle.com
Slightly off-topic, but it's important to point out that in modern C, implementing such a function as a macro is always a mistake. You'd do it as an inline function.
Macros in modern C should only be used for code generation (for DRY). If the language supports doing it without a macro, then do it without a macro.
Don't use a macro where a function will do.
fn sqr<T: Copy + Mul>(x: T) -> T::Output {
x * x
}Unfortunately there are always tradeoffs using macros for something like that, you either end up evaluating the parameters more than once or you have to introduce variables in the macro that may shadow existing variables. Some compilers have extensions to help with that: https://gcc.gnu.org/onlinedocs/gcc/Statement-Exprs.html
You certainly don’t have to introduce the risk of shadowing existing variables. You just use the standard
#define MACRO(x) \
do { \
/* Define whatever variables you like */ \
/* Do things... */ \
} while (0)Second, you still risk shadowing existing variables, even if you make an effort to use "unlikely" names for the locals such as "___my_macro_secret_variable_1", if the macro might be used in one of its own arguments. For example MACRO(MACRO(x)).
If that sounds unlikely, consider MAX(a,MAX(b,c)), which is likely to happen eventually, if your codebase uses such macros or if they are part of a library.
#define maxint(a,b) \
({int _a = (a), _b = (b); _a > _b ? _a : _b; })
>Note that introducing variable declarations (as we do in maxint) can cause variable shadowing [...] this example using maxint will not [produce correct results]: int _a = 1, _b = 2, c;
c = maxint (_a, _b);I hope I get to use C2x* before I retire.
*postmodern C?
AFAIK the only non-supported C99 features in the C compiler are VLAs (optional since C11 anyway), and type-generic macros (those would be good to have though).
Of course it would be nice if Microsoft gave the C compiler a bit more love, especially since it's much less work keeping a C compiler uptodate than a C++ compiler, but at least we got "most of C99".
Although inline was added in C99, it was already an extremely widely supported extension in mainstream compilers, even since the C89 days, when we just called it "ANSI C".
MSVC has supported inline for a long time, long before it started supporting other C99 features.
I could understand the concern if it was about portability to other target platforms, or keeping the option of doing so. But in that case, the public standard your current target supports is irrelevant.
So they add a feature not supported by MSVC and don't learn that it doesn't work until someone else tries to build on Windows.
If you choose to use features based on whether they work or not, you don't need to choose a standard at all. But that loses you all of the guarantees a standard provides.
For projects that will never have more than ~1 million LOC it's probably fine. Less than ~100K, definitely preferable.
Basically just wondering if anyone has first hand experience using this kind of thing on a large project.
Sometimes you do want to debug a release build because the optimizations have gone wrong, though. In that case, it's helpful if release builds build quickly, and especially if they rebuild quickly as you turn various sub-optimization flags on and off.
No need to move things to headers (no headers at all, actually), worry about forward declarations and exposing internal bits of the API etc... Just write the code normally and naturally and let the compiler figure it out. No need to consider subtle performance tradeoffs when deciding where to write the code, just put it where it makes sense.
In what way do you think it is more limited?
Non-portable, of course, I get.
Not all intrinsics are memory safe
I believe there are plans to expose higher-level functionality as a safe interface at a later date (the API design work just hasn't been done yet). For now you can get this functionality as a 3rd-party crate https://github.com/AdamNiederer/faster
1. Using instructions that don't exist on a given CPU will cause the program to crash.
2. `core::arch` exists so that stable Rust programs could make use of SIMD instructions for performance (in particular, `regex` crate). They are very low level, and there is a lot of those, so whether any given SIMD intrinsic is safe or not wasn't considered as those functions exist for safe SIMD abstractions to use - which is advantageous, as Rust cannot really do breaking changes, while a library could easily release a new version without breaking the users (which will still use the old version).
But surely this also applies to a Rust program that only uses safe code but it's compiled with the right compiler switches that is allowed to use SIMD instructions?
AFAIK the main way to manage this right now is throwing the optimized stuff into a plugin and then loading it at runtime— but that ends up having huge implications on your project structure, source layout, etc.
Is important to note what is the POV of "unsafe". Is not unsafe for you or for your setup.
Is unsafe from the COMPILER ie: it can't PROVE is safe!
Static analysis can only decide a subset of the possible "safe" interactions. In the time of COMPILING. Rust can't decide if use of SIMD is safe AT COMPILE TIME because at RUNTIME maybe the cpu don't have it!
Or at least the other way around: if a function is declared with dynamic feature detection and a default path exists why can't it be declared safe by the compiler?
https://github.com/jackmott/simdeez
The thing is that is not yet in-build in the lang. Let the crates sort some problems before commit it in the lang is like part of the philosophy of rust.
For example, rust change the hash implementation after certain crate prove it worthy.
The fact that default C ABI prefers to work without a condom here doesn't mean that Rust has to! This is one of the most avoidable forms of undefined behavior I'm aware of.
I know this isn't a 100% panacea, as some programs will have multiple bits of inline asm for different system and some sort of better or worse way of determining at run-time what to run. However, I've written plenty of programs where it pretty much says on the tin: "you must have AVX2 to run this".
All up this seems like a pretty trivial problem compared to the problem of 'lack-of-safety' in general, and I'm not thrilled to see SIMD intrinsics being put in the "OMG So Dangerous" box as if they are a bunch of wild pointer ops (except, of course, in the case where they actually are - e.g. scatter/gather)...
[0]: https://doc.rust-lang.org/std/primitive.f64.html#method.powi
That said, I can't speak to whether the Rust compiler wouldn't just optimize that away -- it seems like unrolling exponentiation into multiplication for small constant powers would be a very safe and easy thing to do.
Not just small constant powers either. I tried .powi(1000000) and it compiled into a sequence of 25 vmulss instructions.
I suspect then that defining it at x*x is just just because it's easier to type than x.powi(2).