Is this article really confused, or did I misunderstand it?
The thing that makes C/C++ a good language for SIMD is how easily it lets you control memory alignment.
Is this article really confused, or did I misunderstand it?
The thing that makes C/C++ a good language for SIMD is how easily it lets you control memory alignment.
Not sure for C
Restrict does too.
In short, while your statement is true in general, I believe that it is not applicable in the context of the discussion.
The fact that exp is implemented in hardware is not the argument. The argument is that exp is a library function, compiled separately, and thus the compiler cannot inline the function and fuse it with an array-wide loop, to detect later on an opportunity to generate the SIMD instructions.
It is true however that exp is implemented in hardware in the X86 world, and to be fair, perhaps a C compiler only needs to represent that function as an intrinsic instead, to give itself a chance to later replace it either with the function call or some SIMD instruction; but I guess that the standard doesn't provide that possibility?
I gesture to this in the blog post:
> In C, one writes a function, and it is exported in an object file. To appreciate why this is special, consider sum :: Num a => [a] -> a in Haskell. This function exists only in the context of GHC. > ... > perhaps there are more fluent methods for compilation (and better data structures for export à la object files).
This is especially surprising given that very few architectures have an `exp` in hardware. It's almost always done in software.
The function exp has an abstract mathematical definition as mapping from numbers to numbers, and you could implement that generically if the language allows for it, but in C you cannot because it's bound to the signature `double exp(double)` and fixed like that at compile time. You cannot use this function in a different context and pass e.g. __m256d.
Basically because C does not have function dispatch, it's ill-suited for generic programming. You cannot write an algorithm on scalar, and now pass arrays instead.
It is limited to cases where you know all overloads beforehand and limited because of C’s weak type system (can’t have an overload for arrays of int and one for pointers to int, for example), but you can. https://en.cppreference.com/w/c/language/generic:
#include <math.h>
#include <stdio.h>
// Possible implementation of the tgmath.h macro cbrt
#define cbrt(X) _Generic((X), \
long double: cbrtl, \
default: cbrt, \
float: cbrtf \
)(X)
int main(void)
{
double x = 8.0;
const float y = 3.375;
printf("cbrt(8.0) = %f\n", cbrt(x)); // selects the default cbrt
printf("cbrtf(3.375) = %f\n", cbrt(y)); // converts const float to float,
// then selects cbrtf
}
It also, IMO, is a bit ugly, but that fits in nice with the rest of C, as seen through modern eyes.There also is a drawback for lots of cases that you cannot overload operators. Implementing vector addition as ‘+’ isn’t possible, for example.
Many compilers partially support that as a language extension, though, see for example https://clang.llvm.org/docs/LanguageExtensions.html#vectors-....
_Generic((typeof(x), int*: 1, int[]: 2)
\_Generic with type argument is a new feature, it is a bit less elegant without it: _Generic((typeof(x)*)0, typeof(int*)*: 1, typeof(int[])\*: 2)void free(void *ptr);
You use it a lot more than you think ;)
You are however correct, that doesn't mean that C cannot be useful for scientific computing, but I think that there are better alternatives. Although, one can write the core functionalities in C for speed and efficiency, and use another language as a driver.
https://stackoverflow.com/questions/73269625/create-memory-a...
Also structlayout and fieldoffset