I mean that a naive loop that traverses your data does not take several minutes like in Python. But there's a guy in my lab who takes my stupid loops, avxs the fuck out of them, and they become 8 times faster. Now, that's what you would call unreasonably fast (short of rewriting the whole thing for the GPU).
GHC with `-fllvm` is not going make Haskell any easier to compile just because it's targeting LLVM. Fortran is (relatively!) easy to make fast because the language semantics allow it to. Lack of pointer aliasing is one of the canonical examples; C's pointer aliasing makes systems programming easier, but high-performance code harder.
(See https://github.com/llvm/llvm-project/blob/main/flang/docs/Al... for the full story.)
Thanks for the precise Fortran terminology; that's what I meant to say in my head but you're correct. From the linked website for everyone else:
"Fortran famously passes actual arguments by reference, and forbids callers from associating multiple arguments on a call to conflicting storage when doing so would cause the called subprogram to write to a bit of that storage by means of one dummy argument and read or write that same bit by means of another."
Just an example on how to declaring a couple of input float pointers in a fully optimized way, avoiding aliases and declaring it readonly (if I remember correctly, it's been a few years:
Fortran:
real(32), intent(in) :: foo, bar, baz
C: const float *const restrict foo
const float *const restrict bar
const float *const restrict baz
of course what happens then is that people will do a typedef and hide it away... and then every application invents its own standards and becomes harder to interoperate on a common basis. in Fortran it's just baked in.Another example: multidimensional arrays. In C/C++ it either doesn't exist, is not flat memory space (and thus for grid applications very slow), or is an external library (again restricting your interop with other scientific applications). In Fortran:
integer, parameter :: n = 10, m = 5
real(32), dimension(n, m) :: foo, bar, baz
Again, reasonably simple to understand, and it's already reasonably close to what you want - there are still ways to make it better like memory alignment, but I'd claim you're already at 80% of optimal as long as you understand memory layout and its impact on caching when accessing it (which is usually a very low hanging fruit, you just have to know which order to loop it over).Can you clarify what you mean by this? A naively defined multidimensional array,
float foo[32][32] = ...
is stored in sizeof(float) * 32 * 32 = 4096 consecutive bytes of memory. There's no in-built support for array copies or broadcasting, I'll give you that, but I'm still trying to understand what you mean by "not flat memory space".However, even then it's not quite true given C99 VLA.
int a, b;
int arr[a][b]; // ERROR: all array dimensions except for the first must be constant
//or even:
auto arr = new int[a][b]; // ERROR: error: array size in new-expression must be constant
All you can do if you want a dynamic multidimensional array is: int a, b;
auto arr = new int*[a];
for (int i = 0; i < a; i++) {
arr[i] = new int[b];
}
But now it's a jagged array, it's no longer flat in memory.Could you say like,
auto arr = new int[a*(b+1)];
int \* arr2 = (int\*)arr;
for(int i=0;i<a;i++){
arr2[i]=arr+(i*(b+1))+1;
}
and then be able to say like arr2[j][k] ?(Assuming that an int and a pointer have the same size).
It still has to do two dereferences rather than doing a little more arithmetic before doing a single dereference, but (other than the interspersed pointers) it’s all contiguous? But maybe that doesn’t count as being flat in memory.
The thing is that, in Fortran (and even in C), you don't need any of this because the construction is part of the language itself.
However, this would still mean that you need to do 2 pointer reads to get to an element (one to get the value of arr2[i], then another to get the value of arr2[i][j]). In C or Fortran or C++ with compile-time known dimensions, say an N by M array, multiArr[i][j] is a single pointer read, as it essentially translates to *(multiArr + i*M + j).
foo: &f32,
bar: &f32,
baz: &f32,
There are no dynamically-sized multi-dimensional flat memory arrays though, at least not in the core language or standard library.See also CUDA. Sure you can write a C to CUDA converter auto-vectorizer but its likely to have all sorts of bugs and usually never work right except in rare hand-tune cases. May as well just write CUDA from scratch if it is to be performant. Same for array processing, wanna array process? Use compiler for array processing like Fortran or ISPC.
Most people, even good C programmers, don’t write that kind of C, though.
The point of Fortran is that you can write Fortran code that is almost as fast as that nightmare C, but you can do it and still get your degree on time.