They jump through hoops so that C programs and the programs of all the other languages designed to run in the same execution model run as fast as possible. It's not to indulge C programmers but to support the vast body of existing software.
Hardware engineers need optimization targets just liker anyone else, and it's an eminently reasonable one.
Also, it's not like loads of improvements not related to the C execution model haven't been made. Demonstrably, the C model is not holding us back.
Anyway, if you want to move on from the C execution model, great. Now you need to introduce a practical transition plan, which should include such details as how and why we should rewrite all of the existing performance sensitive software designed to work in the old model.
We need to move beyond memory-unsafe paradigms already. Assembly intrinsics aren't great but they're certainly better than C.
Never use it, myself, except configuring the date format in my menu bar.
Because the speed advantage is so huge, people went through the pain of learning this new model and redesigning algorithms to better fit it.
So even with GPUs you have a similar situation..
The current microarchitectures for general purpose computing, that is OoO execution, cache coherent shared memory with a flat memory model, and mostly transparent and coherent caches seems to be optimal. In fact, far from being designed around C, C and siblings had to evolve to support the model well (a memory model, explicit SIMD builtins, support for vectorization and offloading etc).
It is entirely possible this is only a local optimum, but I have yet to see a plausible model for a better architecture.
It is more likely that, now that silicon is cheap and CPU designers are struggling to find ways to use it, extensions to support higher level languages might be added: more fine grained cache control for message passing, extra tags to help GCs, and more stuff that I can think of.
If anything, more modern languages (eg Java) have to contort their data representations to build, say, “structs of arrays” that are highly efficient on modern CPUs (and trivial to express in C!)
I’m very curious to hear of any recent languages that are designed expressly with “mechanical sympathy” in mind.
[2] https://ispc.github.io/ispc.html#the-ispc-parallel-execution...
[3] https://ispc.github.io/ispc.html#structure-of-array-types
Without cache coherency and ILP, programming in any language would be insane. It's not like programming in x86_64 assembly suddenly opens a world of possibilities because its "truly low-level". You gain very little extra control over a given platform by switching to assembly over C, that's what we mean by "low-level".
C maps cleanly onto the instruction sets provided by chip manufacturers. It provides the option to the programmer to optimize structures for use in vectorization if they so choose, or to optimize for some other objective like size for a binary wire protocol or limited memory space.
The very nature of having a choice about memory layout of structures and the ability to cleanly link with the platform ABI is what makes C low-level. Obviously the inner-workings of a modern CPU don't map cleanly to the C virtual machine. However, there's no convincing evidence that greater control over cache invalidation is what's holding back performance of those CPUs.
It certainly opens a world of possibilities regarding the vector units. That's one place where inner loops hand-written in assembly still have an edge. On the other hand, with contemporary CPUs, high quality assembly coding is a highly specialized job skill by itself, so for general purpose engineers, learning it is most likely to yield a bad cost/benefit.
The chips move heaven and earth to maintain the fiction that those instructions actually direct what they do.
The number of fundamentally different kinds of cache, and specialized state machines not directly accessible by instructions, in a modern chip would boggle your mind.
So C is too high level for GPU programming, and needed to be extended.
e.g.
>Why is there no pointer arithmetic?
>Safety. Without pointer arithmetic it's possible to create a language that can never derive an illegal address that succeeds incorrectly. Compiler and hardware technology have advanced to the point where a loop using array indices can be as efficient as a loop using pointer arithmetic.[1]
I wish this meme about Go and generics would die already. C allows you to define an array of int and and an array of char and have those be two different types. Go allows you to do this with slices and maps as well, because slices and maps are also built in types in Go. This is not "generics".
[1] With the largely irrelevant exception of the C11 _Generic stuff.
That isn't to say that the article doesn't have a point though, far from it.