1,247 karma · joined September 24, 2017
In short, the out of order instruction buffer can do some amazing stuff for code that would otherwise run much slower, that doesn't mean you can't gain or lose performance by reordering instructions. For non-trivial code the best composition is almost certainly different between CPUs.
Because it is the only thing we have sufficiently sandboxed so that we can actually run it untrusted.
I totally get what you want, but we already have a bunch of X-to-JavaScript compilers that pretty much all deliver the experience of: "It works, but it is a bit gnarly, so why would you use this when you can just write JavaScript". And I don't think WebAssembly is going to change that experience much, mainly because it is a pretty poor VM, seemingly designed around the misconception that a lack of features makes it a good compilation target.
I honestly think we could give compiled languages a much better time on the web with a new intermediate language, but someone has to design it for that purpose rather than cargo cult Assembly similarity.
In any case, the DOM and CSS is like 90% of the reason to hate the basic web stack. The other 90% of the reason that people hate JavaScript for is the myriad of frameworks that some people opt to use. You don't have to, and if you come from C you will probably feel right at home not doing so.
The current state of the art is trying to prove that the thing did something quantum. Programmability is at best swapping some wires around to change that something.
In general modern military seems to be focused on delivering as much power as possible in a single unit, never bothering to ask if multiple lesser units could do a better job combined. And so we get destroyers and cruisers that look mighty, but can nevertheless be rendered inoperable, if not outright sunk, by a single torpedo.
JavaScript is not machine code, but still a good deal harder to make fast than a language designed for fast sandboxing. Of course there have been bugs, but mostly I think the JS VMs have done a pretty good job of protecting browsers.
The switching between read and write part doesn't make much sense, there is no such switch in cache, and the memory controller takes care of chunking memory access reasonably. Reading and writing is double the operations of just reading of course, but we have covered that already.
What I'm suggesting that the code might do is: At every cycle, read data on location i and location i-1, add them together, store the result on location i. This is problematic because not only do we get an extra read, it also happens on a cache location that was just written to. There is a guard for preventing this unnecessary read, but it might have failed due to not being accessed as the same pointer with the same offset. We still keep it in cache, so a cost of 2 extra cycles per iteration seems reasonable.
Not knowing what cpu compiler and settings are used I can't do much to replicate the test.
The one you should use, AES-GCM, wasn't even in the original box, but was MacGyvered later.