Viable ROP-free roadmap for i386/armv8/riscv64/alpha/sparc64
marc.info
marc.info
"Control-Flow Integrity" can have a general sense, or a specific sense. In the general sense, it means that function pointers and return addresses are protected — and WASM protects the raw function pointers.
In code compiled to WASM, a function pointer is represented as an integer indexing into a single global list with all functions. The "call_indirect" op checks only that the index is within bounds and that the function type signature of the function you look up matches.
In the specific sense, as in the 2005 paper titled "Control-Flow Integrity: Principles, Implementations, and Applications", it refers to the code also enforcing that a right function pointer is used at each indirect call site. Each site has a list of allowed function pointers. WASM does not do this.
(Sorry for the long post but I just don't want people to get confused and believe that WASM is safer than it is)
It is also the default on Fuchsia, which therefore supports shared libraries. <https://fuchsia.dev/fuchsia-src/concepts/kernel/safestack>
The problem with these software-based approaches is that it is security-by-obscurity, which breaks if the address to the safe stack would leak. Like ASLR, it is considered more or less broken on 32-bit systems where it is easier to allocate significant portions of the address space to find where it is not and then do educated guesses. However, there have been a few papers using Intel's MPK or even CET to protect it properly, at some performance cost of course.
It was also the model that Itanium used. You got the return address in a register and because register windows were saved to a separate stack, it thus saved the return address there too.
And if you want to be picking nits, it's actually EM64T or at the very worst, IA-32e. Then again, there are actually two versions of this ISA and Intel is currently calling its version "Intel 64" (and AMD used to call their pre-release version "x86-64", by the way), but definitely not "64-bit x86".
Edit: Oh hey, you're the same guy who made that silly argument 11 months ago [0]. Never mind me then.
Good job ;)
The point is that in a discussion that includes Alpha, "x64" obviously isn't clear. Doubling down with "but everyone does it" isn't a good look.
Now you want to pretend it's not an attempt at pedantry? I'm not sure why you WANT to be an asshole, but either disagree with me TECHNICALLY (that is, point out how "x64" is NOT and never was a reference to Alpha, and show evidence how "x64" somehow has always referred to amd64 outside of Windows-centric circles, and we can discuss that), or admit you're trying to be as wise-ass and don't be pedantic then get upset about me supposedly accusing you of being pedantic.
In other words, what do you really think you're bringing to the discussion? If you don't like when people point out incorrectness, then just say so.
There's a time and a place for making incorrect generalizations. Technical people shouldn't make incorrect generalizations when talking about technical things.
I've implemented a shadow stack for Virgil, so I am aware of how much it sucks. But that doesn't suck as bad as coming up with a stack walking protocol and then modifying every Wasm engine in existence to support that.
We usually base our intuition on the older and simpler branch predictors of 10-20 years ago, where the location of the branch and it's taken/not-taken history are tightly coupled. In those, a never-taken branch is either unknown (which counts as a non-taken prediction) or the branch is known and correctly predicted.
But modern branch predictors decoupled the branch location and branch history. They hash the sequence of the last few branches and use that hash to index into the branch history table, allowing the a single branch to have multiple different histories depending on the control flow leading up to it. That significantly improves branch prediction for hot code, the predictor can now track things like a branch that will always be taken during the first iteration of a loop, or a branch will always be taken if the function was called from one location but not another. They can even track the relationships between indirect branches in vtables.
But Hash collisions are expected, especially outside of hot code. So it should be quite common for never-taken branches in warmish code to be known, but predicted as taken due to a collision.
Trying to make software mechanisms like this that span multiple architectures feels like a relic of the past. These days you need to use processor-specific features like ARM pointer authentication.
Also worth noting: the mov being Turing complete paper.
> So amd64 isn't as good as arm64, riscv64, mips64, powerpc, or powerpc64.
On the embedded side, NXP (former Freescale former Motorola) makes them, as does 'Macom' (former AMCC via some hedge fund gymnastics...never heard of them either). Others like Xilinx probably still have licenses to produce. And you can still get the rad-hard RAD750 from BEA systems, for your outer space or post-apocalypse needs.
I'm not criticizing here, but actually seeking to understand.
Here is an old HP-UX machine running on PA-RISC:
# grep ntoh /usr/include/netinet/in.h
#ifndef ntohl
#define ntohl(x) (x)
#define ntohs(x) (x)
On x86_64, this is a byte-swap: $ grep bswap /usr/include/netinet/*.h
/usr/include/netinet/in.h:# define ntohl(x) __bswap_32 (x)
/usr/include/netinet/in.h:# define ntohs(x) __bswap_16 (x)
/usr/include/netinet/in.h:# define htonl(x) __bswap_32 (x)
/usr/include/netinet/in.h:# define htons(x) __bswap_16 (x)
SPARC is big-endian, and the memory is not in the same order as it is on x86. This can coax bugs out of software that are otherwise not seen on little-endian systems.Alpha is litte-endian, but it has its own exotic problems.
At least some models of the Alpha were bi-endian, selectable via a pin on the package, IIRC. I believe the Cray T3E ran the chips big-endian.
Apple Arm chips however are LE only, but all of Arm's Cortex-A/Neoverse and NVIDIA's cores support both LE and BE operation.
Or somebody may just think it’d be a fun challenge to do and so it supports it. That’s all.
Also, the environmental cost of manufacturing new hardware is often greater than even a decade of running less efficient, older hardware longer.
To add to that, Alphas and older Sun systems are actually proper server hardware. There are differences that you may never learn about in the x86 world. I've been running an AlphaServer DS25 for years now, and it's worlds better hardware quality than any x86 server product you can buy now.
The power supplies on that Alpha are 500W and even fairly low end Supermicro systems have 600W supplies.
The biggest differential is probably the efficiency of the power supply.
How would you characterize the differences?
Nothing is broken as I see it.
Reader mode would help with that.