RAM-less Buffers
emsea.github.io
emsea.github.io
There have been cases where autovectorization in compilers have produced good SIMD versions of code (which is still considered a hard thing to do in some cases), but was still slower during certain benchmarks that did high thread counts (ergo, absurd amounts of CPU time wasted on context switching), but beat non-vectorized in less loaded situations.
>The reason why we won't be considering other operating systems is because the System V ABI doesn't preserve any of the XMM registers between calls and puts the burden on the caller to save them on the stack. If you think about it, this sort of defeats the purpose of using a register buffer if we're always going to be pushing our bytes to memory in user space.
as opposed to windows? regardless of whether it's the caller/callee's job to preserve registers, the result is the same.
Nope, because SSE operations are explictly defined as setting the upper bits of the ymm/zmm registers to zero.
> as opposed to windows? regardless of whether it's the caller/callee's job to preserve registers, the result is the same.
There is a huge difference, with caller saved registers, the caller must compulsively save all registers it's using to the stack before it calls a function. With callee saved registers, if the callee doesn't use the registers, then it doesn't need to save them and a bunch of extra push/pops are saved.
It is optimal to have a mixture of caller and callee saved registers, so the compiler pick what type it uses for each function/variable.
What immediately came to mind was perhaps you want to hide a secret key from entering memory. If this was done in kernel mode, you would be able to disable interrupts/task switching execute the "secret" stuff, and continue on your merry way...
You do understand that toys like this are meant to simulate the mind towards different ways of thinking, right? Great things are born from kennels of innovative thought.
HN has really gone downhill...
Not sure which "sourpuss comments" you mean. Perhaps a direct reply to one of those comments would be more helpful than this general one, which lumps together all fellow commenters.
I only saw comments from people to appreciate the effort, have lots of experience in similar areas, and share their knowledge about the pitfalls they see. All very helpful and polite, as far as I can see.
for (secret_key_t *p = 0; p < RAM_SIZE; p++) {
decode_with_key(p);
}
This does require significant physical access, but it works. I seem to remember reading ~1.5 years ago about a turnkey forensics kit (bottle of refrigerant included) for doing cold boot attacks? Regardless, more ways to protect keys is could be really useful.For Intel, look up SGX. For AMD, look up SEV. Each of these is way more secure than reliance on registers as secure scratch memory.
> So that's what we do: each untrusted thread has a trusted helper thread running in the same process. This certainly presents a fairly hostile environment for the trusted code to run in. For one, it can only trust its CPU registers - all memory must be assumed to be hostile. Since C code will spill to the stack when needed and may pass arguments on the stack, all the code for the trusted thread has to [be] carefully written in assembly.
> The trusted thread can receive requests to make system calls from the untrusted thread over a socket pair, validate the system call number and perform them on its behalf. We can stop the untrusted thread from breaking out by only using CPU registers and by refusing to let the untrusted code manipulate the VM in unsafe ways with mmap, mprotect etc.
(I don't know if that technique is still used)
So this reads to me as a programmer who isn't familiar with SIMD discovering one small part of why it exists and then writing an article about that as if it was a new idea outside the scope of normal SIMD usage.
It's nice that this may give exposure into some low level details to those unfamiliar, but it isn't an innovation.
Downvote me as much as you want for saying so but it is the literal reason why the registers and instructions exist.
I didn't mean any offense to OP, if that's the cause for bad response. I've been using these instructions for well over a decade and the description given is without exaggeration the precise reason they exist and is what I've always used them for.
The purpose of XMM buffers is for SIMD instructions, which implies that the buffers store multiple atoms and does the same instruction on each atom in the buffer at once.
This interesting bit of this little hack is how it breaks that expected data level parallelism and instead provides a (relatively) high-level interface for the new usage. It has nothing to do with shuffling data to and from those registers.
You're getting downvoted because your initial comment ignored the bit that used the registers in a way that wasn't intended and focused on the bits that were the same as the intended usage.
And then you're being condescending about the points you're trying to make. You might not have intended to be condescending, but you are being condescending nonetheless.
Let those who havent worked with SIMD reject an authentic attempt from a vet to elucidate err and be happy with that click. Whatever.
If you can drain my remaining few hundred karma I'll have an excuse to stop clicking another Bitcoin rehash article every day and will thank you.
You're not being downvoted for a technical issue.