Linux to begin tightening up ability to write to CPU MSRs from user-space
phoronix.com
phoronix.com
For flexibility, users should be able to opt out of this (as with the rest of kernel lockdown). Sometimes I just want a holster full of foot-guns. Blanket R/W access to MSRs is that.
There are a lot of fairly powerful MSRs. Some of them might already be restricted for all I know. Here are some examples:
IA32_SYSENTER_EIP -> the instruction pointer after calling SYSENTER. Similar MSRs exist for the syscall instruction.
IA32_DS_AREA/PEBS MSRs -> DS_AREA is a pointer that points to a buffer of pointers that are written to by the perf subsystem. Manipulating DS_AREA to point to a user page should let you write ~arbitrary data to the kernel.
IA32_SPEC_CTRL -> disable some of the speculative execution mitigations.
There are less obviously powerful msrs as well, which probably let you get away with some sketchy-ness: APIC_BASE -> you could try to get it mapped into userspace then control the APIC directly. Or you could hide particular kernel writes by moving the apic over another page (there was an attack against SMM based on this a while back).
X2APIC MSRs -> no need to move the apic base when you could just write to the apic through msrs more directly.
IA32_EFER -> Unexpected switch from long mode to protected. This can't do anything good.
MTRRs -> Changing memory types underneath the kernel (potentially restricting access to kernel memory) can't do anything good to the kernel. You could also change the caching behavior, which sounds messy.
Edit: Formatting. This is probably still unreadable on mobile.That sort of functionality is typically handled by the platform controller by sending commands over LPC or I2C busses. CPU MSRs are generally restricted to controlling the behavior of the CPU itself.
> changing CPU frequency and voltage, etc.
These are often controlled by MSRs, but should be under the direct control of the kernel, not userspace software.
> Think performance counter MSRs, MSRs with sticky or locked bits, MSRs making major system changes like loading microcode, MTRRs, PAT configuration, TSC counter, security mitigations MSRs, you name it.
I think the only legitimate uses of MSRs in userspace today is changing power-management related settings and performance profiling, which should be managed by the kernel anyway.
I welcome this move. PaX/grsec had an option to block MSR accesses for years.
"Plundervolt is an attack on Intel SGX enclaves, which are trusted execution environments. We showed that we can get secrets out of SGX enclaves by lowering the CPU voltage while it is performing calculations. Plundervolt is a problem because SGX has an attacker model that says, “Even if you’re root, you should not be able to look inside my encrypted area.” "
https://research.redhat.com/wp-content/uploads/2019/12/RRQ-V...
Afaik, if you need the best security on Intel chips, disabling Hyper Threading is still a must :/
Wtf are they doing...
Step 2: Run 1 process per pi.
Step 3: Enjoy freedom from spectre-related problems!
Optional step 4: Move everything back to Intel/AMD when management realizes why they went for multiple processes per box in the first place. Money rules everything around us.
Or more likely a chip with a ton of ARM A5* cores on it due to availability.
EDIT: Wait, it looks like I could actually get a Drive AGX for $700 on Newegg[1] and they seem to have about 40% the single thread performance of a Ryzen[2]. So actually not out of the question if I was willing to give up more for security than I actually am.
[1]https://www.newegg.com/nvidia-jetson-agx-xavier-32gb-256-bit...
[2]https://www.phoronix.com/scan.php?page=article&item=nvidia-j...
> This behavior right now can be toggled via the msr.allow_writes= kernel module paramrter with on/off/default. Should legitimate use-cases come up where writes to MSRs from user-space are still desired, they may add the infrastructure to selectively grant/deny access to specific MSRs and ensure they are sanitized by the kernel.
Similar hardware restrictions already exists in the kernel, for example, by default the kernel restricts access to I/O memory since it's a dangerous, low-level zone, but if you really need to for some reasons (e.g. reflash your BIOS), you can boot with "iomem=relaxed" to turn it off. Treating MSR registers in the same way is very reasonable.
On recent kernels, MSRs writes are disabled by default even by the upstream kernel when the kernel is in lockdown mode (like when using Secure Boot).
https://kinghajj.github.io/blog/undervolting-with-secureboot...
But as a design point, this is being rolled back. There are just so many MSRs in the modern world and so many are trivially system-breaking that it's just not possible to sanely administer an interface like that.
If operating systems start to beef up security more aggressively then I won't have to worry so much about having old microcode. OpenBSD is ahead of Linux on this too afaicr, sadly though it doesn't like the Nvidia GPU in my computer.