There are a few relevant things about interrupts that matter here:
1. Interrupts, in general, don't have useful payloads. This is because the CPU and its interrupt controller keep track of all interrupts pending so that they can make sure to invoke the interrupt handler when it's next appropriate to do so. To avoid having a potentially unbounded queue of pending interrupts, the pending interrupt state is just a mask of bits indicating which vectors are pending. On x86, NMI is vector 2. It is pending or it isn't. This is like like a regular device IRQ, which can be pending or not. In the case of regular device IRQs, the OS can query the APIC to get more information. If the IRQ vector is shared, the OS can query all possible sources. In the case of performance counter NMIs, the kernel can read the performance counters. In the case of magic AWS NMIs, there is nothing to query. This is why I think it's a poor design.
2. The x86 NMI design is problematic: NMIs are blocked while the CPU thinks an NMI is running, but the kernel has insufficient control of the "NMIs are blocked" bit, and the CPUs heuristic for "is an NMI is running?" is inaccurate. The result is extremely nasty asm code to fix things up after the CPU messes them up. This is nasty but manageable.
3. x86_64 has some design errors that mean that the kernel must occasionally run with an invalid stack pointer. The kernel developers have no choice in the matter. Regular interrupts are sensibly masked when this happens, so interrupt delivery won't explode when the stack pointer is garbage. But NMIs are non-maskable by design, which necessitates a series of unpleasant workarounds that, in turn, cause their own problems.
The kernel can definitely have bugs that cause NMI handling to fail. I've found and fixed quite a few over the years. Fortunately, no amount of memory pressure, infinite looping, or otherwise getting stuck outside the NMI code is likely to prevent NMI delivery. What will kill AWS's clever idea is if the kernel holds locks needed to create a stackdump at the time that the NMI is delivered. The crashkernel mechanism tries to work around this, but I doubt it's perfect.
Also, regarding point two, is the cumulative result of those factors you describe (other NMIs being blocked, lack of kernel control, the unreliable heuristic) that a diagnostic NMI may end up never being run?
Look up SYSCALL in the manual (AMD APM or Intel SDM). SYSCALL switches to kernel mode without changing RSP at all. This means that at least the first few instructions of the kernel’s SYSCALL entry point have a completely bogus RSP. An NMI, MCE, or #DB hitting there is fun. For the latter, see CVE-2018-8897. You can also read the actual NMI entry asm in Linux’s arch/x86/entry/entry_64.S for how the kernel handles an NMI while trying to return from an NMI :)
To some extent, the x86_64 architecture is a pile of kludges all on top of each other. Somehow it all mostly works.
> is the cumulative result of those factors you describe (other NMIs being blocked, lack of kernel control, the unreliable heuristic) that a diagnostic NMI may end up never being run?
No, that’s unrelated. A diagnostic NMI causes the kernel’s NMI vector to be invoked, and that’s all. Suppose that this happens concurrently with a perf NMI. There could be no indication whatsoever that the diagnostic NMI happened: the two NMIs can get coalesced, and, as far as the kernel can tell, only the perf NMI happened.
Once all the weird architectural junk is out of the way, the NMI handler boils down to:
for each possible NMI cause
did it happen? if so, handle it.
If no cause was found, complain.
Amazon’s thing is trying to hit the “complain” part. What they should do is give some readable indication that it happened so it can be added to the list of possible causes.I don't know if that's enough to make an NMI "unreliable", but the kernel being able to do nothing for NMIs to fail might be a bit strong.
A triple fault will reboot or do whatever else the hypervisor feels like doing.
Though arguably, a hypervisor can dump out some useful state on a triple fault, as they actually tend to do. But that's not an in-band kernel panic, and in VirtualBox at least you can get a dump of that state without sending an NMI or otherwise disturbing the VM.