There are a few relevant things about interrupts that matter here:
1. Interrupts, in general, don't have useful payloads. This is because the CPU and its interrupt controller keep track of all interrupts pending so that they can make sure to invoke the interrupt handler when it's next appropriate to do so. To avoid having a potentially unbounded queue of pending interrupts, the pending interrupt state is just a mask of bits indicating which vectors are pending. On x86, NMI is vector 2. It is pending or it isn't. This is like like a regular device IRQ, which can be pending or not. In the case of regular device IRQs, the OS can query the APIC to get more information. If the IRQ vector is shared, the OS can query all possible sources. In the case of performance counter NMIs, the kernel can read the performance counters. In the case of magic AWS NMIs, there is nothing to query. This is why I think it's a poor design.
2. The x86 NMI design is problematic: NMIs are blocked while the CPU thinks an NMI is running, but the kernel has insufficient control of the "NMIs are blocked" bit, and the CPUs heuristic for "is an NMI is running?" is inaccurate. The result is extremely nasty asm code to fix things up after the CPU messes them up. This is nasty but manageable.
3. x86_64 has some design errors that mean that the kernel must occasionally run with an invalid stack pointer. The kernel developers have no choice in the matter. Regular interrupts are sensibly masked when this happens, so interrupt delivery won't explode when the stack pointer is garbage. But NMIs are non-maskable by design, which necessitates a series of unpleasant workarounds that, in turn, cause their own problems.
The kernel can definitely have bugs that cause NMI handling to fail. I've found and fixed quite a few over the years. Fortunately, no amount of memory pressure, infinite looping, or otherwise getting stuck outside the NMI code is likely to prevent NMI delivery. What will kill AWS's clever idea is if the kernel holds locks needed to create a stackdump at the time that the NMI is delivered. The crashkernel mechanism tries to work around this, but I doubt it's perfect.