It won't simulate the block storage data loss, which is a key part of testing a database, filesystem or similar for robustness to those events.
Even killing a VM instantaneously on a host only simulates the loss of the guest OS's cache. The host OS still has its cache.
And even killing the host OS only simulates loss of the host OS's cache. The drives still have their caches powered, so might behave differently on power loss than host OS crash.
And killing a drive while the host keeps running is different again.
(And in terms of high-level software-observable data effects, abruptly killing power to a drive is not the same as slowly lowering the voltage or limiting the current to the drive, or seeing corrupt data on the I/O bus while the host system loses power (which is never instant), or... you can go quite deep with this.)
In the networked storage environment at AWS, who knows what events are possible to observe on EBS in the event of a system-level failure such as power loss in a data centre.
For example, EBS is a distributed system with replication, so if it has an implementation error that only affects certain untested, complex failure modes, it might lose recent writes in flight (as expected), but then some time later they might come back, or some of them might come back if it's sharded. Both are bad if software using it has already started state recovery after an outage. Distributed systems often have surprising incorrect recovery patterns, because it's a complex and subtle problem; that's why the Jepsen tests keep finding bugs in databases people have been using for years.
The observable events can differ depending on whether it's VM hosts, switches, block storage units, drives failing and in what ways.
For Linux guests we have abstracted away most differences we care about under fsync() or various combinations of O_DIRECT and virtual HDD cache disabling, or relying on known features of filesystems, but there's no easy way to be certain those abstractions actually provide the expected semantics on different kinds of system failure. Stress testing higher layers in the virtualization stack would go some way to verifying that they do in practice, not just in theory.