Your example is a worst-case scenario, where only one very small write operation happens at a time for a very small size, yet there are a lot of things to write.
It would seem that Linux swap uses the block device layer to handle actual I/O for swap. So, I think that if it only needs to write or read one single page at a time, then sure, it will be slow. But if you need to swap out a sizeable quantity of memory (meaning multiple pages), it will need to write or read multiple pages "at a time". Linux also tends to group pages together. Since reading and writing is async [0], as soon as the io operation has been sent, the next page will start to be handled, even if the actual io to disk hasn't finished. So, in practice, it's very likely that there are several IO operations happening at the same time from the drive's point of view.
The system is also highly likely to be doing some unrelated I/O, which is why parallelism helps a lot.
This is where the SSD will shine, because it will absolutely be able to handle multiple small IOs in parallel [1]. And also where, presumably, queue-ordering in the case of spinning drives can help here, too, by optimizing the order in which the paging operations will occur.
I'm not familiar with Windows, but I'd expect things to work somewhat similarly, especially since the swap is backed by an actual file system (but I don't know if the swap files have any properties that make it "special" - though I wouldn't be surprised for that to be the case).
---
[0] https://www.kernel.org/doc/gorman/html/understand/understand... look for heading 11.7
[1] See RND4K / Q32T16 in your benchmark, where the factor is only 2.