Here's a different thought experiment: how much of that data movement is actually necessary? If data movement was taking up so much time, why not investigate why and if that's even necessary, instead of thinking it's the hardware that needs to be faster? As the old saying goes, "the fastest way to do something is to not do it at all." One style of code that seems to be particularly exacerbated by the widespread use of HLLs is making lots of unnecessary temporary copies of data. The biggest inefficiency isn't the hardware (unless your goal is to sell more hardware, but that's a different rant...) --- it's the insanely complex and bloated layers of software that gets layered on top of what is otherwise perfectly adequate hardware. From that perspective, faster hardware only leads to even more wasteful and inefficient software. Making memory copies faster will only cause software to make more unnecessary copies.
But after reading the whole thing, I can say that this is certainly a very strange article. It's written by someone who is presumably quite knowledgeable about hardware, yet seems completely oblivious to the fact that x86 has had a dedicated instruction for this since the 8086, and 16 bytes per cycle was already achievable with the P6 (1996) "fast strings" enhancement:
https://stackoverflow.com/questions/33902068/what-setup-does...