64 B / inst * 2 (read and writeback) * 3 inst / cycle * 3 Gcycle/sec > 1000 gigabytes/sec, which is insane considering 50GB/s is the norm to RAM right now.
*edit: Haswell L1 cache seems to allow a peak bandwidth of 64 byte loads + 32 byte store per cycle, so like other comments point out, this would match the CPU capacity. On the other hand, there is also latency. But it is hard for me to tell how latency factors into this.
Could someone explain how they would decide whether copying packets is cheap or not?