CPUs do minimize latency by:
- Register renaming
- Out of order execution
- Branch prediction
- Speculative execution
They should not be over subscribed as they have to context switch by storing / loading registers and the cache coherence protocols scale badly with more threads.
GPUs on the other hand maximize throughput by:
- A lot more memory bandwidth
- Smaller and slower cores, but more of them
- Ultra threading (the massively over subscribed hyper threading the video mentions)
- Context switching between wavefronts (basically the equivalent of a CPU thread), just shifts the offset into the huge register file (no store and load)
The one area in which CPUs are getting closer to GPUs is SIMD / SIMT. CPUs used to be able to apply one instruction to a vector of elements without masking (SIMD). In ARM SVE and x86 AVX-512 they can now (like GPUs) mask out individual lanes (SIMT) for ALU operations and memory operations (gather load / scatter store).