"CPU designers see vastly more benefit to spending area on, e.g.,vectorized floating-point multipliers."
followed by "Intel is doubling the vector size in its newest CPUs---again!---this time from 256 bits to 512 bits."
Seems like they are being optimized to be better at vector math, and games just happen to highly use these pieces of HW.
More likely the other way around ...
It may in fact be that desktop/mobile CPUs are being optimized for their contemporary benchmark suites, thus targeting the ensuing benefits in marketing. The benchmarks themselves were, for a fair amount of their existence, focused on games-related performance, at least from what I recall.
In short, I think "optimizing for games" means nothing to a CPU designer at say Intel, and anyways qualitative differences (expanding the ISA, or integrating new features on the IC) are more expensive to develop and sometimes tricky to market. Instead, marketing quantitative differences is much easier (hence optimize for benchmarks) - though arguably no easier to develop. Witness Intel's first rocky attempt at targeting the benchmarks in the early 2000s: Netburst [0]. Of course, in the past 5 years, things have changed (the rise of smartphones, meteoric rise in GPU performance with expanding markets & new software, CPU "per core" performance stagnation), so Intel is in the process of re-positioning itself.
[0] https://en.wikipedia.org/wiki/NetBurst_(microarchitecture)
The reason for this is simple -- it's twice as slow. The vector width is half the size, and so you can do half as many operations at a time.
This would also explain for instance why many programming languages drop single precision floats altogether: they don't plan to vectorize in the first place.
OTOH rendering of geometry only needs about as much precision as the display offers, which often means 8 bits on older hardware.