> It needs specialized and careful designed software to tap the capability.
I've been suggesting engineers get more cores of lower speed to gain insight on what will be performant a few years down the road since I saw my first Xeon Phi.
It's been a while since clock speeds got higher (IBM has been pushing 5GHz in their highest end for the past couple years now and it doesn't seem likely they'll cross 6 anytime soon), but we get more cores every year. We now have 4-core entry-level machines and 2-core/4-thread ones are the bottom of the barrel, with a decent one being 8-core. Ampère just announced a 192-core server beast.
And then we have another thing: performance for most users has been "good enough" for the past couple decades. I haven't gotten a new computer just because it had a faster CPU since the early 2000's - they usually turn to dust well before they become too slow to use. My wife will need to upgrade her Macbook soon-ish for regulatory reasons (when Apple EOLs and stops patching macOS 12) and her laptop is still going strong. Considering that, there is little advantage in making all but the most demanding software more parallel.
This leaves the high-end, the stuff that needs a POWER10 or a Telum to run at acceptable speeds, and the cloud vendors, who'd kill to be able to serve 1% more VMs per kilowatt because 1% of their revenue is the GDP of a small country.