Regarding HT, 2 threads is really a sweet spot for a 4-wide CPU. More than that and the competition for cache resources, execution units and register file become significant.
POWER8 is special because a factor of x2 is because each power 'core; is pretty much two distinct smaller cores that can gang together to speed up one thread (it also helps on per core software licensing), while the other x2 factor is for very specialized loads (this is also true, or used to be, for SPARC).
IIRC XeonPhi which is also a specialized cpu has 4xHT.
So, if you don't need the high peak clock speeds, 32 half-speed cores would be preferable for datacenters to save money on electricity and cooling design.
Intel does similar, the E5-2650 v4 is 2.2 GHz to 2.9 GHz.
As an example, the 22-core Intel Xeon E5-2696v4 has a base clock of 2.2GHz. With one or two cores active, it can turbo up to 3.7GHz. With three cores active, the maximum is 3.5GHz, and it decreases by 100MHz per active core until ten cores are active. With 10 or more active cores, the limit is 2.8GHz, provided that the chip is still within its power and thermal limits.
As such this would be ideal for things like VDI and web / cloud hosting where the quantity of VMs is very high, but the load from each is typically not.
Which is why every virtualization platform out there lets you oversubscribe CPUs. That's a solved problem, what's the benefit of having 100 VMs run on 32 slow cores vs 16 fast ones?
As mentioned above, context switching is expensive and extra L1 cache is valuable. Time-sharing can also have a huge effect on latency (because requests must wait until their server is scheduled), even when the throughput is still good.
Even if time-sharing performs well most of the time, when it goes wrong, the performance problems can be opaque and hard to debug. In general solution that "really" does something will save engineer-days as compared to a thing that does it at the same price/performance trade-off, but virtually.