Warehouse-scale computing has similar budgets and timeframes. You don't decide how to re-build this month's 10k machines based on this month's benchmarks. You made the decision as far back as the supply chain required you to do it, maybe a year or more.
I'm sure that AMD's sales team has been telling their big customers about this generation's performance improvements for a while. But with their history, decision makers are going to discount the story a bit until they can see it in production silicon.
So the next few month's movements in AWS and the like will all depend on the extent that their decision makers were convinced many months ago.
Like it or not, Google's stamp of approval carries a tremendous weight in this industry. And if that's not good enough for you, Microsoft and AWS are stepping up their deployments of EPYC as well.
As a result, a lot of smaller companies will now require a much lower standard of due diligence when approving an EPYC Rome deployment.
The signs are there, Lenovo has called them T480 and A480 last year, T490 and T495 this year, indicating these are very close.
Also, software is very often the largest cost for these systems. It's not hard to find yourself paying $100k a month for an Oracle license. A one-time expense of $50k for one piece of hardware vs another is barely a blip.
And in fact that hardware is often charged based on spec. So if you have 4x as many cores on Epyc, you will pay more in software costs on a monthly basis as well. That, or the software will simply refuse to use them until you buy an upgrade, meaning those extra cores are sitting there doing nothing.
It's counterintuitive to people whose experience is building a gaming desktop at home, but hardware expenses are not necessarily a big part of total cost of ownership for enterprise operators.
That shouldn't be a problem. They are both fundamentally the same architecture (amd64) and any CPU-specific features are already opportunistically handled by the vast majority of software because otherwise you wouldn't be able to run the same code on different versions of Intel's CPUs.
It works fine if you shut everything down and reboot the system, but that is often undesirable.
The whole point of the feature is that the VM can be migrated around different physical hardware without having to interrupt service. It just suddenly is running on a different host instance. But it has to be the same type of processor... or at least the same feature set. "Close" is not good enough, it needs to be a 1:1 match.
You can manually disable features until you have found the lowest common denominator between the feature sets of the different processors. But obviously the more types of processors you have in your cluster, the more problematic this is. In very few clusters will you find servers of mixed types, you buy 10,000 of the same server and operate them as a "unit". You don't just add in servers after the fact, sometimes you don't even replace failed servers.
And that hardware decision will have been made years ago, very often. The server market is hugely inertial, it's nothing like you putting together a build one evening and then going out and buying parts and putting it together.
I think you can do that on multisocket systems:
https://www.kernel.org/doc/html/v4.14/core-api/cpu_hotplug.h...
A very long time ago, I worked for a then very large company that sold servers. Plain standard 80486 based servers.
My job was to drive around and drop off these servers for evaluation at prospective customers, who would compare them against 80486 offerings from a different vendor.
Your argument about them all being fundamentally the same would be even stronger: it’s the same CPU.
And yet, customers did not take chances and would go through the eval motions. Because their business relied on it.
Now imagine that at a scale of thousands.
Claiming “they are fundamentally the same” is not wrong, but you don’t care about the fundamentals only. You care about the whole picture and you don’t take chances.
Very conservative corporate customers could wait a short time for good BIOS corrections and sufficient supply for all the parts (not only CPUs) they need before shopping for AMD servers, but they would be buying different hardware from the same established suppliers even if they went with Intel.
That sentence above is a contradiction in terms.
A conservative corporate customers spends many months to do evaluations. There is no such thing as “a short time.”
There's a lot more incentive to explore EPYC than there was a day ago.
And exactly what kind of risk are you taking about?
If switching is as easy as you claim it is, you can still do so next year or the year after.
Your comments in this discussion are very small scale, retail oriented.