AMD's 128-core Epycs could spell trouble for Ampere Computing
theregister.com
theregister.com
It is of course always possible that Ampere has completely shit the bed and the AmpereOne won’t be an improvement, but I’ll give them the benefit of the doubt I guess.
Presumably Altra-Next will just use Neoverse N2 (or N3).
I do expect Ampere to succeed in outperforming stock ARM core designs. Otherwise, there'd be no reason to have embarked on a very risky and expensive journey to have their own custom cores.
They're doing it because they THINK they can outdo Arm.
Arm isn't standing still, and they would have had to start years ago.
And in the Datacenter you've also got google and Amazon making their own arm based chips, not sure if MS does the same.
However, Apples M1 Chip had the biggest impact to consumers I think. The tensor chips aren't particularly better and we just don't have any interactions with hardware servers
AWS used Neoverse for Graviton.
Google Cloud is designing their own ARM chips: they're working on one design from Marvell (likely Neoverse), and their own custom design. These are in the design phase, and very unlikely to be produced this year.
Re: Google Tensor chips on phones: Those used ARM Cortex CPU and ARM Mali GPU, and a Google designed Tensor processing unit for machine learning. The google TPU replaces chips like Qualcomm Hexagon. The TPU does not use an ARM instruction set.
Apple doesn't do it with ARM design in this sense.
[0] https://www.macrumors.com/2023/02/22/apple-secures-tsmc-3nm-...
But I thought we were talking Server/Desktop chips, not phone chips (I have no clue about them, my MI11Ultra seems fast enough for everything, I'm not compiling on it).
And the M2 has massive caches (L1/L2/L3) https://en.wikipedia.org/wiki/Apple_M2
How much performance you get with caches are shown with the X3D models that - gaming performance - are faster than all than more expensive Intel/AMD chips.
M2 Max Ryzen 7900
Cores 12 12
L1 320kb 64kb
L2 32MB 12MB
Last 48MB 64MB (32MB/chiplet)The real hubris at these companies is thinking they can build a team that can create a better chip in one generation. Apple is on what the 10th public generation of their own cores, from a team they acqihired that was already producing cpus? How many generations/respins did it take before they replaced the ARM ip with their own designs, and then how many generations was it before they were faster? Not only that but Arm seems to have gotten serious a few years ago and the IPC is within striking distance of the best amd/intel products. They are no longer doing obviously stupid things, so it seems odd a company like ampere doesn't have another respun Altra with a N2/V2 sitting on the sidelines as a fallback when their own design fails.
The specs there sound seriously impressive to me, and like they might be getting ready to leave AMD and Intel behind in terms of IPC (for their highest performing chips).
ARM's clock speeds are much lower, so single core performance will probably be worse. But I'd guess server clock speeds may be similar.
Disclaimer: I don't know what I'm talking about
Note that dispatch doesn't mean vector width, which is harder for software to take advantage of. It means how many uops the pipeline can handle/clock cycle.
As for higher clock speeds, there's a whole lot more that matters such as pipelining instructions, cache sizes, and any number of other things. Clock speed by itself isn't particularly revealing.
I would also not expect the X4 to utilize that full width in any real workload, but it only needs a fraction of its full width to get more IPC than the X64 competition.
But I'd expect it is a reasonably balanced chip (why waste silicon?), and thus the wide pipeline is an indicator of the chip itself being wide with immense out of order capability.
Branch prediction rates tend to be extremely high. The X4's frontend also has 10 frontend pipeline stages. Which means ideally, it'd be correctly predicting all branches at least 10 cycles into the future, so that on clock cycle `N-10`, the frontend can get started on the correct instructions that'll be needed on clock cycle `N`. The difference between 1 basic block/cycle and >1 basic block/cycle is really small; it already needs a long history of successful predictions to get 1. But of course, each mispredict is extremely costly.
As for bringing up clock speeds and the M1, my point there was that the M1 has already left Intel and AMD behind in terms of IPC; it achieves similar performance despite much lower clock speeds. My original comment said that the ARM Cortex X4 looks like it is starting to leave Intel and AMD behind in terms of IPC, and I used the width as an indicator. You responded saying that the software has to actually allow for this. Yet the M1 example shows that existing software does in fact allow for significantly more out of order execution than Intel and AMD CPUs achieve.
So you could argue that, unlike the M1, the Cortex X4 will not be able to realize such an advantage. While plausible, if it does fail to do so, we at least won't be able to blame the software, because the M1 is able to do so despite the software. It'd have to be some deficiency of the X4 relative to the M1 -- such as cache sizes, memory bandwidth... Hopefully it does turn out to be a great chip! But that remains to be seen.
> Yet the M1 example shows that existing software does in fact allow for significantly more out of order execution than Intel and AMD CPUs achieve
My point is, shortening the pipeline needed to execute an instruction would also get higher performance. Perhaps they invented a better cache, perhaps larger, perhaps more associative, perhaps…? There's more than one way of increasing performance besides IPC and clock.
Design of the CPU can also influence all three of these. Things like better cache, better branch predictors and shorter pipelines, will all help IPC.
> My point is, shortening the pipeline needed to execute an instruction would also get higher performance.
This wouldn't increase throughput if 100% of branches are predicted correctly -- except for the the extra cycles before instructions start executing. It'd decrease branch mispredict penalties though, which is a big deal and would help in practice. The Cortex X4 did shave off a frontend pipeline stage relative to the Cortex X3 (11 -> 10). This is better than Intel Alder Lake. One contributor is probably that it is easier to decode ARM instructions in parallel without needing mutliple pipeline stages (e.g., one to find out where variable width instructions end before the instruction byte stream can be sent to decoders [not an issue if instructions are already in the uop cache]).
> There's more than one way of increasing performance besides IPC and clock.
I assume by "IPC" here you mean dispatch width? Things like a better cache for fewer misses, better prefetching, better branch prediction, larger reorder buffers so that it can speculate further ahead before stalling, all help IPC.
Zen1 CPUs (6 uops) were already wider than Intel Skylake (4 uops) and Ice/Tiger lake (5 uops), matching Alder Lake (6 uops). But they were obviously far behind in IPC (and Zen1 in particular also decoded AVX2 instructions into 2 uops).
Zen1 has SMT, which was part of the reason to go wide early on: the frontend wasn't good enough to feed that width with a single thread, but using two threads could mitigate that. Early on, Zen1 (and the Zen family) generally did better in multithreaded than single threaded benchmarks thanks to that approach.
The ARM Cortex X4 doesn't have SMT, so it's taking a different approach to performance.
A single number isn't going to be representative of performance across benchmarks or all the tasks you're interested in.
Unfortunately, I think it'll be more than a year before we can see the Cortex X4 (as it's aiming at TSMC N3E), but I'm definitely looking forward to deep dives into it's performance (and also that of Intel's Meteor Lake, Zen5, etc).
I too prefer colloquialisms and straightforward speaking, and I swear plenty, but if it makes someone uncomfortable, that could be considered bad manners.
Competition is good. We are lucky to have AMD.
Zen4c and Ampere chips are for data centers.
5820K was almost 9 years ago.
8700K was almost 6 years ago now.
If $375 was too much for a hexacore that was your choice for many many years, for a long time now the market simply preferred the quad cores.
Speaking from experience, I don't run that many parallel workloads (I don't know how parallel my compiles are), but when I do, given a specific TDP for a desktop processor, and the asymptotically diminishing returns due to lower clocks speeds, more memory/cache contention and more synchronization, the benefits of going for more than 4-6 cores are negligible, for parallel workloads.
10 years ago Intel was selling Xeon Phis, 60 little cores on a PCIe card.
Coincidentally the CEO/founder of Ampere worked at Intel at the time.
> In terms of power consumption, that is perhaps where the HPE ProLiant RL300 shines. Here is a screenshot of iLO 6 with the server idle. One can see a 136W average. That is fairly good for a 128-core server (~64-core EPYC or Xeon equivalent.)
https://www.servethehome.com/hpe-proliant-rl300-gen11-review...
Your statement is correct, but I'm not sure the comparison really means much?
Nothing datacenter AMD is efficient under idle condition.
>While a clear win for AMD's Bergamo, it doesn't take into consideration other elements like power consumption. AMD's part is rated for 360W and can be configured up to 400W, while Ampere's has a TDP of just 182W. So yes, it may be 2.5x times faster, but it potentially uses 2-2.2x more power.
For 2 or 4 sockets per RU, you could be dealing with hundreds of kW per rack, and the heat that entails.
For certain deployments that could make sense. For others, a lower power, but still high core count solution could be better.
I think OP's point is that you'd need over twice the number of the lower density cores to get the same performance, thus by going that route you'd end up needing more power to get the same computational resources.
To put it simply, with ARM you'd need a 4U to get almost the same compute as a 2U of AMD.
Public cloud providers will sell you that amount of compute for about $50-$100 per day.
The lower end of that is the Ampere processor and the high end would be the AMD 128-core processor.
In other words, they can make $49 more by using the AMD CPU per day.
The AMD part's power budget is about 2.7W per core, and Ampere's, 1.8W. None of these cores look like very high speed. Their performance will also likely depend heavily on efficiency of caches and access to RAM, not just cores' internals.
It depends on how the memory channels work, what kind of optimizations you use in your benchmarking tools, how power delivery is handled on the motherboard, and that's before you consider the way the CCDs or cross-core communication happens (which is unique on each of the major architectures (Xeon/EPYC/Ampere)).
Anandtech had a great article comparing previous-generation server chips and found that memory access, L2/L3 access, and other cross-core communication varied dramatically on different chips, and even when BIOS was configured certain ways!
If two chips are hard to tell apart in terms of performance, that just means that they are practically equivalent. This means other factors come in play, such as ISA and cost.
Your point stands though.
In a lot of cases when machines are within 20% of each other, its quite possible the real difference between the two results is how many hours a perf engineer has spent tuning the workload.
In a way this is why the simple microbenchmarks are as important as the system level ones. The simpler ones are easier to tune and give a speed of light for a given operation.
Macbooks are fantastic but there is no magic. Manufacturing is extremely important
If we're just running several processes there's not much difference from just running 8 computers each with 16 cores.
You need interprocess communication that works across machines.
You need to orchestrate the deployment of software.
Communication within a machine is much faster.
The big benefit of MSSQL and Oracle SQL is that your database server is almost completely detached from the operating system the likelihood of a system update changing how your database works is nill. With Postgres and other databases it’s not a given. Ironically on Windows you get to extract quite a bit of that back since Postgress on Windows comes with most of its own libraries, however that’s also part of the problem where you can easily have material differences between running Postgress on Windows and Linux.
I am guessing that you don’t want to validate PostgreSQL on all operating systems, but you could always stick to one using software containers.
Postgres is far more dependent on OS libraries which makes it far less predictable.
Of course you could skip patching the dependencies and ship unsafe software the oracle db way.
Postgres is dependent on a lot of OS level C libraries that can materially change how things work.
This means that there will have to be more testing with Postgres and there will be higher uncertainty between different deployments.
All of these can be mitigated and for many organizations the benefits of Postgres might outweigh these downsides but they do exist.
I don't know you or your workload, so I can't comment on its suitability for you. I would hazard a guess, though, that your last proper due diligence knowledge is seven years out of date, or more.
Edit: it’s kinda ironic you think my knowledge of sql server is outdated when you don’t understand the features supported in sql server to begin with.
You are correct that there is no formal SQL type in SQL Server for JSON. And I am sorry for implying otherwise. Type safety requires a constraint on the column intended to hold JSON.
There is no equivalent to Postgres's GIN indices which would allow indexing on an array in the JSON column. Such a requirement would need a normalized table holding the array's values in SQL Server. Whether this is a limitation or a lack of support for JSON, full stop, seems to me a matter open to debate.
I have seen many successful projects (and participated in quite a few of those) that utilize JSON in SQL Server databases. I will amend my former statement, though, because it obviously lacked nuance: SQL Server's JSON functionality has covered all use cases that I have had and personally seen, but my experience is obviously much lesser than some others', so you can take this experience with as much salt as you like (:
It's definitely a shortcoming of MS-SQL, which iirc is being addressed to an extent in the next release.
There are plenty of points where MS-SQL Server is very nice. JSON support is far from one of them. Even if the JSON is still a step above XML support in SQL Server.
MSSQL also has great tooling - another part of the ecosystem.
If I was starting today, pgsql would be my choice for licensing costs alone. But, if I already had a system built on MSSQL, it would be hard to make a business case to move (I have tried).
And Altra's are old PCIe gen 4 and DDR4, while AMDs are PCIe 5.0 and DDR5.
I'm thinking like Cloudflare or Twitter. There's not much compute going on at this level (all the compute would be at the SQL-server or Application-servers or other backend equipment handles).
What's needed for frongend / proxy code is AES for handling the HTTPS / TLS connection, and lots of threads to handle all the different connections.
At least for now, the fastest ARM CPUs are not immune to Spectre, so the recent ARM cores have introduced a set of similar workarounds to the recent Intel and AMD CPUs, in order to serialize the execution and flush internal state around context changes.
Searching for "Speculative Processor Vulnerability" on the Arm developer site will find many resources describing possible attacks and the mitigations that must be implemented for various Arm cores.
If the price is less and the cost to run is less, then the bottom line max performance may not be the leading factor in making a decision. It really depends on a few factors, and this article does a poor job even spelling that much out.