Ampere’s Product List: 80 Cores, up to 3.3 GHz at 250 W; 128 Core in Q4
anandtech.com
anandtech.com
I feel like in 20 years from now we’re gonna be using intel as a cautionary tale of hubris and mismanagement. Or whatever it is that caused them to fail so spectacularly.
Aka yes it’s a cautionary tail and time to run from that ship.
They bent over backwards for cloud providers and offered them special deals that helped finance the cloud providers transitioning to own silicon. They fused off features to create false product "differentiation" like the IBM of old and failed to deliver technology after technology in working form (SGX, TSX, 10nm, ...) They held the performance of the PC platform back by trying to capture all of the BoM for a PC. (e.g. tried to kill off NVIDIA and ATI with integrated 'graphics')
Customers are angry now, that's their problem. Intel is like that Fatboy Slim album, "We're #1, why try harder?" They still think they are the #1 chipmaker in the world but now it is more like #2 or #3.
I'm pretty sure they have their 14nm business, and are working really hard to get a 7nm manufacturing business? A quick search gives me news articles about Intel hoping to have 7nm working by 2021.
Also, oss instruction sets like Risc-V are going to be interesting.
What went wrong at Intel is that they forgot to take the appropriate steps 10 years ago to avoid running out of options right now.
Ten years ago it was already obvious that mobile CPUs were a thing and Intel's attempts to penetrate that market failed around that time. From that moment they were living on borrowed time.
I’m not saying that Intel’s not in trouble, just that the conclusions here are far from obvious. I have some skepticism for people who say that they saw this coming. AMD laid off a lot of top engineers before its recent resurgence. Intel’s failure to ship its 7nm node in volume was a surprise to a lot of people.
Everyone knew that the new process nodes were more difficult, but outside a few experts, hardly anyone was in a position to predict when the move to smaller nodes would slow down.
Not that long ago, people were praising Intel for their superior SSD controllers, or talking about how they would be making 5G modems.
ARM has a much weaker memory model with significant performance implications for multithreading as well.
If simply having a better, and cheaper product will equal to immediate success then Mid-High End Fashion, as well as gazillion of other products in many other industry would not have existed as there are always competitors that offers something better at cheaper price.
Marketing / Discovery and Channels / Distributions. And that also excluding the software advantage Intel has.
AWS only just had their GA on Zen 2, nearly a year since they first made the announcement. Compared to Intel. I dont have any insider information. But I guess AMD has a lot to learn with regards to dealing with HyperScalers.
And you may have notice, Intel has way more leaks than usual in the past 12-18 months. That is part of the PR play to keep people from buying AMD while they try to Catch up.
Intel as of today is still operating at 100% capacity with back-log orders to fulfil, and a new record revenue in the last quarter. So yes, technically Intel is inferior, but until all of those disadvantage materialise into financial numbers it is far too early to call the death of Intel.
I dont hold any Intel Stocks but speaking as an AMD shareholder.
I never said it did. I was just questioning if the facts were indeed those.
To your original question, the simple answer is yes.
Intel had the twin defensive 'moats' of x86 and the leading process technology.
But now Intel has stalled relative to TSMC on the process lead (and possibly lost it) and the last few days have shown that the x86 moat is crumbling. The world will not move to TSMC manufactured ARM overnight but a significant shift could happen quite quickly I think. Intel will / has defended with lower prices but that will ultimately mean a big shift in business model.
but ....
Intel is the last firm with leading edge technology manufacturing in the US. If it starts to falter I can see a concerted effort to maintain that from the US government.
We can make good cars, though the majority of the good cars made in the US are made under the management of Japanese car companies. Tesla has some great technology and product design but on the manufacturing front they are behind the majors.
Sucks too, because otherwise they still are the best bang for your buck when it comes to performance.
Between Windows Update and crappy javascript, the theory that modern desktops are usually idle is also increasingly untrue, and under load the power consumption for the Intel parts is worse.
On the other hand, few applications scale efficiently to more than just four cores. Yes, of course, AMD delivers more Cinebenchpoints-per-Dollar and usually more Cinebenchpoints overall, but that's not necessarily an interesting metric.
Personally I find that if I'm waiting on something to complete that the application in question tends to use only a tiny number of cores for the task at hand. Usually one.
Another significant weakness of AMD's current platform is idle power consumption.
These factors leave me with a much more nuanced impression than "Intel is ded" or "HOW IS INTEL GOING TO CATCH UP TO THIS????"; CPU reviews these days are just pure clickbait.
Clock speed advantage -- Most Zen 2 CPUs don't overclock to 4.5 GHz on any core, let alone all-core. The boost numbers are reached with current firmware, but only for tiniest fractions of a second and never under any real load. Sustained single-core boost frequencies are 200-400 MHz lower than the specified boost frequency. On the other hand, Intel CPUs consistently reach their boost frequencies under load, and most CPUs can do their single-core boost as an all-core overclock under load (with much greater power consumption of course).
In practice this means that for equivalently priced parts (e.g. 3900X vs 10900K) the AMD part will have about a GHz lower clock for lightly threaded workloads, which are most workloads. With Intel settings, the Intel and AMD parts have about the same sustained clocks (3.8-4 GHz) under all-core load, but with the defaults of many motherboards the Intel part will run at 4.8-5 GHz, depending on the cooling.
They're only "equivalently priced" when you're talking about MSRP. Right now the 3900X sells for $413 and is in stock, whereas the i9-10900k sells for $530 and is out of stock.
Meanwhile, pointing at memory latency as the flaw in Ryzen has been a popular misdirection for a while now. People warned me about it being a performance pitfall since before I bought my first Ryzen processor. In practice it doesn’t show up in even the most complexity intensive workloads as a serious issue. For example, Zen 2 performs very well on hardware emulation. This is possibly because where it takes a hit in memory latency it makes up in caching and prefetching, but honestly I don’t know and I am not sure how to measure. In any case it’s certainly favorably comparable to Intel’s best chipsets in single core workloads even if not on top. Factor in price and multicore workloads and you now have the exact reasons why people like me have been singing the praises... Intel’s single core lead may exist in some form but it is not what it once was, it is not an unconditional lead where an Intel core beats an AMD core. Not even close.
None of this means Intel’s dead of course, but IMO thats mostly because they have a lot more going on than just being the best CPU. They’ve got their dedicated GPU coming out, and plenty of ancillary technology as well. It does seem like for a company like Intel having to take a backseat in CPUs for a while will be painful; unlike AMD, this is a new position for Intel and maybe not one they will handle well.
Of course, this is all a factor of Amazon's supply of instances and their chosen on-demand pricing level, but the trends are certainly interesting, and show steady demand for fast Xeon's and increasing demand for ARM's. I have run some compute heavy workloads on the best AMD's I could find on AWS and the speed difference per core for my particular workload was nearly 50%, which got worse as it scaled up to bigger instances because my workload uses a lot of L3 cache. I hear about EPYC's with 256MB of L3 cache but I can't seem to find those on AWS -- only ones with 8MB of cache.
C6g instances only launched on June 11. I'm not sure what information can be gleaned from the spot prices regarding Arm demand at this time.
The C5a instances powered by AMD Rome processors have 192 MiB of L3 cache per socket total (16 MiB L3 slice per compute complex, 12 CCX per socket). You can observe this from the cpuid(1) output:
L3 cache information (0x80000006/edx):
line size (bytes) = 0x40 (64)
lines per tag = 0x1 (1)
associativity = 0x9 (9)
size (in 512KB units) = 0x180 (384)
384 * 512 KiB = 192 MiB(you can download cpuid from http://www.etallen.com/cpuid.html)
For L3 cache, we try to optimize for the best overall performance for the majority of the time. Smaller instance sizes share L3 cache with other instances. I wouldn't call it a "free for all" given some changes in how the cache hierarchy has been shifting over time (e.g., Skylake-SP L2 cache per core was increased, and the L3 cache is now 'non-inclusive')
How is it a misdirection? The data is accurate and memory latency scaling is a well-known issue for simulations like e.g. games (which is a huge market for high end desktop CPUs and also the market 90 % of reviews address), where you can't really explain the performance differences just by higher clocks. It's considered the main reason why much older Intel CPUs can still outperform Ryzen CPUs in games.
On the other hand, if you take something like Cinebench you can literally turn XMP off (thus using JEDEC timings and bus speed) and still get almost the same score (within, say, 2 %). That's because Cinebench is benchmarking pretty much only ALU throughput. That's obviously an important factor for performance, but just as obviously not the only one.
Youtube is possibly the single largest root cause for users upgrading laptops over the past 10 years. They made a silent transition to 60 FPS videos last year which cut hundreds of millions of users from watching HD.
https://www.youtube.com/watch?v=ef1wAfrMg5I is ~10% of 1 cpu on my desktop using chrome.
OTOH, I know what your talking about, my linux machine hates youtube, but that's because even with the chromium freeworld fork with some codec acceleration its still burning CPU like crazy.
So, a big part of this isn't a hardware problem so much as a software one combined with the constant fights over who's codec is the one true choice. AKA its a youtube and !windows/android+chrome problem.
Those tasks are IO-bound, not CPU-bound.
Your concerns have no basis whatsoever.
> Youtube is possibly the single largest root cause for users upgrading laptops over the past 10 years.
No one in the whole world feels the need to upgrade to a high-end workstation because of YouTube videos.
2. Like I said, in my experience Ryzen also competes just fine in single core. It just also decimates in multicore. I’d rather have some tasks run significantly faster than have some run very slightly faster. But that is disregarding the fact that not all tasks are the same and it does in fact win some categories. These CPU architectures are more divergent than usual for lately.
3. Things you think aren’t parallel are. Video games using modern graphics APIs are in fact able to exploit multicore CPUs. Browsers absolutely exploit multicore CPUs. Your system in general will exploit multicore CPUs so during general usage when you are doing more things and have more software running, single core performance will be hurt less. And so on.
Compiling code isn't embarrassingly parallel unless you're building some project with lots of files from scratch. Video rendering and compression also don't benefit as well as you may think:
https://www.phoronix.com/scan.php?page=article&item=3900x-39...
Meanwhile, single-threaded performance affects pretty much 100% of what you do.
In the end, I don't think there's a big difference either way.
This is already only marginally true, the difference is only about 5% depending on the application, and in some applications AMD comes out ahead anyway. Expect the remaining difference to disappear when Zen 3 releases in a few months.
>Another significant weakness of AMD's current platform is idle power consumption.
AMD seems to have caught up here almost entirely. They've done a lot of work to improve idle power consumption lately and the node advantage probably helps, too.
That's a big PR hit, but in terms of sales - Apple isn't really significant customer of x86. But, if it'd trigger Microsoft to double down on Windows on ARM, that could then become a threat. But MS is playing with ARMs for years, and nothing significant came out of it yet.
https://youtrack.jetbrains.com/issue/JBR-2310
https://bugs.openjdk.java.net/browse/JDK-8248315
No, no, no it is not an OS bug, a hypervisor bug, or a JVM bug. Read the whole thing if you have an hour to kill. It's confirmed to be a CPU bug, and Intel knows about it.
I am really looking forward to the postmortem. The behavior reminds me of the old F00F bug.
[1]: https://www.anandtech.com/show/15578/cloud-clash-amazon-grav...
[1] https://www.phoronix.com/scan.php?page=article&item=epyc-vs-...
Still would be interesting to know what differences caused the gap in results, but their setups were pretty different.
My personal opinion is that the Phoronix way places quantity over quality. Measuring performance is an important part of shining a light on where we can improve the product, but I get little practical information from those numbers, even when they are reported as non-synthetic. There are HPC workloads that are showing significant cost advantages when run on C6g, like computational fluid dynamics simulations. See [1].
I expect the scalability of HPC clustering to improve on C6g in the future, like C5n improved cluster scalability compared to C5 with the introduction of the Elastic Fabric Adapter. The Phoronix and Openbenchmarking.org approach doesn't give much insight into workloads like this.
My advice for an audience like folks on HN is is to test it for yourself. For me, being able to run my own experiments is how I come to understand infrastructure better. And the cloud lowers the barrier of running those experiments significantly by being available on-demand, just an API call away. I'd love to hear what you think, either in a thread here or you can contact me via addresses in my user profile.
[1] https://aws.amazon.com/blogs/compute/c6g-openfoam-better-pri...
Will we need to recompile? Will it be almost-100%-binary-equivalent-with-some-hidden-bugs ?
MIPS: https://gcc.gnu.org/onlinedocs/gcc/MIPS-Options.html AArch64: https://gcc.gnu.org/onlinedocs/gcc/AArch64-Options.html Arm: https://gcc.gnu.org/onlinedocs/gcc/ARM-Options.html
Also, without pricing, all these efficiency claims are totally meaningless.
Will this be the same? It seems possible. Does it really get more work done per watt than x86?
And why does the article say "These Altra CPUs have no turbo mechanism" right below a graphic saying "3.0 Ghz Turbo"?
ARM has well-thought out NUMA support, probably a system this size or larger should be divided into logical partitions anyway. (e.g. out of 128 cores maybe you pick 4 to be management processors to begin with).
These chips obviously have variable clock speed, but apparently nothing like the complicated boost mechanisms on recent x86 processors. My guess is that Turbo speed here is simply full speed, and doesn't depend significantly on how many cores are active, and doesn't let the chip exceed its nominal TDP for short (or not so short) bursts the way x86 processors do.
Either that, or 3 GHz always exceeds the envelope and the chip is throttling clocks down all the time to keep itself inside the allowed power envelope.
That's more of an artifact of how TDP is defined than anything else. I doubt that this could peg even a single core at 3.0 GHz given a reasonable cooling setup, let alone run all cores @ 3.0 GHz.
You can get ~3.4GHz average sustained all-core speed on a 3990x (64-core, nominal 280W TDP). This is with an off-the-shelf AIO cooler.[0]
Note: top-end air coolers are often competitive with AIOs, and can be had for $80-$100.[1]
If you're buying a several thousand dollar CPU, dropping even $500 (much higher than you'd need for closed loop liquid or high-end air cooling) on cooling doesn't seem unreasonable.
[0] https://www.anandtech.com/show/15483/amd-threadripper-3990x-...
250W is not an absurd figure to shove in a server CPU. If you're buying an 80-core CPU, you're not going to be skimping on the cooling solution.
Especially given the target market for Ampere is cloud providers, you can expect these to be racked in enclosures that provide sufficient cooling for their operational needs.
Consider code which linearly goes through a list of points in 2D space and does some calculation on the coordinates.
In Rust, the list is a Vec<(f64, f64)>. The Vec is a small object containing a pointer to a large block of data which contains all the points packed tightly together. Once the program has dereferenced the pointer and loaded the first point, all the others are immediately after it in memory, in order, containing nothing but the coordinates, and so the processor's cacheing and prefetching will make them available very quickly.
In Java, the list is an ArrayList<Point2D.Double>. The ArrayList is a small object containing a pointer to an array of pointers to more small objects, one for each point. Each of the small objects has a two-word object header on it. The pointer plus header means that for every two words of coordinate, there are three words of overhead, so the cache is used much less effectively. The small objects aren't necessarily anywhere near one another in memory, or in order, so prefetching won't help.
There are a couple of ways the Java situation can be improved.
Firstly, today, you can replace the naive ArrayList<Point2D.Double> with a more compact structure which keeps all the coordinates in a single big array. This gives you the same efficiency as Rust, but requires programming effort (unless you can find an existing library which does it!), and may give you an API that is less efficient (if it copies coordinates to objects on retrieval) or convenient (if it gives you some cursor/flyweight API).
Secondly, in the future, the JVM could get smarter. In principle, it could do the above rewriting as an optimisation, although i wouldn't want to rely on that. A good garbage collector could bring the small objects together in memory, to improve locality a bit.
Thirdly, in the near-ish future, Java will get value types [0] which behave a lot more like Rust's types. That would give you equally good density and locality without having to jump through hoops.
No, but the higher the memory bandwidth, the sooner those processes can get back to their inefficiency.
Typically only multiple cores running optimized software that will run through memory making heavy use of the prefetcher will exceed memory bandwidth.
PS: don't worry, my upgrade is on it's way :p
With non blocking IO and async processing that can be good enough but to fully utilize dozens/hundreds of CPU cores from a single process, you basically want something that can do both threading and async. But assuming each core performs at a reasonable percentage of e.g. a Xeon core (lets say 40%) and doesn't slow down when all cores are fully loaded, you would expect a CPU with 80 cores to more than keep up with a 16 or even 32 core Xeon. Of course the picture gets murkier if you throw in specialized instructions for vector processing, GPUs, etc.
The best (efficient) way to utilize that many cores IS to have 1 pinned process/thread per-core: https://www.scylladb.com/ https://github.com/scylladb/seastar/
That would be the same in Python too. A problem is that you can't share the kernel pages for the code, and you need to have a shared-cache. Probably 0-mem-copy with no deserialization example: lmdb + Flat Buffers.
Nicer is to have 1/2x cores, but each core being 2x faster ;)
I know it's a bit of a cliche but it feels to me like Apple might have got its timing right on this one.
- When Apple founded ARM in 1990 with Acorn and VLSI, they didn't know that silly cacheless chip would become a world-beating juggernaut. But hey, as founders, they now have a license to mold the microarchitecture however they like.
- When Apple bought NeXT in 1997, they didn't know the sun was slowly setting on PowerPC. But they secretly nursed along the (already built) NeXTSTEP x86 port for years, until the time came to dust it off and start shipping product with Intel Inside.
- When Apple forked KHTML in 2001 and started building WebCore/WebKit, they didn't know that MS was about to leave Internet Explorer to wither on the vine, nor release the final Mac IE only 2 years later. But they quietly invested in building such a konquering (sorry!) product that (with Google/MS help) we're now at risk of an entirely different browser monoculture.
I now unplug my computer when I'm done, which is kind of annoying.
I'd say heat is not a concern, but noise can be. It takes some time to figure out good fan curves to balance cpu temp vs noise. There may be some companies who do pre-built and well configured machines, but I haven't researched that at all.
The profit margins on the Mac Pro are just incredible. (Yes, I'm sure that equivalent professional workstation brands also have huge profit margins... no, that doesn't make me want to pay those lofty prices more.)
The only real value the Mac Pro provides is that it's the most powerful computer you're allowed to run macOS on legitimately. If you can do your work from Windows (with WSL) or Linux, you can save upwards of tens of thousands of dollars by building your own workstation, and that workstation can be significantly more powerful than any current Mac Pro at the same time.
For video professionals who rely on FCPX or similar macOS-only software, they don't really have a choice, and they get the opportunity to essentially pay $10k to $20k just for a license of macOS, which is fun.
It is pricey but if it is something you want to buy for 5 years, it is about $100/month cost. Some people might want to buy it.
Currently, the Intel Xeon is used in both high-end workstations and servers. If one x86 design can be suitable for both of those, presumably one ARM design could do the same.
If they could sell server CPUs at a profit, then Apple could get more return on its design investment by getting into two markets. And they'd get more volume. Though apparently they'd be facing competition from Ampere and Amazon's Graviton.
* There's one "Ampere Computing" [1], but I guess I'm not "in the know" since it is the first time I heard about it :-/
* There's one Ampere [2], "codename for a graphics processing unit (GPU) microarchitecture developed by Nvidia".
Are both things related? Is "Nvidia's Ampere" developed by "Ampere" the company?
Also, I think Ampere is kind of a bad name for a processor line... just makes me think it of high current, power-hungry, low efficiency, etc. :-)
So the question is whether they can land Google, Microsoft, and/or Alibaba as customers for an alternative to AWS M6g instances.
These high density configurations were key when rack space was at a premium, but these days, power is the limitation, so this is interesting to provide more low power cores, i'm just not sure who is going to get the most benefit from them though...
Where this might get interesting, depending on how the pricing stacks up, is that if you're in the cloud function business, this will increase the number of function instances you can afford to keep warmed up and ready to fire. In those situations you're not bottlenecked on the total bandwidth for the function itself (usually), your constraint is getting from zero to having the executable in VM it's going to run in, and from there getting it into the core past whatever it's contending with. If there's nothing to contend with and it's just waiting for a (probably fairly small) trigger signal, execution time from the point of view of whatever's downstream could easily be dominated by network transit times.
I want a workstation with one of these.
Aesthetics is a big thing - rackmount servers are ugly and, unless there are panels covering it, they are horrendous deskside workstations.
Another one is noise. These boxes are designed for environments where sounding like a vacuum cleaner is not an issue. Because of that, they sound like vacuum cleaners, with tiny whiny fans running at high speeds instead of more sedated larger fans and bigger heat exchangers.
HP sold, for some time, the ZX6000 workstation that was mostly a rack server with a nice-looking mount. If someone decided to sell that mount, it'd solve reason #1, at least.
E.g. i suppose computer animation rather takes a GPU than 32-64 universal cores, and compilers are still not so massively-parallel.
Are they significantly cheaper per GHz*core? If so, how hard is it to make use of that power, will a simple recompile work?
In the context of AWS.
They are cheaper per some / specific workload* on AWS.
Especially when ARM Graviton 2's vCPU on AWS are actual CPU core while Intel / AMD instances are CPU thread.
And in general AWS offers the G2 instances with the same vCPU core at 20% discount compared to AMD / Intel instances.
Thank you for that information. Is there a reference that documents this somewhere?
Each vCPU is a thread of either an Intel Xeon core or an AMD EPYC core, except for M6g instances, A1 instances, T2 instances, and m3.medium.
Each vCPU on M6g instances is a core of the AWS Graviton2 processor.
Each vCPU on A1 instances is a core of an AWS Graviton Processor.
> deliver significant cost savings over other general-purpose instances for scale-out applications such as web servers, containerized microservices, data/log processing, and other workloads that can run on smaller cores and fit within the available memory footprint.
> provide up to 40% better price performance over comparable current generation x86-based instances1 for a wide variety of workloads,
From what I read, it's not terribly hard to tell your compiler to compile for a particular instruction set, you just need to do it. Cost savings and better performance are great incentives, as well as Apple moving their Mac platform to it will drive more market share for developers to take the time to recompile.
Edit: Forgot to add the source of those quotes: https://aws.amazon.com/ec2/graviton/
Once it is fixed you are fine. Most of the big programs you might use are already fixed. Some languages give you gaurentees that make it just work.
So e.g., on x86 if you store to A then store to B, then if another core sees the store to B it is guaranteed to see the store to A as well. This guarantee does not exist on ARM.
But on x86, many of these things don't matter, if I understand correctly.
Two threads reading and writing to the same memory area do not necessarily give problems. In fact, many software is built to exploit several facts about how memory accesses work with respect each other.
ARM processors give very few guarantees, so code has to workaround that.
It's good to be skeptical. I always encourage folks do experiments using their own trusted methodology. I believe that the methodology that engineering used to support this overall benefit claim (40% price/performance improvement) is sound. It is not the "benchmarketing" that I personally find troubling in industry.
But We can't measure power consumption/heat, possibly noisy neighbor exists while benchmarking, and can't know real price for cloud instance. I don't blame but it's difficult to comparing a hardware.
Of course different cpus can do different amounts of work per amount of electricity used, but arm generally works out better on a watt per unit of work basis.
>they are operating on the same market/segment
They are. Called computation.
>willful intent to defraud customers
Is it a requirement? I doubt it.
btw, I only clicked the link because I thought of the Nvidia's product, so they are definitely getting eyeball traffic due to the name.
UPD: I recognize that I'm unlearned in trademark law, so I'm not insisting on anything.
Ampere had products on sale in 2019.
If there is a case I can't see Nvidia winning it.
How long were Ampere planning to use the name before 2017 and then does Nvidia using it on a slide in a presentation force them to change it? Still think Nvidia would lose on this one.
I mean it's probably not the fault of either, and a huge coincidence we're getting a flurry of news articles about both in summer of 2020, but come'on (can we have some kind of edits in the titles of HN posts to make the distinction clear?).
As we approach the critical velocity (supply / demand) for parallel architectures, the prospects of bootstrapping a CPU manufacturing company will become extremely feasible. IMO currently it's mostly the specialized knowledge needed to design CPUs that keeps this mostly out of reach today.
I'm no expert, just have an interest in the space, so any dissenting opinions / facts welcome.
The x86_64 ISA is absolutely insane. The only way to implement it in hardware efficiently is to "compile" the super complicated instructions into micro-ops which can actually be decoded and executed on the CPU.
Said another way, Intel has to implement a compiler in hardware which compiles the machine code before it gets executed. The extra complexity means more power and less performance.
You can read more about how microcode and micro ops work here: https://en.m.wikipedia.org/wiki/Intel_Microcode
So this would only be a significant bottleneck for hot loops that are large enough that they don't fit in the uop cache.
It's definitely a real issue but it seems wrong to pin all or even most of Intel's stagnation on that.
> Said another way, Intel has to implement a compiler in hardware which compiles the machine code before it gets executed. The extra complexity means more power and less performance.
This is a sadly prevalent misconception.
There is no "compiler in hardware". There are two kinds of instructions; the simpler ones (which are the most common) are expanded directly into a fixed sequence of micro-ops, while the more complicated ones act like a subroutine call to the microcode. The closest software analogue would be a macro assembler, not a compiler.
AFAIK, the extra complexity for efficiently decoding x86 instructions comes mostly from their variable length without an explicit length indication, and from the variable number of prefixes which can change the interpretation of the following byte, which makes decoding an instruction a serial task. IIRC, both Intel and AMD have a couple of tricks to reduce the impact this has on both power and performance: caching the already decoded micro-ops, and storing the instruction boundaries in the L1 instruction cache.
The whole RISC philosophy was a huge mistake. Yes, x86 instructions do not map well to transistors, and they have to be unpacked into uops to be executed. This is a form of compression. Having a compressed program image turns out to be a massive advantage. RISC proponents thought that x86 was so complicated they could beat Intel with their simple instruction decoders. That almost, but not really, made sense in 1990 but since then has made increasingly less sense, until today where the amount of sense this makes has hit zero. The x86 instruction decoder is a very small part of the floor plan of a modern CPU and every time they rev the microarchitecture it gets smaller. The number of transistors needed to decode the VEX prefix is like a speck of sand on the beach of a 512x512-bit multiplier.
The RISC philosophy wasn't a mistake. Our architectures have just become more sophisticated so that we don't have to make a binary choice. The hybrid is good. The internal uops get the pipelining advantages of RISC, while we get the encoding compression of a CISC instruction set.
If you look at how you can get computing gains going forward after the end of Moore's law, of course the glib answer is "parallelize across more cores!" but the more interesting path is to notice that behemoth single cores like x86 spend a ton of silicon area trying to optimize straight-line execution with things like speculative execution. If you saved all that silicon area, making each single core slower but smaller, but packed more cores on the die as a result, you would most likely come out faster. [1]
[1]: https://science.sciencemag.org/content/368/6495/eaam9744/tab...
There is no proof that these outperform traditional CPUs at all. That is the reason you don't see them being used anywhere other than niche use cases or for cost reasons.
By 2017 there were three times as many smartphones than PCs, all running ARM chips.
The top four supercomputers all use RISC, and the fastest uses ARM.
As for supercomputers, >90% of them are Intel/AMD.
Will there be Razor laptops that last less than an hour on battery that can beat them? Sure.
Will there be people who complain that the Mac isn't fast enough when plugged in? Already happening: the recent Macbook Pros have had complaints about thermal throttling, that obviously slightly larger Dell with a decent fan doesn't have.
But Apple will build performance laptops, using ARM chips, and they will be faster than the equivalent Intel Macbooks if only because they aren't throttled.
The person you replied to said:
> There is no proof that these outperform traditional CPUs at all.
To which you replied talking about embedded market share and supercomputer which have nothing to do with that.
Since now you mention Apple and MacBooks, which haven't been even mentioned, I think you are answering to the wrong thread/post.
Does TSMC have the capacity to support AMD / AWS / Ampere etc making a significant dent in the server market alongside longstanding commitments to Apple etc?
Given how much they spend on Intel CPUs to what extent is it worth AWS / Oracle etc making low hundred million dollar investments in their own silicon or startups like Ampere just to keep Intels pricing competitive?
TSMC never had capacity problem. Which mainstream media likes to run the story. You dont go and ask if TSMC has another spare 10K wafer capacity sitting around. TSMC plans their capacity based on their client's forecasting and projection many months in advance. They will happily expand their capacity if you are willing to commit to it. Like how Apple was willing to bet on TSMC, and TSMC basically built a Fab specifically for Apple.
This is much easier for AWS since they are using it themselves with their own SaaS offering. It is harder for AMD since they dont know how much they could sell. And AMD being conservatives meant they dont order more than they are able to chew.
>Given how much they spend on Intel CPUs to what extent is it worth AWS / Oracle etc making low hundred million dollar investments in their own silicon or startups like Ampere just to keep Intels pricing competitive?
I am not sure I understand the question correctly. But AWS already invested hundreds of millions in their own ARM CPU called Graviton.
From what I've heard, they didn't try very hard. Apparently they thought all they had to do was make chips, and that the sheer "technical superiority" of their process meant that they could treat their customers as second-class stakeholders, withhold information about their production timelines, etc.