Ampere switching to in-house CPU designs
anandtech.com
anandtech.com
My message for people trying to make "ARM Server" CPUs a thing is that you have a real chicken or egg problem to solve, getting these things into the hands of people doing Linux development on x86-64 right now. Without a robust ecosystem of all the Taiwan-based motherboard manufacturers behind you, it's never going to get more than a few percent market share.
I've been a CTO for a decade+ and did try to get one the last time they made a splash.
The problem with this stance is that they can also keep using the non elitist platform, considering the elitist platform is inferior for many use cases.
No sparks...
Basically there was no way of getting your hands on them. They were interested only in "Serious customers".
Having seen this play out a number of times I don't think targetting that segment and only that segment is a great idea.
From a stack perspective, you can't really go it alone anymore. The total necessary software universe is too large to rebuild on a reasonable timescale, and too expensive to rebuild by oneself.
IBM is about the smallest scale that can do this -- and even they're leveraging decades of effort.
If the market isn't established then you may need to get traction with the dreamers and misfits.
I think the lack of availability certainly didn't help, but the architecture itself was probably the biggest downfall of the IA-64 platform. When it was introduced, the hardware was certainly advanced, but relied too much on software to get proper performance out of the chips. The product was doomed from the start.
Ampere will likely have entirely different problems bringing their architecture into the market, but I don't think Itanium ever stood a chance even if they'd sent out free chips to as many influential people and companies as possible.
The architecture was probably a fatal mistake on its own (even if Intel had disbanded its x86 line, the “older” RISC processors outperformed, often substantially) but ensuring that almost nobody saw anything close to what performance it could deliver definitely made it fatal. If Intel hadn’t been greedy trying to charge for compilers, the gap wouldn’t have been nearly as bad.
This must have been a huge part. Maybe ICC could squeeze more MIPS per dollar out of Itanium than x86, but if GCC can't, people are going to not buy Itanium rather than buy ICC.
I think that's the kind of issue they're talking about. If it had been easier for nerds in bedrooms to run Itanium, we'd've seen much better support for it in GCC etc. much sooner.
Back when I supported computational scientists, the total cost of the servers + compilers was so steep that even if everything had worked as well as the Itanium salespeople promised it would have needed to be many times faster to be worth the hassle of changing toolchains, especially since you had to pay up front for the hope that with a fair amount of work you'd at least hit the point of breaking even.
If Itanic had won, we'd still be basically where we are today from a hardware perspective, but with even less competition.
Apple made it happen on the client, but Amazon is making it happen on the server. The people who want to run their own DCs or hardware will end up part of the niche left on x86. I don't see Apple doing anything in the server space beyond providing minis like they do now.
iServer with 64 cores and 256 GiB DDR4 for only $ 699,999.95
You are off by an order of magnitude :)
If they are to be commercialized, I’d expect them to be part of a development stack related to Xcode, or some kind of AWS like services for applications developed for the Apple ecosystem.
Offering generalized compute would invite too much initial cost performance comparison and they couldn’t keep up with demand.
(I’m not sure whether that includes servers at large cloud providers, but I don’t think that matters. Unless there’s a huge performance advantage, I don’t see Amazon, Google or Microsoft buying Apple server CPUs)
I never did deliver for the OLPC project, but I tested a lot of things other people were working on. I also finally grokked the free software idea and became and engineer instead of a history teacher.
The $250 or so of prototype hardware has been repaid in testing, bug reporting other things I was able to do then. It sits on a shelf in my office to remind me how I got started, and what it mean to really be a part of a community that takes risks to invest in newcomers.
Once Raspberry Pi CM4 support is added, I assume that you would be able to work on drivers for PCIe devices as well. Besides, non-server vendors like NXP have released official boards with ARM ServerReady support.
I'm not affiliated with Avantek in any way. I just did a bunch of research looking to build/buy a desktop, and ARM was a contender, but I ended up building a beefy x86 machine with a Ryzen instead.
"I can pretty much guarantee that as long as everybody does cross-development, the platform won’t be all that stable. Or successful.
Some people think that “the cloud” means that the instruction set doesn’t matter. Develop at home, deploy in the cloud.
That’s bullshit. If you develop on x86, then you’re going to want to deploy on x86, because you’ll be able to run what you test “at home” (and by “at home” I don’t mean literally in your home, but in your work environment).
Which means that you’ll happily pay a bit more for x86 cloud hosting, simply because it matches what you can test on your own local setup, and the errors you get will translate better…
Which in turn means that cloud providers will end up making more money from their x86 side, which means that they’ll prioritize it, and any ARM offerings will be secondary and probably relegated to the mindless dregs (maybe front-end, maybe just static html, that kind of stuff).
Guys, do you really not understand why x86 took over the server market?
It wasn’t just all price. It was literally this “develop at home” issue. Thousands of small companies ended up having random small internal workloads where it was easy to just get a random whitebox PC and run some silly small thing on it yourself. Then as the workload expanded, it became a “real server”. And then once that thing expanded, suddenly it made a whole lot of sense to let somebody else manage the hardware and hosting, and the cloud took over. (All emphasis original)"
For example, if I'm building on Firebase with Cloud Functions I'm so far abstracted from the underlying hardware that these architecture specific issues are less important.
Similarly, if I'm working in Django/Python on AWS RDS, using an ARM server for my EC2 instance might be an issue ... but moving my RDS Postgres instance onto ARM might not make a difference.
I have no specific knowledge here but I imagine the x86-64 builds are far better tested, regarding correctness. As cmrdporcupine points out, the performance of the x86-64 builds will also have been better studied. Whether it's an issue in practice, I don't know, but if you're looking to minimize risk, there are good reasons to stick with x86-64.
I can certainly imagine a 'serverless' platform like Lambda or something where running on ARM is just an implementation detail (if it's not already).
Whether you're right or not, I think I agree with the idea.
It also seems to me that, in general, working with a different CPU architecture and instruction set implies that the new & improved stuff has to be significantly better to be worth the cost and risk.
At this point, is one really that much better than another? Are large improvements happening at the margin?
Another camel that'll poke it's head in the tent, in large/expensive embedded (which I realize that the parts in the article aren't aimed at) is whether a CPU and it's battlegroup of other stuff will be supported right off the bat by software infrastructure and later on by actually being available in the long run. You have to wonder how out on a limb anyone who used a Motorola 88k felt.
From 1984 to the iPhone, I believe it was the norm to develop Mac applications on Macs, except for a brief period when the very first models were starved for memory.
It is, but it introduces enough differences that you're not "at home"; indeed I'd say OSX-to-Linux is a bigger difference for development than x86-to-ARM.
Even more people use macbooks to develop for the browser, or to develop using platform-independent languages: Python, Ruby, Node, maybe even golang.
Quite a few people then package their backend code in a container and test it in a Linux VM which "Docker for Mac" helpfully provides.
But indeed the fact that Macs used to run on Intel, so the libraries you used in development were mostly that same versions of the same code as in production, must have been helping.
People only develop on MacOS because of how similar it is to Unix, not in spite of it. Go ahead, look at how many developers were using MacOS before OSX. I'll wait. The reason why people develop on Macs is because it's "close enough" to Linux to deploy with, and "good enough" as a desktop to keep tabs on most of your other apps.
> So while there may be some premium to being able to "develop at home", clearly many developers value other considerations more highly.
Let's clear this up: there is no development without development at home. If you work for a company as a developer, you'll be expected to maintain a machine that can consistently build and deploy your software. Unless you're working as a web developer (in which case, may god have mercy on your soul), you're probably going to end up using an x86 machine. Hell, my best friend is a die-hard Apple fan, and even he has admit that he can't really upgrade his Xeon iMac unless Apple offers him an upgrade within the x86 ecosystem.
Why can't he upgrade? (Genuine question)
There were plenty of other problems with pre-OSX MacOS.
> The reason why people develop on Macs is because it's "close enough" to Linux to deploy with, and "good enough" as a desktop to keep tabs on most of your other apps.
The differences between x86 OSX and x86 Linux are bigger than the differences between x86 Linux and ARM Linux, in my experience. So these machines are also "close enough" to ARM servers.
> there is no development without development at home. If you work for a company as a developer, you'll be expected to maintain a machine that can consistently build and deploy your software.
Most developers have a local development environment and a separate production environment. Developers place some value on having them be similar, but they're also happy to introduce differences between local and prod for the sake of other tradeoffs.
I wouldn't be at all worried about deploying my Ruby, Python, or Java app on ARM, having worked with x86 locally.
That is, as long as the Ruby, Python, and Java development teams themselves have been able to test on ARM hardware themselves!
I think Linus' comments still apply to those higher level languages. They all rely on platform dependent code.
Java less so, but the behavior and performance characteristics are definitely different.
I know Java has C extensions for some things too. I'm curious how much code is actually "pure Java", but I've actually rarely seen Java used at any of the places I worked at.
And then there are all the other native and system packages that you depend on. For example, I do a lot of work in Python, and I would be very hesitant to have anything but an Intel CPU in my workstation. Because that's what our servers run, and I know that a lot of packages I use include native code that has been carefully tuned for Intel CPUs. If I choose AMD or get an M1 Macbook, I'm going to have a much poorer sense of how the things I'm working on would behave in production.
That said, I'd be willing to at least try it. On the other hand, if work started talking about switching to ARM servers in production, I'd instantly be clamoring for a (non-M1) ARM workstation to develop on, because I need to know that I'm not going to get into trouble with packages that get a lot of love in the x86 versions of the native libraries but not so much on ARM.
This was a good argument a couple of decades ago. Now? Keep in mind that most devices in this planet are ARM-based, ever since we started talking about 'smartphones'. They have been running these interpreters for a long time now. Most applications and libraries of any importance have been ported.
> If I choose AMD or get an M1 Macbook, I'm going to have a much poorer sense of how the things I'm working on would behave in production.
I thought that's what staging environments were for? Plus scale tests? A laptop provides a pretty poor proxy for a server environment behavior. For starters, they are not really the same spec, are they? Are you running a Xeon-based laptop?
Or, for example, there's one library I use that includes some inline assembly language routines. Which target SSE, not SVE. So, if we're targeting different ISAs, then my dev box isn't even executing the same code as what would be running in production.
That seems crazy to me. What bizarre excuse does your organization have for that policy?
To me, part of the point of a staging environment is to provide something as close as practical to prod in which to test, reproduce, evaluate, and investigate issues, be they performance, correctness, or what have you.
Even if it were allowed, it would still be awkward. I want to, as much as is practicable, find problems as early as possible, and not wait for them to get all the way to staging. I'm not really interested in making myself less able to do my job effectively just because I can.
Nowadays a browser takes that role.
Having a workstation with the same CPU is definitly not an issue as Linus makes it to be.
I develop and test against docker containers that run on a specific image + arch + base OS.
Even though my environment is supported on ARM, I am still not keen on using ARM servers even if I think it will work. There's just too much risk involved and I don't have the time to set it up and test it thoroughly.
With Python, some of the really useful libraries are written in C so you really need consistency between your environments if you don't want gotchas in your deployment pipelines.
Most web stuff dont deploy on x86, they deploy on VM or more specifically JVM. Which is someone else's problem. Amazon and Oracle are both spending resource making sure it works flawlessly.
I mean seriously, most doing Web Development dont even want to touch the VM.
You may want to ask Amazon, who is trying to get as many Graviton 2 produce as possible from TSMC because the demand for them are outstripping supply.
And all of that was before Apple even announced their switch to ARM on Mac.
For some workload it is half the cost on R6g than x86 instances. And for some like Twitter, they will invest into making sure it works with ARM. Those works get filter down the pipe. Medium-Large size companies see ecosystem matures and clear cost benefits they will move.
"Hyperscale cloud" is just buzzword for "huge data center somewhere else" so there's not much of a difference except packaging and power (machines for hyperscale data centers often take direct DC power, not AC into one PSU per machine.)
ARM's dominance in the cell phone market means it is ample popular enough, got lots of eyes on it, and increasing amounts of software is picking up dedicated optimisations for it.
I'm not sure what else you're really expecting. The hardware is there. The software is there. Very likely you could switch to ARM without noticing.
[1] Last time I mentioned this to someone involved in POWER architecture, they told me there were going to be cheap POWER based systems coming out to meet the "home hackers" market. I'm not paying particularly close attention, but I haven't seen anything yet?
Pay-As-You-Go is great for trying out new architectures: it's easy to justify spinning a few instances to test whether something can work under ARM. And if it does then convert existing instances to save money. No (hard to get unless you are a large business) dev boards to purchase.
It's available but as you said, not at any reasonable price. The motherboard alone seems to be $5,000.
When I first started developing on AWS, I bought the company a beefy server so we could run virtual machines on it to "test it locally". We bought a dual socket supermicro board and put our choice of Intel Xeons on it. It ended up being north of $2000 (ten+ years ago). At the time AWS c1.xlarge were $0.80 per hour. Paying per hour just felt wrong. And then, within a few months, we had a working system, and we were running 250 c1.xlarge for hours on end, and charging 50x for the end result. The $2000 machine sat idle and was a total waste of money. We developed on our local machines (a Mac), and deployed to test machines on AWS.
Now I pay some hosting company $x per month to host my eldest's minecraft server. I don't want to manage that shit! I have money but not time, and I use my time effectively. That means I don't waste my time buying and configuring servers.
There have been many failures in the past to outdo ARM's cores: Intel, AMD, Nvidia, Samsung, Qualcomm. Granted, some of those are ancient history or were more management failures than technical failures. It seemed to me that with ARM finally focusing on developing high performance cores, and the success of N1, the industry would coalesce around that. ARM has control over the ISA development which can influence and be influenced by microarchitecture designs, though they do work with Apple and Fujitsu somewhat. ARM gets to combine license fees from many customers into a single large R&D budget, much like TSMC. There's a huge risk that you spend billions for years and end up with custom cores that are worse than what ARM puts out.
So why design custom cores?
Maybe it's a necessity for a viable CPU company. Maybe it's too easy to create a custom CPU with stock ARM cores so unless you can do substantially better than that, all your biggest potential customers will just make their own: Amazon, Microsoft, IBM, Oracle, Baidu, Google, Tencent, Alibaba, are all perfectly capable of putting some stock cores together. Yes, the devil is in the details but I don't think the details are enough to base a company around. Maybe even if making a better microarchitecture is unlikely you have to try or be doomed to failure in the long run.
Maybe they believe their engineering is better. It's tempting to discount the value of individual contributors in projects involving hundreds or thousands of people. Maybe Ampere believe their engineers and/or managers are just better than the competition, and can beat them with fewer resources.
Maybe it's not as hard as it used to be. The received wisdom after decades of Intel dominance was to not even try to compete. Now we seem to be reaching a plateau in microarchitecture innovation. Yes, CPUs are still getting faster but the designs seem to have converged a lot. It certainly seems easier to catch up given that Apple, ARM, and AMD have all succeeded recently. Ampere also starts with an A so they're in good company.
Maybe it doesn't actually have to be substantially better. Maybe the N1 license fees are high and they just need something similar for less money. Maybe they just need something different that their sales team can sell as being better.
I don't think it's the wrong option for most of these companies; the ARM designs are really designed with a bit too much of a focus on power consumption and die area for server applications. The Neoverse V1 is the first ARM design really targeted at that market, and is shipping late this year AIUI. (That said, the Neoverse N2 is shipping not much later, and is a generation newer design, so we might end up with more N2 than V1, despite the V1 being a higher end part.)
And using a custom uarch is also a statement to your customers, shareholders, competitors etc. that you mean business rather than just going the well-trodden path with ARM cores.
That's because a consequence of the second paragraph is that it's potentially business-profitable to make a "custom uarch" even if you make no use of the enormous scope for design changes.
The comment you replied to said "with a given ISA, how custom can the uarch be", implying that it would not be very custom.
You replied, in your first paragraph, that on the contrary, it could be pretty custom, with the implication being that it probably would be, because that's the value proposition that Ampere is presumably offering.
So your paragraph 1 is suggesting that the sentence "Ampere will now use a custom uarch" is an argument that Ampere's custom arch might well be pretty highly customized.
ON THE OTHER HAND....
Your second paragraph points out that simply being custom is a social signal (the signal of being, in your words, "that you mean business rather than just going the well-trodden path"). You argue (and I agree) that this social signal might make you more attractive to some customers. I.e. the signal might increase sales. This implies that it would be reasonable for Ampere to spend a certain amount on a barely-custom custom uarch, purely for sales/marketting reasons, as long as it didn't substantially harm performance. You know, because of the implication.
So your paragraph 2 is suggesting that the sentence "Ampere will now use a custom uarch" is NOT an argument that Ampere's custom arch has any particular need to be highly customized.
Well, Apple uses the ARM ISA for their A and M series chips, and but not ARM's microarchitecture, and their chips have significantly higher IPC than comparable parts from Qualcomm or Samsung. Similarly, Intel and AMD both implement the same x86 ISA, but the performance advantage has swung repeatedly between the two. Microarchitecture might be "in large parts" dictated by the ISA, but there's still enough left over for hardware designers to have a significant influence on the overall performance.
This is very much not the case. Just consider that, for example, Intel still supports running the same ISA that ran on 8086. The microarchitecture now running those instructions isn't at all like the 8086 microarchitecture. (Its running a bunch of other instructions too, but the points stands)
An example in the reverse direction, is that a some MIPS chip teams saw that MIPS was dying, and reused their MIPS microarchitectural designs to build ARM processors.
Sure, ISA can be designed assuming some things about microarchitecture, and then it can be inconvenient to change it. But not impossible. And during the transition to ARM64, ARM took the opportunity to remove things that were inconvenient for different microarchitectures (eg, directly accessible PC reg, changes to how condition codes work, etc)
Refer to RISC-V unprivileged spec, chapter one. Not dictating microarchitecture is one of its goals.
Personal take: it would be the end of them if they weren't.
I'm not sure that's true though..
With the way Apple operates, we have ten years before we can reasonably assume that they would change their architecture. Is RISC-V really going to be worth all the headaches of a transition in ten years?
"The company was founded in November 1990 as Advanced RISC Machines Ltd and structured as a joint venture between Acorn Computers, Apple Computer (now Apple Inc.) and VLSI Technology."
https://en.wikipedia.org/wiki/Arm_Ltd.#Founding
More details about ARM licensing:
"Finally at the top of the pyramid is an ARM architecture license. Marvell, Apple and Qualcomm are some examples of the 15 companies that have this license."
https://www.anandtech.com/show/7112/the-arm-diaries-part-1-h...
I can bet $100 on it being wrong. Especially the ARMv8 ( and v9 )
Long story short, it's probably not terribly utilitarian to have strong opinions about the subject when participating in an international forum such as this one.
I know virtually nothing about the internals, but that‘s also what people said about ARM.
> If X86 still competes I see no reason why RISC-V can't also
First RISC-V needs to catch up with the 40 years of x86 improvements that kicked out SPARC, Cray, SGI and others from high-end performace systems.
I'd suspect an ISA designed in the last few years with performance in mind (even if not a priority), wouldn't take as long to get to where the industry is now, but it's still a lot of years of focused effort required, and we'll see.
I don't have a horse in the race, except that I do like my experience to not evaporate; and I sincerly hope real mode segmented address edge cases aren't something I'd need to relearn on another ISA even though they're still (barely) relevant on x86 in early boot.
Saying X86 is 40 years work is also kind of stupid, as a huge amount of the ISA dead weight, and high performance implementations have their roots much earlier than that (e.g. Intel have tried to move on at least once and ended up coming back to the OOO superscalar paradigm).
1. Already ported.
2. Not dependent on modern performance, and thus well served by emulation.
Just shows the power of an establed ISA (all the software that runs on top of it) and how solid ARM's business model is.
The ISA is an interface, and as such may not be copyrightable or even patentable. Even if it was, we've had various interoperability exceptions, and one does need an ARM ISA to run a program compiled for ARM.
Besides, as far as I know ARM doesn't just license the ISA. It licenses cores. There's a good chance that if you make your own cores to implement an ARM ISA, you wouldn't have to pay any fee. You may however have to get around trademark, and not call your CPUs "ARM CPUs". Though I suspect "ARM compatible" would work perfectly (just like "IBM compatible" worked with third party PC vendors).
The terms are secret, but basically yes, AMD had access to 386? designs and what not as part of IBMs second source requirements with Intel, but Amd486 and up designs were AMD original with instruction sets under license from Intel, although the licenses were set up after release under litigation.
> Does Intel pay AMD a fee for using their 64 bit extensions?
Yes / sort of. Use of the AMD64 extentions was negotiated under the broad cross licensing between the two companies.
> Does anybody pay the inventor of the SSE and AVX instruction sets?
Do the persons involved get royalties? I'd guess not, but who knows. Would you have to pay more to get a license for them, almost certainly.
You'll note there's been a lot fewer x86 processor designs not from the big two over the past many years than there were in the late 90s. Some of that is because 'RISC is going to change everything', but a lot of it is becsause nobody can get an x86 license.
> Does anybody pays Nintendo for deploying emulators?
I don't think so. Patents are too old, and there's not a whole lot of non-Nintendo emulation in the commercial market (but I know Capcom has released some collections, of Megaman games, etc). If those types of releases started including systems with system roms, then you would start having copyright claims.
Nintendo has never seriously impeded system emulation; there's just no case to be made. ROM distribution is clear copyright violation though, of course.
I don't agree with that. The base ISA is highly optimised for high performance. Many of the decisions made are precisely because it makes it easier to design very wide superscalar CPUs. The vector extension is also looking like it'll be more efficient than ARMs.
What's missing, is instruction set extensions that makes it comparable to ARM/x86 across all workloads. And obviously it's far less mature in general.
RISC-V has been optimized for a single-purpose, being as simple as possible to be easily taught to students and implemented by them in a limited time.
It lacks many features required for high performance without hardware of excessive complexity, the most obvious being that RISC-V does not have decent addressing modes, so that the highest performance implementation reported until now (by Alibaba at the 2020 Hot Chips Conference) had to add a non-standard ISA extension to correct this.
Nobody remotely competent chooses RISC-V for "high performance" in any definition of that term.
RISC-V can be the right choice in many cases because:
1. No costs for using the RISC-V ISA
2. Easy customization with non-standard extensions
3. An already existing complete software development environment, with compilers, debuggers etc.
4. Acceptable performance at a given implementation cost
There is no need to invent extra fictitious advantages, like "high performance".
It is actually a replacement for MIPS in a more literal sense: The company owning MIPS has recently renamed itself to MIPS, and it abandoned MIPS ISA in favor of RISC-V ISA.
Do you have any proof that additional addressing modes would increase performance? What I remember from college and papers I've read is that more modes are generally a detriment to pipelining.
Maybe you're talking about "needing more instructions". You'd be wrong in this case. Using the compact instruction set gives an average of 15% more dense code compared to x86 and somewhere around 25-35% more dense compared to aarch64 (about equivalent density to thumb, but without the switching overhead that makes thumb slow).
"Xuantie-910: A Commercial Multi-Core 12-Stage Pipeline Out-of-Order 64-bit High Performance RISC-V Processor with Vector Extension"
From "VIII. NON-STANDARD INSTRUCTION SET EXTENSION"
Targeting at various industrial applications, XT-910 enables a set of custom non-standard instructions, additional to the standard RISC-V instructions. The non-standard instruction extension can be categorized into two groups based on the purposes - memory access enhancement, and basic arithmetic operation enhancement.
A. Memory access enhancement
Memory access instructions usually account for a high proportion in the total number of instructions, so enhancing memory access instructions can directly benefit the overall performance. By analyzing the mainstream applications running on RISC-V, we observed that for the basic RISC-V instructions there is still quite some room for improvement in memory access related instructions.
First, we support register + register addressing mode, and support indexed load and store instructions. This type of instruction extension reduces the usage of the registers for calculation and reduces the number of instructions for address generation, thereby effectively accelerating the data access of a loop body. Second, unsigned extension during address generation is supported. Otherwise, the basic instruction set does not support direct unsigned extension from 32-bit data to 64-bit data, resulting in too many shift instructions.
I'm not certain if they break out the results by individual optimization, but in Fig. 20 in the paper it looks like all their tweaks add up to a 20% boost over the standard RISC-V ISA. Pretty huge.
>I'm not certain if they break out the results by individual optimization
They don't, so it's not clear where the 20% comes from. They're using a compiler that optimizes for their microarchitecture, which might account for a lot of that. As for the mention of "extensions", it's not even clear to me whether they mean official RISC-V extensions such as IMAFD or custom extensions.
[0] https://riscv.org/technical/specifications/
(start with chapter 1 of the unprivileged spec)
As a practical counterexample to that FUD, it is also the ISA chosen by the Barcelona Super-computing Center EuroHPC[1] to design a new high performance CPU for future supercomputers.
[0]https://riscv.org/technical/specifications/
[1]https://www.bsc.es/news/bsc-news/bsc-working-towards-the-fir...
Generally, for RISC-V, look to China it seems?
However, RISC-V is so minimal that it might be possible to add it as a second ISA to a chip without much effort. On the other hand, it's so minimal that a JIT translator from RISC-V to the native ISA can be a valid alternative (see for instance https://carrv.github.io/2017/papers/clark-rv8-carrv2017.pdf).
Of course, it helps if the ISAs are at least broadly similar in terms of register space and available operations and memory ordering specifications. And the performance is in the details.
It's good news for open source hardware development. But bad news for USA and for people with security concerns. I don't know how comfortable I'd be with a Chinese designed CPU, Chinese compilers, and a Chinese kernel ... even if they were all open source.
The US Empire is corrupt, but its processes of decision making and power are relatively transparent (open elections, senate trials, etc.). The Chinese system is entirely opaque. In a sense, all citizens in China are ultimately completely vulnerable to the whims of the state ... there is no recourse for arbitrary detention, seizure of property, etc. This is a model China seeks to export in the long run ... either directly, or through affiliates.
America sucks ... but as people often say about Democracy, its the worst system of governance except all the rest.
I’ve studied Chinese politics and have reached the opposite conclusion to you.
Abu Ghraib wasn’t made public through courts at all.
How many other such places exist and still kept secret for now? The past would suggest there are at least several.
If you lived in America, you would see the huge amount of protests on the streets against the Iraq war and other international adventures. Also, US politicians with significant power regularly speak out against such adventures.
This is America's internal control mechanisms. There are no such mechanisms in China.
Anyway, there's no point discussing this further. Either you understand this level of nuance, your you inherently believe in the notion that "all countries are equally bad" type philosophy with no regard to how information is brought to light and how a country is constrained.
Guantanamo bay is still open, despite what happens there being illegal under Cuban, US and international law. Julian Assange and Chelsea Manning are still persecuted.
You either understand that your country is oppressing the entire world or you blindly believe everything your ruling class tells you.
I don't think you can patent trivial details of "API" but you can absolutely patent a given function of an ISA.
You probably shouldn't be able to (Alpha vs. MIPS, x86 etc.) but currently you can.
Google are working on their own ARM Server Chip. And we finally have ( some sort of ) confirmation Microsoft will be using Ampere.
At least Ampere will have fixed their Go to Market problem where no hyperscaler were committed to buying ARM instead of building their own. Along with Tencent Cloud which is a big surprise.
Which got me to question why Marvell lost the battle to Ampere.
Assuming the default mode of operation is having just a few cores operating at 100% with the majority running at 0%, it's wasteful to not let those active cores run as fast as possible.
Obviously the default mode of operation is to run all cores, or very nearly all cores, at full tilt all of the time. This chip is designed to be garbage at few-threaded workloads.
That’s I suppose a simplification they can make by focusing only on “datacenter” while intel needs the same design to work from 1 socket to hundreds.
>In fact, Ampere explains that what the move towards a full custom microarchitecture core design was actually always the plan for the company since its inception, and their custom CPU design had been in the works for the past 3+ years.
So my guess is that they always had the license but never put out the announcement until they actually have their custom core.
Although in modern days PR speak a "custom" core could also be a custom ARM X-1 [1] type of custom. We will have to wait for more details.
Their first product (eMAG) was based on the last APM X-Gene (3?) cores.
They also had two from-scratch developed microarchitectures.
Apple and AWS seem to leave the competition behind.
Subtle but significant difference. The company is sentence case now.