The first two custom silicon chips designed by Microsoft for its cloud
theverge.com
theverge.com
The only thing that's changed is that they're scaling like crazy now and can justify overhead that comes with designing ASICs versus using off the shelf parts.
Right now NVIDIA has the lead because they have the better software, but they can't make the chips fast enough. Will be interesting to see if their better software continues to keep them in the lead or if people are more interested in getting the capacity in any form.
Supporting the common librarires that I use is very important for me to chose the cloud platform.
Because different than the ARM chip also announced in the same Ignite event, Microsoft doesn't exactly "need" nor can fully utilize an AI chip. Google trains its foundational models (e.g. Gemini) on its own TPU hardware but Microsoft's is heavily reliant on OpenAI for its generative AI serving needs.
Unless Microsoft is planning to acquire OpenAI fully and switch over from Nvidia hardware...
They're going to play a modified version of the old Rareware trick.
It's also a pretty great game to buy up OpenAI equity, which ultimately gets spent on Microsoft compute. Two birds, one stone.
Almost nobody in this game cares about profit right now.
[0] Microsoft tried to buy Nintendo very early on
None of this is really strange. It also wasn't strange when Google announced H100 systems while also pushing TPUs they developed. Microsoft has Jensen on stage because customers of Microsoft Azure demand Nvidia products. Customers of Google Cloud demand Nvidia products. So, they provide them those products, because not providing them loses those customers. It's that simple. Everyone involved in these deals acknowledges this.
I'm guessing at this point the ASICs make a lot more economic sense, though. :)
In some ways, Microsoft was 10 years ahead, but they are terrible as an organization at proliferating research projects to production across multiple orgs.
The funny thing is that this fact has been shown inadvertently by NVIDIA:
https://www.servethehome.com/nvidia-shows-intel-gaudi2-is-4x...
There cannot be more than a handful of companies in the entire world that could afford such a huge price (tens of millions of $).
In comparison with a still extremely expensive cluster of 64 NVIDIA H100, the difference in speed would reduce to only two to three times, and paying several times less for the entire training becomes very attractive.
That’s complete pocket change for any of the Fortune 500.
Such a big expense only makes sense for a company where spending that amount would bring hundreds of millions of $ of additional revenue.
I doubt that any of the companies that have already spent such amounts have recovered even a small part of their expenses. It is more likely that they bet on future revenues, but it remains to be seen who will succeed to achieve that.
Sure if there is a plausible ROI, they’d have no issues dropping that much money (actually far more). Revenues for fortune 500’s are going to be in the 10’s of billions anyway, and it wouldn’t be hard to make an argument that random AI project could increase that by a couple percent or decrease costs a couple percent, which would more than provide that ROI.
Their biggest issue is usually having anyone in leadership that has a clue enough to even propose something plausible, let alone get a team together to give it a plausible go.
If they have that, Capital is not the issue.
am i missing something here? just like you'd want to scale an h100 cluster out beyond one box of 8, you'd do the same for gaudi2?
Intel has always published the "training per dollar" because no one else competes.
Even for fine tuning you are almost always better off getting smaller GPU cloud instances.
https://press.aboutamazon.com/2023/9/amazon-and-anthropic-an...
Inferentia (inf1) was GA'ed in December 2019 so it's actually almost 4 years old now. The trainium (trn1) chips and the Inferentia 2 (inf2) refresh is indeed 1 year old though.
Microsoft Azure™ Inference for Cloud Apps© 365 Pro® Live Series X™®
> Azure Active Directory is now Microsoft Entra ID
ok, geez, thanks
Entra ID sounds like a type of ID.
I’m not sure how something could legitimately have each of these names. I assume the functionality changed pretty dramatically over the lifespan of the product?
I was pretty confused the first time I encountered it, though. Could tell from context that it had something to do with accounts, but thought maybe it could be for syncing a user’s home directory or something!
It actually isn’t a terrible name, in isolation, since a directory (like, the non-digital version) was for keeping identities. But “directory” in tech has a pretty strong association with file systems.
It technically builts upon (amongst other systems) LDAP, the Lightweight Directory Access Protocol and X.500
I guess you could think of it more like a telephone directory - which again is also where the file system metaphor has its roots I guess. So the two are not so different in the end.
A month or so ago my laptop was requiring a BitLocker recovery code before it would boot. I spent ages looking for Azure AD, before eventually discovering it was renamed to Entra. They should have had a transition period where it would have “ (formerly called Azure Active Directory)” on its name.
TPU is pretty good but is associated with Google. MTIA is an acronym but still maps to what the chip does. ~~"Cobalt" is worse as it does not mean anything~~ . Cobalt is the CPU chip, MAIA is the accelerator so this matches Meta's naming.
Funny, that's precisely why I think the names are bad. It's like if Google had chosen "Search-ola" as their name. Way too on the nose and/or lazy. Having said that, I don't really care all that much and I imagine that may have been the spirit of those who chose the names.
> So it really brings together many of the things that Intel is doing as an IDM, now bringing it together in a heterogeneous environment where we’re taking TSMC dies. We’re going to be using other foundries in the industry, we’re standardizing that with UCIe. So I really see ourself as the front end of this multi-chip chiplet world doing so in the Intel way, standardizing it for the industry’s participation with UCIE, and then just winning a better technology.
Looking at their balance sheet: plenty.
TSMC is all about pay to play. Apple is first in line because they’re willing to pay to be first in line. I have no doubt Microsoft can justify spending the money they save not buying nvidia into getting some priority access from TSMC.
Also keep in mind Apple is now on 3nm so there’s likely spare 5nm.
It’s very likely they knew something like this was coming, as they’ve been doing FPGAs for more than a decade now.
Why aren't any other companies entering this space? TSMC's growth and profits are immense, it's not like the market couldn't bear more competitors.
Also, the current situation is geopolitical insanity. It feels like China and the U.S. are on a path to war in the next decade or so. China is itching to retake Taiwan... and if you thought the U.S. fought a lot of wars over oil, there's NO WAY we wouldn't go to war to prevent all the world's most important semiconductors falling exclusively under PROC control. It's the 21st century, we would have no choice. I know that TSMC is diversifying and building fabs in the U.S. and Germany, and hopefully that will reduce the risk of war. But it's just nuts that this is even a risk. How does one company control so much of the global market?
The sooner Intel spins out the fab side into its own entity the better positioned it will be to pick up the business that TSMC doesn't have capacity for.
People will overlook a lot, if the product is compelling enough and the sales guys can keep a straight face while promising not to peek.
Apple and Microsoft are trillion dollar companies ( both nearing $3 trillion market caps ). They can afford it. Heck they have the balance sheet to acquire both tsmc and samsung if either were up for sale.
I thought Apple was moving towards being a full-stack company. Where all the hardware and software is developed in-house. If anyone has the resources, it is apple.
Having a modern chip fab binds an insane amount of capital, and makes you much more vulnerable in case of market turbulences. This is exactly a reason why there exist people who like to claim that Intel should spin off their fabs.
TSMC is so dominant because they have very, very, deep institutional knowledge. Even if you have billions to invest it's not easy to create that, and definitely impossible to do it short term.
There's only 3 companies left chasing smaller nodes. Everyone else has given up and are focusing on revenue from high volume parts in older processes.
I get the impression it's mostly physics guys.
Relevant xkcd: https://xkcd.com/435/
TLDR: Math is not. :-)
I almost went to get a CE degree because I find it all so fascinating but ultimately I'm glad I didn't.
the serious system architects, designers, and fab managers are doing just fine, software or not. the dudes in Phoenix who are getting pushed through 2-year degrees to help run a clean room are probably making 65k.
This is the crazy thing. Such fundamental important technology but salaries generally suck. Much more money in cat memes and influencer channels.
In fact for a small project in an elective he programmed a computer vision system that was far ahead of anything me or my peers had made at the time in our CS program. He had pretty much zero programming experience.
I must admit he's a particularly smart person but it was crazy that I went on to make more money than him. His work is way more complicated than mine as far as I can tell.
More importantly, the new chip production lines are dependent on hyper-specialized EUV lithography systems that only company is able to manufacture (ASML). ASML has their own limits to production too.
GlobalFoundries' Fab 1 can do 1 nm, and its Fab 8 is capable of doing 14 nm:
> https://en.wikipedia.org/w/index.php?title=GlobalFoundries&o...
ASML are solving a small subset of physics problems: how to project extremely small features. TSMC are solving many more physics problems: how to structure layers of doped silicon into transistors, how to structure those transistors into logic gates, and those logic gates into functional blocks. This is why silicon process is not just a matter of capital investment, and why nobody is going to show up overnight with 10 billion dollars and change things. It’s not that TSMC are the only ones with a big bag of cash to give ASML.
ASML's machines are very impressive, but they're just one piece of that puzzle. And ASML themselves rely on collaborators like Zeiss.
This idea that it's all just capital and others are reluctant to invest otherwise they could duplicate TSMC simply false.
Just few years ago was four companies doing the bleeding edge: Intel, GlobalFoundries, Samsung and TSMC. Intel was the best. First Global Foundries could not keep up. Then Intel fumbled ant TSMC became the best. Samsung is the only one keeping up with TSMC but they come slightly behind. Intel tries to catch up the two. No companies have "exited" they just can't compete.
The bleeding edge semiconductor node design is like doing Apollo moon program every 4-5 years. Only one can be the best.
There are? INTC and Samsung?
>falling exclusively under PROC control
It wouldnt fall under PROC control.
It would be destroyed.
The market could only bear if they are willing to pay a lot more in the name of diversification. Otherwise most would simply have trouble keeping up.
We really need some basic FAQ on HN for hardware technology related topic.
Taiwan (unlike Ukraine) is far too critical to the world economy.
Take Taiwan. You can’t “retake” something you never had.
Because it would take over a decade to build a comparable chip fab and thats assuming you had unlimited cash to throw at it.
You'd be spending billions with zero revenue for a decade and no guarantees it'll be anything as close to as good as TSMC's setup. Try finding someone willing to invest in that and it becomes clear why theres not more chip fabs out there.
IMO, MAD (mutually assured destruction - if a nuclear power start a war with another nuclear power, both will probably end up destroyed by nukes) makes the prospects of that highly unlikely.
they are literally having discussions in San Fran right now to reduce the likelihood of that happening. And it sounds like China is quite open to discussion, and the US as well.
They are building navy fleet like no other.
Still harassing neighbors ( eg. Philippines).
Nothing changed. They just want a bigger advantage.
Azure Maia is TSMC N5, I think.
Nvidia H200 and H100 are TSMC 4N.
Amazon has had Graviton chips (after they acquired Annapurna Labs) since 2018. Or do you mean specifically AI-oriented chips?
Clearly. All I got was “using ARM IP” and “TSMC N5”
"Manufactured on a 5-nanometer TSMC process, Maia has 105 billion transistors — around 30 percent fewer than the 153 billion found on AMD’s own Nvidia competitor, the MI300X AI GPU. “Maia supports our first implementation of the sub 8-bit data types, MX data types, in order to co-design hardware and software,” says Borkar. “This helps us support faster model training and inference times.”"
Add it to the list of things you can't buy at any price, and can only rent. That list is getting pretty long, especially if you count "any electronic device you can't fully control or modify".
https://ir.amd.com/news-events/press-releases/detail/1168/am...
If any of these companies truly made competitive silicon they absolutely would commercialize it.
I suspect they aren't as competitive as the press releases hold them to be, and this Microsoft entrant is likely to follow the same path. Like Google, Tesla, Amazon and others it seems mostly an initiative to negotiate discounts from nvidia.
It would be great if there were really competition. When Google was hyped about their Tensor chips they did have a period where they were looking to commercialize it, and there are some pretty crappy USB products they sell.
Now, I know that what you actually mean is selling the chips themselves to third parties :) But it's not obvious that there's any point to it given their already existing model of commercializing the chips.
First, literally everyone is already supply-constrained due to limits on high end foundry capacity. Nvidia has a ton of capacity because they're one of TSMC's top two customers. The big tech companies will have much smaller allocations which are used up just supplying their own clouds. Even if the demand for buying these chips rather than renting were there, they just don't have the chips to sell without losing out on the customers who want to rent capacity.
Second, the chips by themselves are probably not all that useful. A lot of the benefit is coming from the silicon/system/software co-design. (E.g. the TPUv4 papers spent as much attention on the optical interconnect as the chips). Selling just chips or accelerator cards wouldn't do much good to any customers. Nor can they just trust that systems integrators could buy the cards and build good systems to house them in. They need to sell and support massive large scale custom systems to third parties. That's not a core competency for any of them, it'll take years to build up that org if you start now. And it means they need to ship the software to the customers, it can't continue being the secret sauce any more.
Nvidia on the other hand has been building up an ecosystem and organization for exactly this for the last decade.
And TSMCs top customer is not even playing in the cloud space.
My bet: if it really becomes clear what capabilities an AI accelerator chip needs and lots of people want to run (or even train) AIs on their own computers, AI accelerators will appear at the market. This is how capitalism typically works.
My further bet: these AI accelerators will initially come from China.
Just look at the history of Bitcoin: initially the blocks were mined on CPUs, but then the miners switched to GPUs and "everybody" was complaining about increasing GPU prices because of all the Bitcoin mining. At some moment, Bitcoin mining ASICs appeared from China and after those spread, GPUs were not attractive anymore for Bitcoin mining (of course the cryptocurrency fans who bought the GPUs for mining attempted to use their investment for mining other cryptocurrencies).
Yet many startups and existing designers anticipated this demand correctly, years in advance, and they are all still kinda struggling. Nvidia is massively supply constrained. AI customers would be buying up MI250s, CS-2s, IPUs, Tenstorrent accelerators, Gaudi 2s and so on en masse if they wanted to... But they are not, and its not going to get any easier once the supply catches up.
Unless there's a big one in stealth mode, I think we are stuck with the hardware companies we have.
Theres also some kind of actual AI crypto project that I wouldn't touch with a 10 foot pole.
But ultimately, even if true distribution like Petals figures out the inefficiency (and thats hard), it had the same issue as non Nvidia hardware: its not turnkey.
As I already hinted in my post: I see a huge problem in the fact that in my opinion it still is not completely clear to this day which capabilities an AI accelerator really needs - too much is in my opinion still in a state of flux.
A good example of this is Intel canceling, and AMD sidelining, their unified memory CPU/GPU chips for AI. They are super useful!.. In theory. But actually, they totally useless because no one is programming frameworks with unified memory SoCs in mind, as Nvidia does not make something like that.
Can you order any of these devices online as a regular person? Anybody can order a $300 Nvidia GPU and program it. This is the reason why deep learning originated on the GPUs. Forget those other AI accelerators, even if you bought something like a consumer grade AMD GPU, you couldn't program it because it's restricted. The reason why Nvidia's competitors are struggling is because their hardware is either too expensive or hard to buy.
My bet: in 6 months jart will have models running on local or server, with support for all platforms and using only 88K of ram ;)
I, personally, am interested in retrocomputing, amateur/hobbyist electronics, and hobbyist computing (including semiconductors [2]). While these techniquess and devices may be light years away from anything resembling a computer that can compete with SotA commercial offerings, they do offer the promise of “keeping the candle lit” as it were. I will note that if you follow Sam Zeloof’s chronicles, he progressed through the earliest phases of semiconductor development far faster than the industry did back when it was pioneering the technology. Of course, he had the benefits of knowing it was already possible and access to the written knowledge of the experts who went before him.
Within a ten minute window:
- Satya announced GPT-4 runs (at least partly) on a new AMD offering
- Satya announced an in-house chip for ML acceleration
- Satya brings NVidia CEO Jensen Huang on stage
they've got every horse in the race, huh
(disclaimer, I work for MS but all the stuff talked about here far is waaaay above my paygrade haha, and all brand new info to me)
Obviously they're going to play every angle.
This should be ringing alarm bells at FTC and DoJ.
You know what's even better than trust busting and breaking up cartels? Preventing the formation of cartels and trusts in the first place.
Would it be better for competition if Microsoft only used one supplier?
If we assume that Microsoft is roughly able to architect compute units of a similar performance-to-number-of-transistors ratio as nVidia is, then having twice the number of transistors should roughly result in twice the performance.
That is very different than it is with typical software. If you give a programmer who needs to write 100 lines of code to solve a given problem 100 more lines to fill, he won't simply be able to copy-paste his 100 lines another time and by that action be twice as fast at solving whatever problem you tasked him with. With GPU compute units, such copy-pasting of compute units is exactly what's being done (at least until you hit the limits of other resources such as management units, memory bandwidth etc.).
It is like knowing the kind of engine a car has. Not all V8 gas engines produce the same power, but knowing that it is a V8 instead of an inline three cylinder does give you an idea of the expected performance characteristics.
In a way you’re right, neither tells you anything about performance.
An sf90 has a v8. So does an 83 mustang, and it even has 25% higher displacement! So clearly the 2023 Ferrari is basically a fastback mustang…
This (wanting higher density) is the opposite of the trade-off that I was expecting. In my (limited and out of date) experience, power was the limiting factor before space, and I believe AI racks have very high power draws already.
I would have guessed this would be because larger nodes would be better for AIs tight communication patterns, but they specifically call out datacenter space as the constraint. Curious if anyone knows more about this
On the other hand, if you are building your own data center, which is the case for Microsoft, presumably you can arrange high power zone to run GPUs.
Is this done as a bridge until/if Nvidia is able to deliver their processors fast enough?
I would think that competitors to Nvidea would have serious competitors on the market already if competing can be done by Microsoft for whom producing hardware is not their main business focus.
So more of an SOC a la AWS graviton or Apple Silicon than a pure GPU?
Several Chinese companies also develop chips based on N2, like Alibaba T-head and ZTE Sanechips. I worked on software tunings for both of them. It's good to see more and more Arm chips.
BTW, I'm not even speaking to whether x86 can compete at the same power per watt... I think it just won't make sense financially to be out of sync with the industry.
Are we reading the same scores?
The top consists of what appears to be an Intel i3-10100 overclocked past 13GHz(!), a Ryzen 7 5800H at 2.8GHz, and then an i9-14900K at just below 800MHz.
The i9-14900K and M3 actually haven't appeared in the official chart, but you can search for them as they already have thousands of test runs[0][1]. Both of them score around 3100 in single core, and around 21000 in multicore (for the M3 Max).
[0]: https://browser.geekbench.com/search?utf8=%E2%9C%93&q=i9-149... [1]: https://browser.geekbench.com/search?utf8=%E2%9C%93&q=mac15
Personally, I do not consider geekbench a viable kit.
ISA doesn't imply performance characteristics.
It's like saying that programming language (syntax) has performance implication.
No, it doesn't. Everything is up to the compiler, runtime and standard library.
Of course there may be some feature that make compiler's life easier, but still things are way, way more complicated than "just take ARM ISA and you'll be king"
https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-...
Microsoft isn't Apple or Google in this regard, dragging developers into new worlds, and it is quite telling that they had now to put up some kind of ARM advocacy action.
https://blogs.windows.com/windowsdeveloper/2023/10/16/window...
For example, for scientific computation and computer-aided design, Fujitsu is the only company that has designed Arm CPUs that can compete with the x86 CPUs, but they do not sell their CPUs on the free market.
For a huge company, the floating-point performance of the CPUs is less important, because they can use datacenter GPUs with even greater throughput, so the existing Arm server CPUs could be good enough even for a supercomputer, as they only have to move the data to and from the GPUs. However the small businesses and the individuals cannot use datacenter GPUs, which have huge prices, so they can use only x86 CPUs and there is not the slightest chance of any alternative that would appear soon.
Another application domain for which no Arm vendor has ever made competitive devices is for cheap personal computers.
Nothing what Apple does matters, because they do not sell computers, they only lend computers that remain under their control and which are much more expensive than their alternatives anyway.
Besides Apple, only Qualcomm, Mediatek and NVIDIA are able to make Arm CPUs with a performance similar to the cheapest of the Intel and AMD CPUs, but all these 3 companies demand for their CPUs prices that are several times higher than the prices of comparable x86 CPUs.
Like for CPUs with high floating-point or big integer performance, there is not the slightest chance for the appearance of any company that would be willing to sell Arm CPUs that are both cheap and fast.
Also for server CPUs, all the companies that have attempted to design Arm-based server CPUs have never designed models suitable for small businesses or individuals, but only models that can be bought only by very big companies.
I would not mind to switch from x86 to Arm, but there is absolutely no perspective for that.
If the x86 CPUs would disappear, that would be a catastrophe for the people who do not want to depend on the mercy of the big companies. That would be a return to the times from before the personal computers, when all computing had to be done remotely, in the computing centers of big companies, which have been renamed now as "clouds".
I agree that Qualcomm/Mediatek/Rockchip/Nvidia pricing is really terrible but I guess prices don't matter when there's almost no demand anyway.
ARM is ok only for reasonable performance at low power (if we forget about VIA).
It's like saying that programming language (syntax) has performance implication.
No, it doesn't. Everything is up to the compiler, runtime and standard library.
Of course there may be some feature that make compiler's life easier, but still things are way, way more complicated than "just take ARM ISA and you'll be king".
https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-...
Also even if you assume that ARM has 1-2% better perf than x86, then how many people will risk transition over that? Some will, but majority will no.
Because right now I'm looking to save up for a majorly spec'd Apple MacbookPro just to be able to do this stuff on a *nix operating system. I have no great love for Apple but the abilities of their chips and the vast software offerings are tempting this Linux guy in that direction.
Something that Microsoft cannot seem to do any more. I used Windows from 3.x-WinME; NT3.51-WinXP, getting off before Vista. What I've seen since then has done nothing to tempt me back to their side. Since I unfortunately must deal with Windows 10 at work, it definitely reinforces my distaste for their systems....
So despite thinking OSX has been rendered ugly for the past ten years now, I'm still thinking heavily in that direction, even with the high costs. Snapdragon X sounds nice enough but I have zero expectations based on past behavior at those getting decent Linux support any time soon. And no one else seems to even be trying, that one Thinkpad aside.
I’d expect a future hypothetical Microsoft ARM laptop to be like a surface-RT; some Windows dropped on a third party ARM chip. Microsoft is a software company, after all. So it is more of a matter of, do they happen to have bought a chip that supports Linux (probably yes, because what hardware manufacturer wants to be dependent on one company for OS support?) and can you get past Secureboot (probably yes, after a couple years at least, when the jailbreak happens).
I'd never tell the higher ups this but it was pretty easy, too. I'll let them bask in my glory of saving the company $60k/month.
Nvidia made some really amazing strides in the past few years, taking over cloud gaming where Onlive and Stadia utterly failed, making DLSS, etc.
I just hope they don't abandon us gamers for their AI stuff :( Probably the entire gaming market is way smaller than the potential AI market, just hopefully not too small to matter.
[1] A 1970 Corvette 427 has a 0 - 60 mph of 5.3 seconds (src: https://www.caranddriver.com/features/g15379023/the-chevrole...) and cost around $44k inflation adjusted dollars. You can buy a 2008 Nissasn 350Z Enthusiast that will do it in 5.2 (src: https://www.zeroto60times.com/vehicle-make/nissan-0-60-mph-t...?) for around $13k today.
[2] I'm too lazy to calculate relative cost / cycle in old warehouse computers vs phones but it's gotten _better_.
Most of the value of chips is in their design, which is owned by different entities. Manufacturing is important too (only TSMC can make these advanced designs at scale and at lower costs than the competition).
The question I have is if Cobalt has any innovations in its design, or if its just bog-standard ARM Neoverse cores. Its not too big of a deal to download ARM's latest designs and slap them into... erm... your designs. But hopefully Microsoft added value somewhere along the road (The Uncore remains important: cache sizes, communications, and the like).
Presumably this means that Cobalt has a much lower power consumption than the current Ampere CPUs used by Azure.
Most of the power consumption reduction for a given performance may have come from using a more recent TSMC process, together with a more recent Arm Neoverse core, but perhaps there might be also some other innovation in the MS design.
But if TSMC is the only company that can do this, they're a bottleneck for the entire world. Not to mention a strategic and geopolitical risk for the West.
It's be nice if some domestic companies invested in fabs again...
You likely became bewitched by their glamorous marketing side. I'd bet that the real work that the team does is very similar to the work that basically every ASIC design team does.
I bet you haven't used any Microsoft product before. /s
My personal preference is to avoid this.
There are real profits in the chip space and considering that there are 3 fabs and one clear leader who will make anything for anyone this is a sign that NVIDIA are doing a great job.
It makes a lot of sense from the point of view of cloud giants.
It wasn’t that long ago that computer manufacturers would build their own chips.
Amazon
Apple
Goole
Microsoft
> Microsoft is currently testing its Cobalt CPU on workloads like Microsoft Teams and SQL server,
Teams is so bloated it needs its own 128-core CPU to run well. /s