Intel 3rd gen Xeon Scalable (Ice Lake): generationally big, competitively small
anandtech.com
anandtech.com
That said, fixing Intel is a three step process right? First they have to get their process issues under control (seems like they are making progress there). Second, they need to figure out the third party use of that process so that they can bank some some of revenue that is out there from the chip shortage. And finally, they need to answer the "jelly bean" market, and by that we know that "jelly bean" type processors have become powerful enough to be the only processor in a system so Intel needs to play there or it will lose that whole segment to Nvidia/ARM.
Yields on larger chips are still substantially worse than smaller chips, even with harvesting.
once you get those easy wins the rest take a bit to chip away at.
while initial growth was 100x, that pace is going to slow down as marketshare increases.
Day by day, more data centers get AMD systems by choice or by requirement (Oh, you want 8xA100 nVidia made modules with maximum performance. You need an AMD CPU since it has more PCI lanes, for example).
You don't see much AMD server CPUs around because first generation and most of the second generation has completely bought by FAANG, Dropbox, et al.
As the productions ramps up with newer generation, we can buy the overflowing parts after most of the production is gobbled up by these buyers.
When you configure the system full-out with GPUs & HBAs, the number of lanes becomes a matter of necessity rather than a spec which you drool over.
A PCIe lane is a PCIe lane. Its capacity, latency and speed is fixed, and you need these with minimum number of PCIe switches to saturate the devices and servers you have, at least in our scenario.
All of the other benchmarks in the PCMark test suite push the bottleneck down to the storage device.
One would think Intel might want to build a storage array that could stress the PCIe lanes but then that might show an entirely different picture than the one Intel is portraying.
PCIe efficiency is such buzz-none-word.
I'm genuinely having trouble understanding what you mean by this.
Rocket Lake being backported to 14nm means that 10nm can be allocated in greater proportion toward higher-priced chips like Alder Lake and Ice Lake SP. Seems like it would be good for production volume.
In laptop processors they can easily show efficiency gains from the 10nm process and IPC improvement, It appears most laptop processors end up running a power envelope lower then the ideal performance per watt efficiency. Server Processors with higher core counts means you can run more workloads per server, again providing efficiency gains. However desktop / gaming tends to be smaller core count + higher frequency with little concern of efficiency outside of quality of life factors (i.e. don't make me use a 1KW chiller). Intel has been pushing 5ghz processor frequency for years, and rocket lake continues that push (5.3ghz boost), when they drop frequency to move to 10nm, its hard to see an IPC improvement that is able to paper over that.
However alder lake CPUs will have a thread count advantage, so at least with 24 threads it should be able to show generational improvement over the current 8c/16 rocket lake parts. That will allow them to at least argue their value with select benchmarks and intel only features. Those 8 efficiency cores will likely be a BIG win on laptop, but on desktop I doubt they will compare favourably to the full fat cores on a current Ryzen 5900x (i.e. a currently available 24 core processor).
Intel is going to have at least 1 more BAD mainstream desktop generation before they can truely compete on the mainstream high end, however there is a chance they have something like a HEDT part that would allow them to at least save face. That being said, given a choice, Intel will give up desktop market share for the faster growing laptop and server markets.
I'm not seeing a good reason for thinking this is the case. Server CPUs are harder to fab (much larger die area) and they need to fab more of them (desktop CPUs are relatively niche compared to mobile and server CPUs).
If anything this is a sign that 10nm is fully ready.
> [1] Jelly Bean chips are those that are made in batches of 1 - 10 million with a set of functions that are fairly specific to their application.
Think Atmel AtMega parts, there are trillions of these in various roles. When you think of something like a 555 timer[1] that is now more cost effectively and capably replaced with an 8 pin micro-processor you can get an idea of the shift.
While these are rarely built on the "leading edge" process node, when a process node takes over for high margin chips, the previous node gets used for lower margin chips, which effectively does a shrink on their die increasing their cost (most of these chips seem to keep their performance specs fairly constant, preferring cost reduction over performance improvement.)
Anyway, the zillions of these chips in lots of different "flavors" are colloquially referred to as "jelly bean" chips.
> Jellybean is a common term for components that you keep in your parts inventory for when your project just needs “a transistor” or “a diode” or “a mosfet”
-----------
For many hobbyists, a Raspberry Pi or Arduino is a good example of a Jellybean. You buy 10x Raspberry Pis and stuff your drawer full of them, because they're cheap enough to do most tasks. You don't really know what you're going to use all 10x Rasp. Pi for, but you know you'll find a use of it a few weeks from now.
---------
At least, in my Comp. Engineering brain, I think N2222 transistors or 3904-transistors, or the 741 Op-amp. There are better op-amps and better transistors for any particular job. But I chose these parts because they're familiar, comfortable, cheap and well understood by a wide variety of engineers.
Well, not the 741 OpAmp anymore anyway. 741 was a jellybean back in the 12V days. Today I think 5V compatibility has become the standard voltage (because of USB). So 5V op-amps are a more important "jellybean".
I think your jellybean op-amps would more likely be TL072, LM358, or NE5532.
Maybe it was fully obsolete by that point, but high school + neighborhood libraries aren't exactly filled with up-to-date textbooks or the latest and greatest.
I remember that Radio Shack was still selling kits with 741 in them, as well as breadboards and common components... 12V wall-warts and the like. Online shopping was beginning to get popular, but I was still a mallrat who picked up components and dug through old Radio Shack manuals into 2005 or 2006.
It was the ability to walk around, and see those component shelves sitting there in Radio Shack that got me curious about the hobby and start researching it. I do wonder how modern children are supposed to get interested into hobbies now that malls are less popular (and electronic shops like Radio Shack are basically disappeared).
------------
I don't remember what we used in college. I knew that I was more selective and understood the kinds of problems various OpAmps had back then. Also you're not really rich enough to invest into a private stockpile of chips, and instead just use whatever the labs are stocked with in college.
LM358 is the jellybean that I keep in my drawer today. If you're curious. Old habits die hard though, I still think 741 as the jellybean even though it really is obsolete today.
https://www.intel.com/content/dam/www/public/us/en/documents...
If you are already putting buck converters everywhere, it makes sense to raise the voltage coming from the PSU, because you can reduce the current supplied by the PSU. Only the power supply will be running at 12V, almost everything else will be running at a lower voltage. (Would you rather supply 120W at 12V/10A or 5V/24A?)
Lets say you make a design that has +/- 1V error. On the 5V circuit, that's 20% error. But on the +/-12V circuit, it is only 4.1% error. (Going much better than 5% or 1% error is nonsensical. Most resistors you'll buy are 5% error, and capacitors are maybe +/- 20% error). So you can see, +/-1V error is acceptable on +/-12V circuits, but maybe unacceptable on 5V circuits.
If precision is important, there are 1% or 0.1% resistors available, as well as matched-resistors (which have say 1% error, but all the resistors have the same degree of error, so your "proportions" remain consistent. You can manually-match resistors together with an accurate ohm-meter as well)
As such, +/- 12V is simply easier to make "precise" electronics on. Of course, the downside is that you've got leakage all over the place (more voltage means more power-draw)
---------
But now its 2021. Most parts, even extremely cheap "Jellybean" parts like the LM358, have low errors and low-bias. And USB's popularity as a 5V delivery mechanism has grown, everyone has a USB plug somewhere to use as the basis of electricity experiments.
So while electronic engineers way back played with +/-12V, today the assumption is that you play with 0V to 5V.
It also has speed and power advantages.
I think this release is excellent news on many levels.
That said you can read some interesting details here: https://en.wikichip.org/wiki/7_nm_lithography_process
It seems like nobody is talking about this, could anyone shine some light?
And immediately, we see the problem about dropping to 10nm: that's literally smaller than the distance that photons vibrate on their way to the final target.
And yeah, 10nm and 7nm is a marketing term, but that doesn't change the fact that these processes are all smaller than the wavelength of light.
-------
So there are two ways to get around this problem.
1. Use smaller light: "Extreme UV" is even smaller than normal UV at 13.5nm. Kind of the obvious solution, but higher energy and changes the chemistry slightly, since the light is a different color. Things are getting mighty close to literal "X-Ray Lasers" as they are, so the power requirements are getting quite substantial.
2. Multipatterning -- Instead of developing the entire thing in one shot, do it in multiple shots, and "carefully line up" the chips between different shots. As difficult as it sounds, its been done before at 40nm and other processes. (https://en.wikipedia.org/wiki/Multiple_patterning#EUV_Multip...)
3. Do both at the same time to reach 5nm, 4nm, or 3nm. Either way, 10nm and 7nm is the point where the various companies had to decide to do #1 first or #2 first. Either way, your company needs to learn to do both in the long term. TSMC and Samsung went with #1 EUV, and I think Intel though that #2 multi-patterning would be easier.
And the rest is history. Seems like EUV was easier after all, and TSMC / Samsung's bets paid off.
Mind you, I barely know any of the stuff I'm talking about. I'm not a physicist or chemist. But the above is my general understanding of the issues. I'm sure Intel had their reasons to believe why multipatterning would be easier. Maybe it was easier, but other company issues drove away engineers and something unrelated caused Intel to fall behind.
It seems like the diffraction pattern of the "long" wave laser would give you exactly what you want on the chip... and if you are putting hundreds of chips on a single wafer, it seems like you might not even need to worry too much about ringing at the edges.
The only ways to defeat diffraction are reducing it with shorter wavelengths and compensating it with multiple exposures with different patterns in which light and dark fringes compensate each other: exactly the two general approaches (EUV and multipatterning) taken by the semiconductor industry.
In the far field with a narrow bandwidth coherent light source (i.e., a laser), the projected image should be the FT of the aperture. That limiting condition is sometimes known as Fraunhofer diffraction, and generalizes to arbitrary apertures (not just a single hole).
Consider a narrow-band laser incident upon a diffraction grating for example. It produces a single point (well, two or three points, mirrored across the grating), not the uniform smudge that you'd expect by naively adding up the diffraction patterns of a bunch of slits. You should actually try this experiment for yourself!
The only trick is that you need a collimated laser and you need it to illuminate the entire inverse-FT-chip-grating aperture at once.
This deck has some nice examples of multi-hole apertures on slides 20 and 22: https://www.brown.edu/research/labs/mittleman/sites/brown.ed...
Making a device at a specific technology node (e.g. 14nm, 10nm, 7nm) isn't just about the lithography, although litho is crucial too. In effect, lithography is what allows you to "draw" patterns onto a wafer, but then you still need to do various things to that patterned wafer (deposition, etching, polishing, cleaning, etc.). Going from "we have litho machines capable of X nm spacing" to "we can manufacture a CPU on this node at scale with good yield" requires a huge amount of low-level design to figure out transistor sizings, spacings, and then how to actually manufacture the designed transistors and gates using the steps listed above.
Could we simplify this roughly into "ASML makes the machines to shine light at the right nm, foundries makes the Silicon and packages it into an useful device, architecture designers give you the layout to etch" if my mom were to ask?
EDIT: More seriously though: https://www.youtube.com/watch?v=NGFhc8R_uO4
> Its pretty simple. We zap sand with lightning until it starts thinking for us.
I think I kept working the joke in my brain, and it was too brutally simple at first. And then I worked it over, and now it reads too complicated. Ah well.
I think I was overthinking the joke.
Here's a neat video where they use an electron microscope to actually compare the transistor sizes for Intel 14nm and AMD 7nm: https://www.youtube.com/watch?v=1kQUXpZpLXI
As to why it's not just a matter of buying a bunch of ASML lithography machines and plugging them in: In addition to what the other replies have noted, there is so much complexity and precision required in a fab. Consider all of the challenges that would be involved with starting with a bunch of industrial robots, and trying to build a fully automated assembly line that manufactures cars. Then scale precision requirements up by many orders of magnitude.
I mean, when you put it that way, it's genuinely astounding that they managed to find themselves technically so far behind, and so quickly. This was not a minor mismanagement fuck up, this is complete incompetence. Things clearly need to be burnt to the ground, they need a lot of churn.
If they can get back to being engineering-driven rather than financially driven, it will all fall into place.
I think this is often stated as a cause, but more usually a symptom.
Isn’t the situation you described precisely where we would expect stagnation to be very likely?
I'm generally happy to blame MBAs for quite a bit. But isn't this also part of the nature of monopolies? A lack of competition means that in the short term, strategy doesn't matter at all. There was a period where Intel execs could have run the company via astrological chart or coin flip and done just as well for themselves. Managerialism and other MBA ideologies make that worse for sure.
I'd like to think I'm a good leader. Probably not the best one - there are plenty of folks I look up to. Probably not the worst one, but that's not for me to decide. In any case, as middle management, [0] I see three major components to my role:
1. Hire and retain really talented people who are good at what we do;
2. Decide what we'll work on now, what we'll do later, and what we won't do; [1] and
3. Sponsor my team and provide air cover for #2
I generally have a pretty good idea of _what_ to do, but not always exactly. I used to know _how_ to do it, but I'm probably not the best at that today. My goal is to hire or develop a team that can put a fine point on the _what_, or challenge me when I'm wrong, and to own the _how_.
In terms of sponsorship and air cover, it's taking 1 and 2 above, and keeping it on the rails with minimal disruption. I see good middle management as needing to be very team-focused, caring for the health and composition of the team above all. Sometimes that means absorbing a lot coming from above, and sometimes it means being really pushy upwards to get what we need to be successful. But the team should be free to execute if we're doing our jobs well, and their end of the bargain is to deliver what we agreed to do.
This is not the philosophy of all middle managers, and I'm sure there are other models that work. This is the one I can speak to.
--
[0]: "Middle management" defined as someone who manages other managers but is not C-level.
[1]: This doesn't mean I cook this shit up. I'm listening to the team I manage, as well as the priorities and directives coming from above, and applying my filters. Gotta figure they hired me because I know generally what to do in my space, and the team agreed to work with me because I have the ability to point them toward the critical path and push the debris out of the way before they run into it. Without management, there are either too many ideas to harmonize, or you get lucky and the team converges on the right one. I think good middle management helps us converge on the right ideas faster.
It doesn't make you experienced, or a leader. But it does give you the essential skills to run a business, that are so diverse that they are hard to acquire organically.
A good MBA is perhaps like a good coding Boot Camp. It will basically teach you enough that you can survive 3 months of a junior dev job, and if you're smart, bootstrap yourself up from that. But if you staff a company with no-one but people from such Boot Camps with no other experience, who will just then resonate against each other, that won't go anywhere.
More often than not it teaches them the arrogance to believe they can do everything on their own and that they know the law because they have an MBA. I'm currently closely observing a sequoia invested startup run by a bunch of TOP 3 MBA program invested people.
Questionable legal decision, terrible HR, abusive bullying management. It's all a big circle jerk.
Let's not kid ourselves. The real thing MBAs learn in school is how to network and score their next company they will destroy.
My wife is doing an MBA right now, on a top-100 uni globally; so not Harvard but good. I am genuinely impressed at the breadth and detail they go into.
As I say, these don’t turn you into experts or experienced business people, but from my perspective, very valuable.
Equally, running a large corporation without people who understand so many complex business dynamics (and crucially, you'll need some who appreciate how they interact) seems crazy-suboptimal to me.
A good example is Toyota versus the US Big 3 automakers. GM basically invented the modern American management the MBA came out of. Toyota was run by engineers. Toyota came out of the post WWII decimation of Japan to become the world's largest automaker, kicking the asses of US automakers in terms of cost and quality.
Toyota was happy to teach how they did it, and even set up a joint venture with GM, the NUMMI plant. The plant worked fine, but GM was unable to adopt the lessons. This American Life tells the tail well: https://www.thisamericanlife.org/561/nummi-2015
I think this is directly related to the MBA ideology, an important part of which is that you can take privileged people, give them a short general education covering a number of disciplines, and then make them part of the managerial caste that controls everybody else. A core notion of the MBA is that management is a universal skill. This is both historically novel and distinct from elements of the Toyota Way, which strongly honor the people doing the actual work and their accumulated knowledge.
MBA is a good starting place to understand complex business problems. It won't make you a good engineer. And yet even engineering companies have supply chains to manage, contracts to negotiate, labour force to grow and keep motivated, finance deals to secure, multi-year strategic planning etc. I know for a fact engineers are, by and large, not good at that.
Doing an MBA won't magically make you good at that, but at least gives you a head start thinking about those problems and standard industry solutions (if you take more than arrogance and networking out of it). Similar to how an engineering/science degree doesn't tell you how to build a Tesla, but gives you the basics to bootstrap yourself from that point. Even if many a Stanford engineering grad thinks they'll build the next Decacorn because they are so great.
I really can't see what's controversial about it.
If you really can't see what's controversial about giving power to a bunch of young people with a smattering of training and a very specific ideology, you haven't been looking at the harm caused. Maybe start there.
It reminds me of a friend's complaint about having done debate in high school. He says it took him years to unlearn the habits it gave him. Including the notion that whatever take first occurred to him in a discussion is one that must be defended at all costs. Or treating all discussions as gladiatorial to-the-death combat.
Learning that sort of faux certitude is certainly helpful in persuading people. So I get why people invest so much in learning it. But it seems to be to be double edged.
They teach you about those topics. But it's not clear how much they teach actual skills beyond finance and presentation, in that there's no actual business to run. One can read about negotiation, for example, but actual negotiation skills come from negotiating deals.
A friend of mine went back for a high-end MBA after a few years working for large companies as a (non-software) engineer. His basic take was that the curriculum only had one thing that he couldn't have figured out just by thinking about the topic for a bit. (It was the principle of comparative advantage, which is a bit of a mind-bender.) He says what he really got out of it was strong relationships with a network of people that were soon highly placed in a bunch of major companies.
Again, skills are not replacement for experience, or even good ideas. They won't make you a pro in those things either.
It's a bit like saying a CS degree is useless, because it's impractical, and with programming, you learn by doing anyway. And you can't build a business out of a CS degree alone. Undoubtedly so, and yet a CS degree shows you how large the computer world is, so to speak.
And I'm not saying an MBA is useless. As I said above, it's very useful for establishing a network. It's also an entry ticket into a bunch of high-caste jobs and a thorough indoctrination into the ideology of managerialism. The starting salary for MBA is proof of its usefulness to their holders.
I agree there's some analogy with a CS degree. The main difference being that nobody puts a fresh CS grad in charge of anything important, because we understand they don't know much yet. An MBA, on the other hand, puts people in positions of significant influence despite them knowing little or nothing about the actual work. It's as if somebody went through the book-learning part of a medical degree and then was turned loose to practice on real patients.
What can I say, this is not consistent with my experience. Perhaps it says more about the motivations and attitude of the people attending.
Perhaps a core confusion is thus: running a technology-based business requires, well, technological and business acumen. If you try to supplement one with the other, things go wrong, regardless of the direction. When I hear Toyota is ran by engineers, I'm sure the engineers don't negotiate finance deals, manage HR, career development, accounting etc. These boring things are important. Equally, when I hear about Boeing or IBM becoming corporate quagmires, whose core product is developed without a proper engineering input, I'm not suggesting this can work. My personal experience of companies ran by technical folk was, as said, awful.
The frustration about MBAs I hear in this thread seems to be when they encroach on technical territories they don't know about, and I agree this cannot work.
I think where MBAs work best is where you have mid-career people, already with experience, using this to boost their skills in diverse areas (alongside all other benefits like networking etc.), and I maintain it is great at that. If you take someone who doesn't know anything about anything, give them an MBA, then a random start-up to run, and it all goes wrong, I wouldn't blame the MBA.
You keep being sure about Toyota without any knowledge. That's very MBA of you, but I don't think it advances the discussion. The executives who turned it into a global power were engineers. And that engineering orientation was vital to the creation of the Toyota Way. Deeply influenced by W Edwards Deming, an engineer and statistician.
Nobody is denying that "boring" things like finance are important. Indeed, I'm saying the opposite. My point is that giving somebody a quick education that covers many "boring" things in a shallow fashion is not a great way to create executives. It is a great way, though, to create a glib executive caste that is good at extracting cash. Which is why executive comp has risen massively over the decades even though few would argue that executives are massively more competent.
You might be right that the MBA is better for people who have already developed actual expertise. But the average age of Harvard MBA applicants is 27, and only 12% are over 30. So if you're correct, then we can at least agree that most MBA grads shouldn't be in charge of anything.
As an occasional manager, there's some value to the work that some managers do. But managerialism itself often extractive (that is, part of its function is to justify siphoning a lot of cash into the pockets of managers and execs) and value-destroying (both through optimizing for value-negative things and empowering bad decisions).
A good example is what private equity does to companies: https://www.nytimes.com/2009/10/05/business/economy/05simmon...
Or the way MBA ideology tends to look at everything through a quarter-by-quarter financial lens, leading to underinvestment in R&D, staff, etc.
The pattern is typically something like:
1) Management awards themselves a bunch of stock and sets annual cash bonus targets tied to short term financial metrics.
2) Management then focuses entirely on optimizing for short term metric attainment (e.g., quarterly EPS) and current stock price, by minimizing investment in R&D and new products, reducing worker headcount, engaging in financial engineering like stock buybacks, etc.
3) Management extracts as much value for themselves as possible through the above and then leaves before it all falls apart, to repeat the process somewhere else.
MBAs might be over-represented within this group, but this isn't something taught in MBA curricula, not all MBAs are engaged in this abhorrent behavior, and non-MBAs do it as well.
This conclusion is extremely premature. Intel has greater market share now than it did in the Athlon 64 days.
Intel pulled in revenue of $77.9 billion last year. I don't think total collapse is a realistic outcome.
MBA or not, technologist or not, you can still be good or bad at what you do. Smart or stupid. I think simply looking at a mismanaged tech company and saying “ah too many MBAs in charge” is missing the point.
My guess is, if you look at the well-ran large companies, these will also have an MBA-rich management core.
I worked at a mid-sized company completely driven by technological people, with no one having a solid business background. It wasn’t good.
Is it though?
Or is it just the hardware of equivalent of the risky move of throwing it all away and rewriting from scratch for version 2? Intel have just incrementally improved their design, whereas AMD threw theirs away and almost disappeared in the process. In a parallel dimension, Intel could be exactly where they are now, with no AMD.
But... I agree that there was/is clearly a huge amount of incompetence letting non-engineers run an engineering company.
Large corps do gather intelligence on the competition. Even small ones do.
For example, when a new pizza place opened in town, the one I worked for when in College, immediately ordered one to see what they were like, pricing, taste, etc.
Large corps go way beyond this, and information does leak. Bribes, promised promotions and roles at the new company, theft of code, processes, trade secrets.
You do not have to use trade secrets, IP, or processes in house, and can silo that info from engineers / the main company. Yet having an awareness of these things, gives you an idea of how much to spend on research.
I wonder, did Intel not do any of this? Did they do so, but AMD security/counter espionage worked to provide a fake picture?
I wonder how large corps operate in this realm, these days.
A MBA Story.
Once upon a time, a MBA joined a company, worked his way from Finance Director to C grade. He could probably be described as pure breed MBA. He made friends with other rare MBA in the company and improve their career path along the way. Why? Because that is what you do, MBA always help another MBA. Somewhere along the line, the company's first non-technical / engineer MBA CEO was born.
As the company grow there are more MBAs, but not all MBAs are bad, you just need engineers and product people around these MBA. But some of these engineers or product genius [1] are pain in the ass to manage, and since many of them dont have MBA, they dont understand MBA speak and they get driven out of company decision making forum. So what happened was these product engineers got rotted out over the years even when they are C grade and out very publicly.
With leading edge tech company there is a very long lead time in their product R&D roadmap, so they were fine for the first few years, and continue to make record breaking results. Then MBA move his way up to the board and become chairman just in the time before his MBA CEO friend retired. Which means MBA Chairman gets to pick a new CEO again, guess which CEO would a MBA pick? Another MBA! Since MBA always help another MBA, new MBA CEO promoted more MBA to various rolls within the company. And MBA cult now officially dictate the company decision making.
Everything was great for a few more years, the market was growing, they have a monopoly and most important of all, they are doing great because they are MBA. At least that is what those MBA thought. Until one day cracks started appearing, MBA CEO thought it was all fine, but with time external pressure were piling in and the whole industry moved on. The market looked very differently to what company were 10 years ago. And somehow the MBA CEO already had an on going internal investigation which shows he had an inappropriate relationship and was let go. Which is the MBA speak of CEO being fired. And a recently arrived CFO was named as interim CEO while MBA Chairman and Board continue their CEO search. They could hire other MBA from outside the company, but since the leading edge tech company are a gigantic company with so many different specific domain knowledge and culture. The suitable choice outside the company are virtually nil. MBA Chairman looked at CFO, despite not being his first choice and is fairly new to the tech company, but he has a MBA! And the third MBA CEO was born.
The market is moving at a much faster pace than before. To give credit to the new MBA CEO he was trying very hard to steer the ship to a new path and direction. The tech company are already behind. These new direction and requirement are completely new, MBA middle management have no idea what they are doing and cant turn the ship fast enough. And while they are still turning, the industry has moved further along. I guess this is an example of MBA being killed by MBA.
Finally shareholder were so fed up, some of them asked the same question;
>how they mismanage the company to the point that they squandered a two-decade long global monopoly on computing, in an era where computing absolutely dominated global economic development.
There are very little choice of what shareholder could do. They need those product engineers back. And a new CEO. But the perfect candidate had a fallout with MBA Chairman a decade ago, the only option was to rally enough support to get rid of him. In the end, MBA Charimand retired, a new Board Chairman was named, and those product engineers are back to save the company.
As far as the MBA story goes, this is the end. We dont know what happen to the leading edge company, I guess that part is To be continued.
Yet those chips are apparently made on either TSMC's or GloFo's 12nm nodes (which one I can't find). Whilst Intel's 14nm is still superior to either TSMC/GloFo 12nm by all accounts, they're definitely the same ballpark. Competition for Intel even looks rough for cheaper high volume chips. You'd need to consider that Intel still needs some time to adapt to being a third party foundry too.
To me, it seems like the same story is playing out again.
But more importantly, Intel actually has way more foundry expertise than Global did/does. In the case of AMD selling off Global it was more like cutting loose an under performing part of the company rather than empowering the a significant asset of the company.
I'm not sure that is a worthwhile segment for them to try to compete in though. It's high mix, has a ton of competition, and doesn't necessarily leverage their fab strength.
Key question will be how quickly Intel will shift to the next architecture, Sapphire Rapids. Will this release be like the consumer/desktop Rocket Lake? E.g. just a placeholder to essentially volume test the 10nm fabrication for datacenter. Probably at least a year out at this point since Ice Lake SP was supposed to be originally released in 2H2020.
Alder Lake is meant to be a consumer part contemporary with Sapphire Rapids, which is server only. They're likely based on the same (performance) core, with Adler Lake additionally having low-power cores.
Last I heard the expectation was still that these new parts would enter the market at the end of this year.
M1 is proof that it can be done, however you can absolutely make a bad CPU for a good ISA so I wouldn't take it for granted.
Intel is already behind AMD -- they have no product segment where they are absolutely superior. The means AMD is setting the market pace.
On top of this Apple is switching to ARM designed CPUs. This also looks to be a vote of no confidence in Intel.
The consensus seems to be that Intel who have their own fabs -- never really nailed anything under 14nm and are now being outcompeted.
There are some who would argue this claim, but I think it's at least a defensible one.
Still, availability is an important factor that isn't captured by benchmarking. AMD has had CPU inventory trouble in the low-end laptop segment and high-end desktop segment alike.
> The consensus seems to be that Intel who have their own fabs -- never really nailed anything under 14nm and are now being outcompeted.
Intel has done well with 10nm laptop CPUs. They were just very late to the party. Desktop and server timelines have been quite a bit worse. I agree Intel did not nail 10nm, but they're definitely hanging in there. It's one process node at the cusp of transition to EUV, so some of the defeatism around Intel may be overzealous if we keep in mind that 7nm process development has been somewhat parallel to 10nm because of the difference in the lithographic technology.
Intel is still biggest player in the game, because even though they are stuck at 14nm, AMD isn't able to manufacture enough to take bigger chunks of the market. Apple won't sell it to PC/Datacenter space, rest are still niche.
I think this isnt quite fair, their laptop 10nm chips have been shipping in volume since last year, and their server chips were released today, with 200k+ units already shipped (according to Anandtech). The only line left on 14nm is socketed Desktop processors, which is a relatively small market compared to laptops and servers.
To put it in numbers alone, look at this benchmark. Flagship vs Flagship: https://www.cpubenchmark.net/compare/Intel-i9-11900K-vs-AMD-...
They are, in a certain sense, suffering from their own success in that their competitors have basically been nonexistant up until Zen came about (and even then only until Zen 3 have Intel truly been knocked off their single thread perch). This has led to them getting cagey, and a bit ridiculous in the sense that they are not only backporting new designs to old processes but also pumping them up to genuinely ridiculous power budgets. With Apple, AMD, and TSMC they have basically been caught with their trousers down by younger and leaner companies.
Ultimately this is where Intel need good leadership. The mba solution is to just give up and do something else (e.g. spin off the fabs), but I think they should have the confidence (as far as I can tell this is what they are doing) to rise to the technical challenge - they will probably never have a run like they did from Nehalem to shortly before now, but throwing in the towel means that the probability is zero.
Intel have been in situations like this before, e.g. When Itanium was clearly doomed and AMD were doing well (amd64), they came back with new processors and basically ran away to the bank for years - AMD's server market share is still pitiful compared to Intel (10% at most), for example.
Which one? I dont believe they missed a die shrink, it just took a long time. Intel 14nm came out in 2014 with their Broadwell Processors, and the next node, 10nm came out in 2019 (technically 2018, but very few units shipped that year).
From the 2015 10-K [0]:
"We expect to lengthen the amount of time we will utilize our 14nm and our next-generation 10nm process technologies, further optimizing our products and process technologies while meeting the yearly market cadence for product introductions."
Spoiler alert: In the five years after shelving the tick-tock model, Intel also missed the yearly market cadence for product introductions.
[0] https://www.sec.gov/Archives/edgar/data/50863/00000508631600...
It's hard to see what they can do. AMD is winning on every measure I can see on x86, and M1/Ampere feel like a huge curveball thrown at them. I don't even think they could switch to using TSMC fabs to help them shrink the process size as all the capacity is booked up for years (especially by Apple).
I think also the dynamics have changed in the server market in the last 10 years or so. You now have a much more condensed market with this big cloud operators buying millions of CPUs and having much more leverage over them before.
I would be surpised if by the end of the 2020s if ARM wasn't standard in nearly all server workloads. Most software can run on it unmodified now - databases all work great on ARM, and most applications are interpreted or run on a VM, so it is pretty easy to move to.
Also, the first 10nm "Ice Lake" mobile CPUs were not really an improvement over the by then many times refined 14nm chips "Comet Lake". It's been a faecal pageant.
So what you're seeing isn't really anti-Intel, it's probably often more like bitter disappointment that they haven't done better. Though I'm sure there's a tiny bit of fanboy-ism for & against Intel.
There's definitely some of that pro-AMD fanboy sentiment in the gaming community where people build their own rigs: AMD chips are massively cheaper than a comparable Intel chip.
its back to where everyone design its own chip for their own product but don't need a fab 'cause of foundry like TSMC and Samsung.
For instance, you can now get an i7-10700K (which is roughly equivalent in single thread and better in multi thread) for cheaper than a R5 5600X.
My cherry-pick is where the AMD chip is 30% more expensive, but multi-threaded performance is 100% better in this example: https://www.cpubenchmark.net/compare/Intel-i9-11900K-vs-AMD-...
Edit: picking individual processors to compare (especially low volume ones) is often not useful when talking about how well a company is competing in the market.
I think that summarizes it pretty well in that one graph.
Many people, I'd say especially enthusiasts, were quite happy when AMD was able to compete on a performance/$ basis and then outright beat Intel.
Of course, now the tables have turned and AMD is able to extract that price premium while Intel cut prices. Who knows how long this will last, but Intel is still the 800 lb gorilla in terms of capacity, engineering talent, and revenue. I don't think we've heard the last from them.
"As impressive as the new Xeon 8380 is from a generational and technical stand-point, what really matters at the end of the day is how it fares up to the competition. I’ll be blunt here; nobody really expected the new ICL-SP parts to beat AMD or the new Arm competition – and it didn’t. The competitive gap had been so gigantic, with silly scenarios such as where a competing 1-socket systems would outperform Intel’s 2-socket solutions. Ice Lake SP gets rid of those more embarrassing situations, and narrows the performance gap significantly, however the gap still remains, and is still undeniable."
This sounds about right for a company fraught with so many process problems lately: Play catch up for a while and hope you experience fewer in the future to continue to narrow the gap.
"Narrow the gap significantly" sounds like good technical progress for Intel. But the business message isn't wonderful.
1. https://www.anandtech.com/show/16594/intel-3rd-gen-xeon-scal...
"At the end of the day, Ice Lake SP is a success. Performance is up, and performance per watt is up. I'm sure if we were able to test Intel's acceleration enhancements more thoroughly, we would be able to corroborate some of the results and hype that Intel wants to generate around its product. But even as a success, it’s not a traditional competitive success. The generational improvements are there and they are large, and as long as Intel is the market share leader, this should translate into upgraded systems and deployments throughout the enterprise industry. Intel is still in a tough competitive situation overall with the high quality the rest of the market is enabling."
"As impressive as the new Xeon 8380 is from a generational and technical stand-point, what really matters at the end of the day is how it fares up to the competition. I’ll be blunt here; nobody really expected the new ICL-SP parts to beat AMD or the new Arm competition – and it didn’t. "
That sounds "competetive enough" to me in the datacenter world, given the existing market lead Intel has.
AMD’s idle power consumption is a bigger issue for desktop, laptop, and HEDT.
I doubt this applies to HPC (the target market for this part) as they either schedule jobs closely or could, I imagine, shut them down. But I'm not in that space either so this is merely conjecture.
* I am sure there are corner cases where the load is uniform, but they are by definition few.
Some bigger companies have a lot of batch jobs that can run overnight and steal idle cycles, but you have to be gigantic before that's realistic. (My experience with writing gigantic batch jobs is that I just requisitioned the compute at "production quality" so I could work on them during the day, rather than waiting for them to run overnight. Not sure what other people did, and therefore not really sure how much runs overnight at big companies.)
Cloud providers have spot instances that could take up some of this slack, but I bet there is plenty of idle capacity precisely because the cost can't go to $0 because of electricity use. Or I could be completely wrong about workloads, maybe everyone has their web servers and CI systems running at 100% CPU all night. I've never seen it, though.
Based on the article they're targeting high performance compute, i.e. "application codes used in earth system modeling, financial services, manufacturing, as well as life and material science."
Not all workloads are CPU-bound. Cloud providers have many servers for which the CPUs are idle most of the time, because they're disk-bound, network-bound, other-server-bound, bursty, or similar. They're going to aim to minimize the idle time, but they can't eliminate it entirely given that they have customer-defined workloads.
Just a few servers could be kept idle, not off, to enable a sub-second start-up time for some new work.
Even purchased on the used market, a decent server could cost $4000; if you spread that cost across 4 years, that's $80ish per month just to pay for the server, which is probably around your power / rent costs.
The costs are more stark if you're buying new, where it is easy to sped $10-30k for a decent spec server.
Don't buy a server to leave it turned off, you're wasting money.
For example, the Ryzen 3000 desktop chips seemed to have the issue[0], but the same Zen 2 cores seem to have found some improvements in the Ryzen 4000 mobile chips[1].
I didn't want to just rely on Reddit forum comments, so I found this measure of the Ryzen 3600[2].
> When one thread is active, it sits at 12.8 W, but as we ramp up the cores, we get to 11.2 W per core. The non-core part of the processor, such as the IO chip, the DRAM channels and the PCIe lanes, even at idle still consume around 12-18 W in the system.
My interpretation was expect ~12 W or more idle consumption (just from the CPU package), but I'm not sure I understand it correctly.
I couldn't find the same information for Ryzen 4000 laptops, but the same APU is tested in a NUC, where the total system draw (at the wall) at idle was about 10-11 W, still nearly double that of a Core i7 U-series NUC[3], but certainly lower than that of just the CPU package in the Ryzen 3600.
Anecdotally, my 45W Ryzen 7 4800H laptop with 15.6" 1080p screen lasts about 4 hours on 80% of the 60Wh battery with 95% brightness, doing various non-intensive tasks. Though I don't know how well the battery holds up on complete non-use standby.
[0] https://old.reddit.com/r/AMDHelp/comments/cfm1xa/why_is_ryze...
[1] https://old.reddit.com/r/Amd/comments/haq4fg/the_idle_power_...
[2] https://www.anandtech.com/show/15787/amd-ryzen-5-3600-review...
[3] https://www.anandtech.com/show/16236/asrock-4x4-box4800u-ren...
I measured an Asus Mini PC PN50 with a Ryzen 4500U. The idle power usage was 8.5 Watt for the system. This with 32GB of memory and a SATA SSD installed. It would be nice if it was lower than this, but it isn't too bad. Interestingly the machine used 1.2 Watt while off after it wasn't on power, 0.5 Watt after starting up and shutting it down.
Recently noticed some people focussing on low power but powerful 24/7 home "servers". Systems that are on 24/7, but often idle. One system used around 4.5 Watt in idle. The "brick" / power adapter often uses too much power, even when everything is off.
Objectively, Ice Lake "isn't too bad" in comparison with Zen 3 either. "It would be nice if it" drew a 20% less power at load, but in practice you could use either product without trouble.
However, because each die is smaller, you can be more aggressive with your binning, potentially having a larger number of chips which you can bin into a higher frequency bucket.
Ultimately, it's a matter of trade-offs, and what's right in one place might be wrong in another (see, e.g., AMD manufacturing mobile APUs on monolithic dies versus all their desktop and server parts using chiplets).
We are seeing more of the industry moving to chiplets and similar solutions, notably Intel with their tiles.
Do you see getting the most of limited 7nm production capacity as "cheating"? The only fair comparison is in the resulting performance and TDP, and Intel remains seriously behind.
There's probably a few more gates still on the AMD side, but it's not the half again larger that you'd expect looking at area alone.
edit oops nevermind, I see my comment was also mysteriously transported from the dupe.
Publicly the problems have been lately, but the things that caused these problems have happened much further back.
I'm cautiously bullish on Intel. From what I gather, Intel is in a much better place internally. They have much better focus, there is less infighting, its more engineering then sales lead, they have some very good people and they are no longer complacent. It will however take years before this is becomes visible from the outside.
Given the demand for CPUs and the competitions inability to deliver, I think intel will do OK even if they are no ones first choice of CPU vendor, while they try to catch up.
And I can bet those prices have lots of room for special discount to clients. Since RAM and NAND Storage dominate the cost of server, the difference of Intel and AMD shrinks rapidly in the grand scheme of things, giving Intel a chance to fight. And there is something not mentioned enough, the importance of PCI-E 4.0 Support.
I wanted to rant about AMD, but I guess there is not much point. ARM is coming.
Only interesting if you care about price (dollars spent per inference)
For raw speed (no matter the price) the GPU wins
GPUs have far more bandwidth, but CPUs beat them in latency. Being able to AVX512 your L1 cached data for a memcpy will always be superior to passing data to the GPU.
With Ice Lake's 1MB L2 cache, pretty much all tasks smaller than 1MB will be superior in AVX512 rather than sending it to a GPU. Sorting 250,000 Float32 elements? Better to SIMD Bitonic sort / SIMD Mergepath (https://web.cs.ucdavis.edu/~amenta/f15/GPUmp.pdf) on your AVX512 rather than spend a 5us PCIe 4.0 traversal to the GPU.
It is better to keep the data hot in your L2 / L3 cache, rather than pipe it to a remote computer (even if the 16x PCIe 4.0 pipe is 32GB/s and the HBM2 RAM is high bandwidth once it gets there).
--------
But similarly: CPU SIMD can never compete against GPGPUs at what they do. GPUs have access to 8GBs @500GB/s VRAM on the low-end and 40GBs @1000GB/s on the high end (NVidia's A100). EDIT: Some responses have reminded me about the 80GB @ 2000GB/s models NVidia recently released.
CPUs barely scratch 200GB/s on the high end, since DDR4 is just slower than GPU-RAM. For any problem where data-bandwidth and parallelism is the bottleneck, that fits inside of GPU-VRAM (such as many-many sequences of large scale matrix multiplications), it will pretty much always be better to compute that sort of thing on a GPU.
-------
It should be noted that most CPUs do use DDR4 for good reason: they have an expectation to use more than the capacity of HBM2. GPUs (and a supercomputer CPU, like the A64FX) can rely upon the fact that their workloads are designed for distributed-compute and/or otherwise fit inside of 32GBs or 40GBs.
CPUs on the other hand, could very well be reading/writing to a 2TB in-memory database these days, even on a single-socket single-node system like AMD EPYC. That flexibility to have as much (or as little) RAM as the customer demands is a major advantage of traditional CPUs like EPYC or Xeon, which the A64FX cannot partake in.
If you need more than 32GBs of RAM, A64Fx is a no-go. That HBM2 is soldered directly onto the package and cannot expand.
Although an A100 still has twice the bandwidth and over 3x the GFLOPS without tensor cores and more than 6x with.
Yeah. And at 3200 Mbit/sec, that comes out to 200GB/s. (3200 MHz x 8-bytes (aka 64-bit) == 25GB/s. x8 channels == 200GB/s).
> where IIRC the biggest NVIDIA part is still at 6.
That's 6x *1024-bit* HBM2 channels. Total bandwidth is 2000GBps, or over 10x the speed of the "8x channel EPYC". Yeah, HBM2 is fat, extremely fat.
----------
*ONE* HBM2 channel offers over 300GBps bandwidth. And the A100 has *SIX* of them. Literally ONE HBM2 channel beats the speed of all 8x DDR4 EPYC memory channels working in parallel.
25600 is the channel rate in (EDIT) MB/sec of the stick of RAM. That's 25GB/s for a 3200 MHz DDR4 stick. x8 (for 8-channels working in parallel) is 200GB/s.
-----------
This has been measured in practice by Netflix: https://2019.eurobsdcon.org/slides/NUMA%20Optimizations%20in...
As you can see, Netflix's FreeBSD optimizations have allowed EPYC to reach 194GB/s measured performance (or just under the 200GB/s theoretical). And only with VERY careful NUMA-tuning and extreme optimizations were they able to get there.
BVH-tree traversals are done on the GPU now for a reason. GPUs are better at latency hiding and taking advantage of larger sets of bandwidth than CPUs. Yes, even on things like pointer-chasing through a BVH-tree for AABB bounds checking.
GPUs have pushed latency down and latency-hiding up to unimaginable figures. In terms of absolute latency, you're right, GPUs are still higher latency than CPUs. But in terms of "practical" effects (once accounting for latency hiding tricks on the GPU, such as 8x way occupancy (similar to hyperthreading), as well as some dedicated datastructures / programming tricks (largely taking advantage of the millions of rays processed in parallel per frame), it turns out that you can convert many latency-bound problems into bandwidth-constrained problems.
-----------
That's the funny thing about computer science. It turns out that with enough RAM and enough parallelism, you can convert ANY latency-bound problem into a bandwidth-bound problem. You just need enough cache to hold the results in the meantime, while you process other stuff in parallel.
Raytracing is an excellent example of this form of latency hiding. Bouncing a ray off of your global data-structure of objects involved traversing pointers down the BVH tree. A ton of linked-list like current_node = current_node->next like operations (depending on which current_node->child the ray hit).
From the perspective of any ray, it looks like its latency-bound. But from the perspective of processing 2.073 million rays across a 1920 x 1080 video game scene with realtime-raytracing enabled, its bandwidth bound.
> (which is what I think ajross keeps alluding to)
It seems like ajross is accusing me of underestimating CPU-bandwidth. At least, that's my interpretation of the discussion so far. As you've pointed out however, I'm overestimating it.
EDIT: But I'm overestimating it on both sides. A100 2000 TB/s is the "channel bandwidth" as well, as the CAS and RAS commands still need to go through the channel and get interpreted.
Whereas Nvidia's A100 has over 2000 GB/s of memory bandwidth. That's 10-fold better.
The two last apps I worked on have been GPU-only. The CPU process starts running and launches GPU work, and that's it, the GPU does all the work until the process exits.
There is no need to "pass data to the GPU" because data is never on CPU memory, so there is nothing to pass from there. All network and file I/O goes directly to the GPU.
Once all your software runs on the GPU, passing data to the CPU for some small task doesn't make much sense either.
However, the famous "Moana" scene for Disney-level productions is a 93GB (!!!!) scene statically, with another 131GBs (!!!) of animation data (trees blowing in the winds, waves moving on the shore, etc. etc.).
That's simply never going to fit on a 8GB, 40GB, or even 80GB high-end GPU. The only way to work with that kind of data is to think about how to split it up, and have the CPU store lots of the data, while the GPU processes pieces of the data in parallel.
https://www.render-blog.com/2020/10/03/gpu-motunui/
Which has been done before, mind you. But it should be noted that the discussion point for GPU-scale compute runs into practical RAM-capacity constraints today, even on movie-scale problems from 5 years ago (Moana was released in 2016, and had to be rendered on hardware years older than 2016).
Moana scene is here if you're curious: https://www.disneyanimation.com/resources/moana-island-scene...
----------
But yes, if your data fits within the 8GBs GPU (or you can afford a 40GB or 80GB VRAM GPU and your data fits in that), doing everything on the GPU is absolutely an option.
If you’re Disney, you can afford boxes with 10+ A100s with NVLink sharing the memory in a single 400+ GB pool. Unknown if that ends up being more economical than the equivalent CPU version, but it’s important to understand in order to evaluate the future of GPUs.
There is always a problem size that does not fit into memory.
Whether that memory is the GPU memory, or the CPU memory, doesn't really matter.
We have been solving this problem for 60 years already. It isn't rocket science.
---
The CPU doesn't have to do anything.
The GPU can map a file stored on hard disk to VRAM memory, do random access into it, process chunks of it, write the results into network sockets and send them over the network, etc.
The only thing the CPU has to do is launch a kernel:
int main(args...) {
main_kernel<<<...>>>(args...);
synchronize();
return 0;
}
and this is a relatively accurate depiction of how the "main function" of the two latests apps I've worked on look like: the GPU does everything.---
> However, the famous "Moana" scene for Disney-level productions is a 93GB (!!!!) scene statically, with another 131GBs [...] That's simply never going to fit on a high-end GPU.
LOL.
V100 with 32Gbs and 8x per rack gave you 256 Gb of VRAM addressable from any GPU in the rack.
A100 with 80GB and 16x per rack give you 1.3 TB of VRAM addressable from any GPU in the rack.
You can fit Moana in GPU VRAM in a now old DGX-2.
If you are willing to bet cash on Moana never fitting on a single GPU, I'd take you on that bet. Sounds like free money to me.
This person rendered the Moana scene on just 8GBs of GPU VRAM. It does this by rendering 6.7GB chunks at a time on the GPU, with the CPU keeping the RAM-heavy "big picture" in mind. (EDITED paragraph. First wording of this paragraph was poor).
------
Its not that these problems "cannot be solved", its that these problems "become grossly more complicated" when under RAM / VRAM constraints. They're still solvable, but now you have to do strange techniques.
------
With regards to a Ray-tracer, tracing the ray-of-light that's bouncing around could theoretically touch ANY of the 93GBs of static object data (which could have been shifted by any of the 131 GBs of animation data). That is to say: a ray that bounces off of a any leaf on any tree could bounce in any direction, hitting potentially any other geometry in the scene.
That pretty much forces you to keep the geometry in high-speed RAM, and not do an I/O cycle between each ray-bounce.
As a rough reminder of the target performance: Raytracers aim at ~30 million to 30-billion ray-bounces per second, depending on movie-grade vs video-game optimized. Either way, that level of performance is really only ever going to be solved by keeping all of the geometry data in RAM.
> A100 with 80GB and 16x per rack give you 1.3 TB of VRAM addressable from any GPU in the rack.
That doesn't mean it makes sense to traverse a BVH-tree across a relatively high-latency NVLink connection off-chip. I know GPUs have decent latency hiding but... that's a lot of latency to hide.
Again: your CPU-renderers can hit 10s of millions of rays per second. I'm not sure if you're gonna get something pragmatic by just dropping the entire geometry into distributed NVSwitch'd memory and hoping for the best.
Honestly, that's where the 8GB CPU+GPU team becomes interesting to me. A methodology for clearly separating the geometry and splitting up which local compute-devices are responsible for handling which rays is going to scale better than a naive dump reliant on remote-connections pretending to be RAM.
Video games hit Billions of rays/second. The promise of GPU-compute is on that order, and I just doubt that remote RAM accesses over NVLink will get you there.
> If you are willing to bet cash on Moana never fitting on a single GPU, I'd take you on that bet. Sounds like free money to me.
The issue is not Moana (or other movies from 2016), the issue are the movies that will be made in 2022 and into the future. Especially if they're near photorealistic like Marvel-movies or Star Wars.
----------
The other problem is: what's cheaper? A DGX-system could very well be faster than one CPU system. But would it be faster than a cluster of Ice-Lake Xeons with AVX512 each with the precise amount of RAM needed for the problem? (Ex: 512GBs in some hypothetical future movie?)
A team probably would be better: CPUs have expandable RAM, that's their biggest advantage. GPUs have fixed RAM. Slicing the problem up so that pieces of Raytracing fits on GPUs, while the other, "bulkier" bits fit on CPU DDR4 (or DDR5), would probably be the most cost-efficient way at solving the raytracing problem.
The GPU-Moana experiment showed that "collecting rays that bounce outside of RAM" is an efficient methodology. Slice the scene into 8GB chunks, process the rays that are within that chunk, and the collate the rays together to find where the rays go.
This is very interesting - do you have a link that explains how it works / is implemented?
https://www.nvidia.com/en-us/geforce/news/rtx-io-gpu-acceler...
https://www.amd.com/en/products/professional-graphics/radeon...
The GPU itself can send PCIe 4.0 messages out. So why not have the GPU make I/O requests on behalf of itself? Its a bit obscure, but this feature has been around for a number of years now. The idea is to remove the CPU and DDR4 from the loop entirely, because those just bottleneck / slowdown the GPU.
--------
From an absolute performance perspective, it seems good. But CPUs are really good and standardized at accessing I/O in very efficient ways. I'm personally of the opinion that blocking and/or event driven I/O from the CPU (with the full benefit of threads / OS-level concepts) would be easier to think about than high-performance GPU-code.
But still, its a neat concept, and it seems like there's a big demand for it (see PS5 / XBox Series X).
I think the Radeon + Premier Pro documentation makes it clear how it works: https://www.amd.com/system/files/documents/radeon-pro-ssg-pr...
As you can see, the GPU is attached to the x16 slot, and the 4x NVMe SSDs are attached to the GPU. When the CPU wants to store data on the SSD, it communicates first to the GPU, which then pass-throughs the data to the four SSDs.
That's the simpler example.
--------------
In NVidia's case, they're building on top of GPUDirect Storage (https://developer.nvidia.com/blog/gpudirect-storage/), which seems to be based on enterprise technology where PCIe switches were used.
NVidia's GPUs would command the PCIe switch to grab data, without having the PCIe switch send data to the CPU (which would most likely be dropped in DDR4, or maybe L3 in an optimized situation).
That's a pretty big "if" and often not a workable.
You just write normal C++, and it works.
Another aspect that seems like a hidden assumption in CPU-GPU discussions is that you have the time-energy-expertise budget to (re)build your application to fit GPUs.
40TBs+ -- Storage-only solutions. "External Tape Merge sort algorithm", "Sequential Table Scan", etc. etc. (SSDs or even Hard drives if you go big enough)
4TB to 40TBs -- Multi-socket DDR4 RAM is king (8-way Ice Lake Xeon Scalable Platinum will probably reach 40TBs). Single-node distributed memory with NUMA / UPI to scale.
1TB to 4TB -- Single Socket DDR4 RAM (EPYC, even if at 4x NUMA. Or Single-node Ice Lake).
80GB to 1TB -- DGX / NVlink distributed memory A100 ganging up HBM2 together. GPU-distributed RAM is king.
256MBs to 80GBs -- HBM2 / GDDR6 Graphics RAM is king (80GB A100 2TB/s).
1.5MBs to 256MBs -- L3 cache is king (8x32MBs EPYC L3 cache, or POWER9 110MB+ L3 cache unified)
128kB to 1.5MBs -- L2 cache is king (1.25MB Ice Lake Xeons L2, this article)
1kB to 128kB -- L1 cache is king. (128kB L1 cache on Apple M1). Note: "GPU __Shared__" is a close analog to L1 and competes against it, but is shared between 32 to 256 GPU threads, so its not an apples-to-apples comparison.
1kB and below -- The realm of register-space solutions. (See 64-bit chess engine bitboards and the like). Almost fully CPU-constrained / GPU-constrained programming. 256x 32-bit GPU registers per GPU-thread / SIMD thread. CPUs have fewer nominal registers, but many "out of order" buffers or "reorder buffers" that practically count as register storage in a practical / pragmatic sense. CPUs just use their "real registers" as a mechanism to automatically discover parallelism in otherwise single-thread written code.
------------
As you can see: GPUs win in some categories, but CPUs win in others. And these numbers change every few months as a new CPU and/or GPU comes out. And at the lowest levels: CPUs and GPUs cannot be compared due to fundamental differences in architecture.
For example: GPU __shared__ memory has gather/scatter capabilities (the NVidia PTX instructions / AMD GCN instructions permute vs bpermute), while CPUs traditionally only accelerate gather capabilities (pshufb), and leave vgather/vscatter instructions to the L1 cache instead. GPUs have 32x ports to __shared__, so every one of the 32-threads in a wave-front can read/write every single clock-tick (as long as all 32 they are on different ports/alignment, or you have a special one-to-all broadcast). CPUs only have 2 or 4 ports, so vscatter and vgather operate slowly, as if a single thread were reading/writing each of the memory locations.
But CPU L1 cache has store-forwarding, MESI + cache coherence, and other acceleration features that GPUs don't have.
GPUs are therefore more efficient at sharing data within workgroups of ~256 threads, but CPUs are more efficient at sharing data between cores, or even among out-of-die NUMA solutions, thanks to robust MESI messaging.
SIMD throughput?
https://www.phoronix.com/scan.php?page=news_item&px=Linus-To...
The problem is that for the workloads that need specialization, you sometimes really need it. You could also delete the vectorized AES units in your Intel machines too and the general purpose performance wouldn't be affected much, but cryptographic performance specifically would tank, and it turns out, that matters a lot in aggregate for many people.
Ultimately there are literally dozens of specialized inactive units on any CPU at any given time that could be "better spent on general purpose units" (which also isn't necessarily true if other architectural choices prevent those units from being utilized effectively). People just like complaining about AVX-512 because it's easily digestible water cooler chat they read about on a blog.
It impedes latency-sensitive applications due to forced down-clocking taking place for AVX-intensive instruction streams. The other side effect is that it quickly produces extreme heat that not only forces further down-clocking of the core but also taints neighbouring cores with dissipated heat and prevents them from going into turbo.
Not everything is about performance in laptops.
"As impressive as the new Xeon 8380 is from a generational and technical stand-point, what really matters at the end of the day is how it fares up to the competition. I’ll be blunt here; nobody really expected the new ICL-SP parts to beat AMD or the new Arm competition – and it didn’t. The competitive gap had been so gigantic, with silly scenarios such as where a competing 1-socket systems would outperform Intel’s 2-socket solutions. Ice Lake SP gets rid of those more embarrassing situations, and narrows the performance gap significantly, however the gap still remains, and is still undeniable."
edit: it also depends on how many dies the mask (reticle) has on it. Intel uses one die reticles, so i. theory their real wafers have no situation in which they have partial dies at the edge.
There is no reason not too so this and the edges are also often used for calibration and (potentially) destructive testing.
The only area of the wafer that will not be exposed is the notch, it’s always on one side of the circle and it’s used for moving the wafer around, this is why you often see wafers with one of the sides cut off giving it the flat tire shape.
The litho step could in theory be optimized by skipping incomplete fields at the edges, but the reduction in exposure time would be relatively small, especially for smaller designs that fit multiple chips within a single image field. I imagine it would als introduce yield risk because of things like uneven wafer stress & temperature, higher variability in stage move time when stepping edge fields vs center fields, etc.
SLI really shines on HEDT platforms, and this is probably the last non-multi-chip quasi-HEDT CPU for a while with this kind of IO.
(Yes, I know SLI is 'dead' with the latest generation of GPUs)