Apple M1 Ultra Meanings and Consequences
mondaynote.com
mondaynote.com
I find it hard to believe that this was a last-minute decision. Rather, I think this pattern of a new core design (using a new process if there is one) releases first for the smallest devices (iPhones), and the gradually moves its way up the lineup all the way up the Ultra before the cycle repeats with a new generation is likely Apple's new strategy going forwards.
My understanding is that this is pretty much what Intel and AMD do too (releasing their smaller dies on new processes first) and that this is a general strategy for dealing with poorer yield number on new process nodes. The idea that Apple would ever have considered releasing their biggest chip as the first chip on a new node seems far-fetched to me.
Incidentally this means that Apple will no longer have a node advantage once Zen4 launches - both Zen4 and A15 will be on the same node, so we can make direct comparisons without people insisting that Apple's performance is solely due to node advantage/etc.
But yeah, that does go to show that 3nm is slow to launch in general - Apple would not willingly give up their node lead like this if there were anything ready for an upgrade. I don't think it's actually falling behind in the sense that it was delayed, but it seems even TSMC is feeling the heat and slowing down their node cadence a bit.
Also, as far as this:
> Second, the recourse to two M1 Max chips fused into a M1 Ultra means TSMC’s 5 nm process has reached its upper limit.
There is still Mac Pro to come, and presumably Apple would want an actual Pro product to offer something over the Studio besides expansion.
marcan42 thinks it's not likely that quad-die Mac Pros are coming based on the internal architecture (there's only IRQ facilities for connecting 2 dies) but that still doesn't rule out the possibility of a larger die that is then connected in pairs.
Also bigger/better 5nm stuff will almost certainly be coming with A15 on N5P later this year, so this isn't even "the best TSMC 5nm has to offer" in that light either.
I said quad-Jade Mac anythings aren't coming because that die is only designed to go in pairs (that's the M1 Max die). Everyone keeps rambling on about that idea because that Bloomberg reporter said it was coming and got it wrong. It won't happen.
Apple certainly can and probably will do quad dies at some point, it'll just be with a new die. The IRQ controller in Jade is only synthesized for two dies, but the architecture scales up to 8 with existing drivers (in our Linux driver too). I fully expect them to be planning a crazier design to use for Mac Pros.
That's a big change from Apple where they've historically put their newest processor in every single phone they launch (even the $430 iPhone SE announced last week has the A15 now).
I wonder if it's purely a cost cutting measure, or if they're not expecting good enough yields to supply them for every iPhone, or if they're holding some fab capacity back to have room for the higher end chips alongside the A16.
There's been rumors that they'll be skipping A15 cores for the upcoming M2 processors.
If they skipped over the best-selling iphone, that would give them a TON of extra space for M2 chips. This would allow them to put a little more ground between the new Air with the M1 and the pro iPads. It would also allow them to drop a new version of the macbook air and drive a lot of sales there. I know I'd gladly upgrade to a M2 model -- especially with a decent CPU bump and a 24/32GB RAM option.
Then again, they could just stick with what people expect. I wouldn't be surprised either way.
The iPhone 12 Pro was perhaps the least differentiated high-end model Apple has ever put out.
I think the 13 Pro has a few features that make it a bit more of a compelling buy:
- The new telephoto lens is a massive improvement (I wonder if your last experience was with the 12 or older? The new camera is actually worth something while the old one had mediocre quality compared to the main lens).
- ProMotion has no tangible benefit, but it makes every interaction with the screen look smoother. When you go back to old phones that don't have it, it's jarring. I can see why some of the Android-using tech enthusiasts have criticized Apple for not delivering high refresh rate for so long.
- The previous iPhone 12 had identical main/wide cameras with the Pro model unless you got the Max variant, which is no longer the case. The iPhone 13 Pro has different/better cameras all around over the 13.
- The GPU of the 13 Pro has an extra core over the 13, which was not the case for the iPhone 12 lineup. Anyone who does mobile gaming on graphically intense games should probably choose the Pro model over the regular one.
- Significantly better battery life over the non-Pro version, which was not the case for the 12 models, which had identical ratings.
I know the value, but I work in tech and spend tons of time digging into hardware as a hobby. Camera matters to some, but most of the rest are pretty bare features compared to the $200 (20%) increase in price.
When I list all the things I can buy with $200, where do these features rank in comparison to those other things? I'm blessed with a good job, so I can afford the luxury. I was poor when I was younger and I definitely wouldn't be spending that for those features. $220 out the door would be almost 20 hours of work at $15/hr (after taxes).
But, there are some other points to consider:
- It seems like most people in the USA who buy mid to high-end phones finance their phones from carriers, and pay 0% interest for it. So, what the consumer is really considering is "is the Pro model worth $5-8/month more to me?" or "Would I pay $200 extra over 2-3 years?" and I think that's an easier justification for many people.
- Carriers offer a number of financial incentives and discounts in exchange for loyalty (there aren't any contracts anymore, but there are "bill credits" that function the same way).
- You did use $15/hour as an example, which around the median US salary, but Pro models are not intended to be the top selling model for the median earner in the US. They're marketed at, I would guess, the top 20% of earners, which lines up with the Pro/Pro Max models only making up 20% of iPhone sales in 2020 [1]. That would mean that Apple would expect individuals buying the iPhone Pro models to make about $75,000/year or greater. About 10% of the population makes a 6-figure salary. [2]
- Smartphones are the primary communication and computing device for many if not most people. I think that there are many people who see the smartphone as the most valuable possession they own.
[1] https://www.knowyourmobile.com/phones/most-popular-iphone-mo...
[2] https://en.wikipedia.org/wiki/Personal_income_in_the_United_...
I find this attitude bizarre. It's not any cheaper! I guess it can make a difference if you have cash flow issues. But an iPhone Pro is decidedly a luxury, so if you have cash flow issues then you probably just shouldn't buy one?
It's not really about the total price paid, it's about whether the item is affordable on a monthly basis.
I agree with your philosophy where a lot of people are way too willing to let easy financing change their budget. I also assume that free financing would inflate prices (e.g.: cars and homes).
Would the iPhone and other flagship phones cost $700-1000 if financing wasn't so common? I don't think so, personally.
The other aspect of this is that, mathematically, $200 paid now is objectively worth more than $200 paid over time thanks to the time value of money. [1]
Let's say we're at a "normal" 2% rate of inflation. If I buy a $30,000 car today with 0% financing over 6 years, by the time I hit year 6 my monthly payment is representing a 12% lower value (inflation compounded yearly) than when I started.
Technically, even if you have $1000 in cash to buy your new iPhone, you "should" just finance it and invest/save/use the rest elsewhere.
[1] https://www.investopedia.com/terms/t/timevalueofmoney.asp
With my ADHD issues, the combination of
>The generally long battery of the ProMax in particular (I had issue with my phone dying midday, forgetting to charge)
>iOS 15 Focus mode
>UWB wireless trackers
>The most effective ANC in any wireless headphones
>Apple watch for notifications, timers, etc
Are all very helpful and the promax is just good for pragmatic and practical reasons. I would definitely say it's a high priced item but I wouldn’t frame it as a luxury. I fundamentally see my phone as assistive technology, it can’t die on me, Apple makes their longest lasting phone their most expensive.
Sure it does, particularly for a direct input device higher refresh rate screens improve input speed and accuracy.
Last year I upgraded from a 6s+ to a 12 and I can confirm this the other way around: my old phone did feel slow for a couple of key apps; turns out they are still slow on the new phone. They are just poorly written, and one is basically just a CRUD app with no excuse.
So my lesson is to likely keep this phone for a decade. It's not like I couldn't afford to upgrade but why bother?
[happy owner of a 6s here]
Another big difference is 5G. I can easily get 500+ MBit/s downstream while outside.
Oh, and about the cameras: the one update I especially like is the night photography, I can now even take photos of the starry night sky.
Besides that: it’s just a phone, it just has an absurdly powerful CPU in it, and now it does text detection in images. Screenshots and photos have selectable text in it.
Yes, the battery was kinda shot but I could have replaced it.
nm notation used to mean the width of the smallest feature that could be made. Even today there are processes such as atomic layer deposition (ALD) that allows singular atom thick features. The difference between nodes now are in shrinking macro features, you don't necessarily make them smaller, more important is density. This is currently done with 3d transistors (finfet) and perhaps in the future going full vertical. When all other optimization have been exhausted it's likely to see multiple layers of stacked transistors simulator to what they are doing with NAND memory chips. Eventually even that will hit a wall due to thermal limitations. Beyond that people have proposed using carbonnano tube transistors. That tech is very early but has been proved to function in labs. If we ever figure out how to manufacture carbon nanotubes chips, it will be truly revolutionary; you could expect at least another 50 years of semiconductor innovation.
That's the problem. We can't. All those technologies are in such a primordial state, if at all, that we don't even know if we will ever be able to use them efficiently 20 years from now.
Although if you just shrink the chips and keep the transistor count the same, then you have a more energy efficient chip. Which is especially useful for portable devices.
The real numbers have been around 20 nm for a decade. They decreased a bit with Intel's competitors achieving better lithography than them before them. And we are in the realm of tons of little tricks that improve density and performance - nothing really dramatic but there are still improvements here and there. The tens of billions of dollars thrown at research achieved them but it is not comparable to the good old days of the '80s, '90s and the '00s
I don't think that's fair. Density is still increasing fairly substantially. Just going off of TSMC's own numbers here:
16nm: 28 MTr/mm2
10nm: 52.5 MTr/mm2
7nm: 96.5 MTr/mm2
5nm: 173 MTr/mm2
Performance (read: clock speeds; but for transistors those are one & the same) are not really increasing, though, those have pretty much plateaued. And the density achieved in practice doesn't necessarily keep up, as the density numbers tend to be for the simplest layouts.Because it's all subjective now, companies went wild with marketing, because consumers know "lower nm => better". But, say, GF 14nm is much more comparable to Intel 22nm, and GF 12nm is still solidly behind late-gen 14++, probably more comparable to TSMC 16nm. Generally Intel has been the most faithful to the "original" ratings, while TSMC has stretched it a little, and GF/IBM and Samsung have been pretty deceptive with their namings. Intel finally threw in the towel a year or so ago and moved to align their names with TSMC, "10nm ESF" is now "Intel 7" (note: no nm) and is roughly comparable with TSMC 7nm (seems like higher clocks at the top/worse efficiency at the bottom but broadly similar), and they will maintain TSMC-comparable node names going forward.
Anyway, to answer OP's question directly though, "what comes after 1nm" is angstroms. You'll see node names like *90A or whatever, even though that continues to be completely ridiculous in terms of the actual node measurements.
future improvement is going to come from the same place it mostly comes from now: better design that unlocks better density and a revolutionary new litho process that as of yet doesn't exist
TL;DR: things get really murky after a notional 2.1nm generation. Past that we'll need a new generation of EUV sources, advancements in materials, etc, that AFAIK are still quite far from certain (but I am not an expert on this stuff by any means).
I personally think we're headed to a stall for a while where innovation will focus mostly on larger packaging/aggregation structures. Chiplets and related are definitely here to stay. DRAM is moving in package. Startups are playing around with ideas like wafer scale multiprocessors or ssds. I think clever combinations of engineering at this level will keep us with momentum for a while.
For GPUs, yield is less of an issue... the chips are manufactured with the expectation of a number of cores of the many thousand small (in terms of silicon area) ones being defective - overprovisioning, basically. That allows them to simply bin the sliced chips according to how many functional core units the individual chip has.
In contrast, even the Threadripper AMD CPUs have only 64 large cores which means the impact of defects is vastly bigger, and overprovisioning is not feasible.
Meanwhile if you're making 6 x 8 core chiplets and one of those cores is defective, well that chiplet can go into a 48 core or be a midrange consumer cpu or something, and you'll just pick one of your many many other 8 core chiplets to go with the rest for the 64 core.
>Qualcomm has decided to switch back to TSMC for the Snapdragon 8 Gen2 Mobile Platform. Samsung’s 4nm process node is plagued by a yield rate of as low as 35 percent.
https://www.techspot.com/news/93520-low-yield-samsung-4nm-pr...
https://www.theverge.com/22972996/apple-silicon-arm-double-s...
The author ran Apple Europe and then moved to the US and was an Apple VP for a long time. If anyone is allowed to have this kind of attitude then it's reasonable in Gassée.
In people in general, it's...weird.
Apple ultimately went with the NeXT / Steve Jobs combo, quite wisely, but for a long time there was a whole gang of BeOS fanboys lamenting that decision.
> Second, the recourse to two M1 Max chips fused into a M1 Ultra means TSMC’s 5 nm process has reached its upper limit. It also means TSMC’s 3 nm process isn’t ready, probably not shipping until late 2022. Apple, by virtue of their tight partnership with TSMC has known about and taken precautions against the 3 nm schedule, hence the initially undisclosed M1 Max UltraFusion design wrinkle, likely an early 2021 decision.
"recourse"... "design wrinkle"... wouldn't something like UltraFusion be an architectural goal at the outset, rather than something grafted on later? Feels pretty fundamental.
I have a vague memory that AMD has/had something similar -- the idea what their entire range would be the same basic core, fused together into larger and larger configurations. Seems like a smart move to concentrate engineering effort. But chip design not even slightly my area.
You are correct - AMD CPUs from 2016 onwards make use of a collection of up-to-8-core chiplets linked by what they call "Infinity Fabric"
> While working on AIC2 we discovered an interesting feature… while macOS only uses one set of IRQ control registers, there was indeed a full second set, unused and apparently unconnected to any hardware. Poking around, we found that it was indeed a fully working second half of the interrupt controller, and that interrupts delivered from it popped up with a magic “1” in a field of the event number, which had always been “0” previously. Yes, this is the much-rumored multi-die support. The M1 Max SoC has, by all appearances, been designed to support products with two of them in a multi-die module. While no such products exist yet, we’re introducing multi-die support to our AIC2 driver ahead of time. If we get lucky and there are no critical bugs, that should mean that Linux just works on those new 2-die machines, once they are released!
https://asahilinux.org/2021/12/progress-report-oct-nov-2021/
I'm not saying that's what happened, but it's a charitable reading of the original post.
If we keep scaling up number of processors rather than clock speed, what is going to be the maximum number useful cores in a laptop or desktop? 20? 100? 1000? At some point adding more cores is going to make no difference to the user experience, but the way we are going we'll be at 1000 cores in about a decade so we better start thinking about it now.
Or to put it another way, what normal workloads will load up all the cores in the new M1 chip?
Being a software developer, compiling things is the obvious choice, except when you come to that rather serial linking phase at the end of the compile job. Already my incremental Go compiles are completely dominated by the linking phase.
There are a few easy to parallelise tasks, mostly to do with media (as it says in the article). However a lot of stuff isn't like that. Will 20 cores speed up my web browser? How about Excel?
Your average user is going to prefer double the clock rate of your processors to doubling the number of processors.
Anyway, I don't want to rain on Apple's parade with these musings - the M1 Ultra is an amazing achievement and it certainly isn't for your average user. I wish I had one in a computer of my choice running Linux!
> Your average user is going to prefer double the clock rate of your processors to doubling the number of processors.
I disagree. The reality is that these days people are running multi-threaded workloads even if they don't know it. Running a dozen chrome tabs, Slack, Teams, Zoom, some professional tools like IDE's, Adobe creative suite, etc. adds up very quickly to a lot of processes that can use a lot of cores.
It’s wild to see that in print.
If your code is bottlenecked on a single thread, or if it doesn't scale well to higher thread counts, Apple is actually great right now. The downside is that you can't get higher core counts, but that's where the Pro and Ultra SKUs come in.
(The real, real downside is that right now you can't get higher core counts on M1 without being tied to a giant GPU you may not even use. What would be really nice is an M1 Ultra-sized chip with 20C or 30C and the same iGPU size as A14 or M1, or a server chip full of e-cores like Denverton or Sierra Forest, but that's very much not Apple's wheelhouse in terms of products unfortunately.)
That's the problem, though -- if you clock yourself much lower, of course you can get higher IPC; you can pack more into your critical paths.
Now, certainly Apple has some interesting and significant innovations over Intel here, but quoting IPC figures like that is highly misleading.
https://images.anandtech.com/graphs/graph17024/117496.png
That's an absolutely damning chart for x86, at iso-power the M1 Max scores 2.5x as high as a 5980HS in FP and 1.43x as high in integer workloads, despite having just over half the cores and ~0.8x the transistor budget per core. So it's a lot closer to the ~2.5-3x IPC scores than you'd think just from "but x86 clocks higher!". And these results do hold up across the broad spectrum of workloads:
https://images.anandtech.com/graphs/graph17024/117494.png
Yes, Alder Lake does better (although people always insist to me that Alder still somehow "scales worse at lower power levels than AMD"? That's not what the chart shows...) but even in the best-case scenario, you have Intel basically matching (slightly underperforming) AMD while using twice the thread count. And that's a single, cherrypicked benchmark that is known for favoring raw computation and disregarding performance of the front-end, if you are concerned about the x86 front-end, this is basically a best-case scenario for it... high code compactness and extremely high threadability. And it still needs twice the threads to do it.
https://i.imgur.com/vaYTmDF.png
Like your "but x86 uses higher clock rates", you can also say "but x86 uses SMT", so maybe "performance per thread" is an unfair metric in some sense, but there is practical merit to it. If you have to use twice the threads to achieve equal performance on x86 then that's a downside, where Apple gives you high performance on tasks that don't scale to higher thread counts. And if Apple put out a processor with high P-core count and without the giant GPU, it would easily be the best hardware on the market.
I just strongly doubt that "it's all node" like everyone insists. Apple is running fewer transistors per core already, and AMD/Intel are not going to double or triple their performance-per-thread within the next generation regardless of how many transistors they might use to do it (AMD will be on N5P this year, which will be node parity with Apple A15). x86 vendors can put out a product that will be competitive in one of several areas, but they can't win all of them at once like Apple can.
And going forward - it's hard to see how x86 fixes that IPC gap. You can't scale the decoder as wide, Golden Cove already has a pretty big decoder in fact. A lot of the "tricks" have already been used. How do you triple IPC in that scenario, without blowing up transistor budgets hugely? Even if you get rid of SMT, you're not going to triple IPC. Or, how do you triple clockrate in an era when things are actually winding backwards?
Others are very, very confident this lead will disappear when AMD moves to N5P. I just don't see it imo. The gap is too big.
But from some quick searching, excel will split out independent calculations into their own threads. So for that, the answer seems to be: it depends. If you're using 20 cores to calculate a single thing, it seems like the answer is "no". But if you're using 20 cores to calculate 20 different things, it seems like the answer is "yes".
It seems unavoidable that you can get more total performance with larger numbers of slower cores than smaller numbers of faster cores. The silicon industry has spent the entire multi-core era - the last 15 years - fighting this reality, but it finally seems to have caught up with us, so hopefully in the next few years we will start to see software actually start to adapt.
A55 is probably 1/8 the performance, but something like 1/100 of the power consumption and a miniscule die area. I wouldn't want to have all A55 cores on my phone though.
Performance per die area is also relative. For example, Apple clocks their chips around 3GHz. If they redesigned them so they could ramped them up to 5GHz like Intel or AMD, they would stomp those companies, but they would also use several times more power.
What is really relevant is something like the ratio of a given core's performance per area per watt to the same value for the fastest known core.
The only interesting area for ultimate low-power in general purpose computing is some kind of A55 with a massive SIMD unit going with a larabee-style approach for a system that can both do massive compute AND not have performance plummet if you need branchy code too.
Since the Intel e-cores still have a relatively wide decoder, e-core designs may be the part where the bill comes due for x86 in terms of decoder complexity. Sure it's only 3% of a performance core, but if you cut the performance core in half then now they're 6%. And the decoder doesn't shrink that much, Gracemont has a 3-wide decoder vs 4-wide on Golden Cove, and you still have to have the same amount of instruction cache (instruction cache hit rate depends on the amount of "hot code", and programs don't get smaller just because you run them on e-cores). A lot of the x86 "tricks" to keep the cores fed don't scale down much/any.
edit:
Intel Golden Cove: 7.04mm^2 with L2, 5.55mm^2 w/o L2
Intel Gracemont: 2.2mm^2 with L2, 1.7 mm^2 w/o L2
Apple Avalanche: 2.55mm^2. (I believe these are both w/o cache)
Apple Blizzard: 0.69mm^2 (nice)
Note that N7 to N5 has roughly 1.6x logic density scaling - so a Blizzard core would be 1.24mm^2 or roughly 73% of the transistor count of Gracemont for equivalent performance! For the p-cores the number is 82%.
This is one of the reasons I feel Apple is so far ahead. It's not about raw performance, or even efficiency, it's the fact that Apple is winning on both those metrics while using 2/3rds the transistors. It's not just "apple throwing transistors at the problem", which of course they are, but just they're starting from a much better baseline such that they can afford to throw those transistors around. The higher transistor count in total is coming from the GPU, the cores themselves Apple is actually much more efficient (perf-per-transistor) than x86 competitors.
Of course, it doesn't help that Intel lists laptop chip turbo frequencies to use either 95w or 115w and Anandtech's laptop review of one had the 12900H hitting those numbers with sustained power draw at an eye-raising 85w. That's 2-3x the power of M1 Pro and only 20-30% more performance.
That laptop also showed that cutting power from 85w to 30w roughly halved the performance. On the plus side, this means their power scaling is doing pretty well. On the negative side of things, it means their system gets worse multithreaded performance at 30w despite having 40% more cores.
Something I don't often see, but it does come up here and there. One nice thing about the M1 is the performance is consistent as it doesn't have a massive auto-scaling boost involved. An Intel or AMD chip might start off at top speed single thread, but then something else spins up on another core, and you take a MHz hit on your primary thread to keep your TDP in spec. The background task goes away, and the MHz goes back up. Lots of performance jitter in practical use.
Interconnects and IO also consume power. You can't just scale small e-core counts without also hitting power walls there too.
All that said, I'd love to see some E-core only chips come out of intel targeted at thin clients and long battery life notebooks.
they exist, that's called Atom. "e-core" is just a rebranding of Atom because the Atom brand is toxic with a huge segment of the tech public at this point, but Gracemont is an Atom core.
There's no Gracemont-based SKUs yet, but Tremont-based Atoms exist (they're in one of the models of NUC iirc), which is the generation before. Also, the generation before that is Goldmont/Goldmont Plus which are in numerous devices - laptops, thin clients, and NUCs.
Keep an eye on the Dell Wyse thin-client series, there are a lot of Goldmont-based units available if you can settle for a (low-priced surplus) predecessor.
Gracemont is such a huge departure from previous atom offerings I don't really consider them as having the same design goals. These new e-cores would be really nice for my use cases, better density and power efficiency than Ryzen.
Maybe with AMD dragging their feet on the V2000 and V3000 lines of low power offerings, I can get these sooner...
I suspect that rather than 1000 cores we might start to see more levels of cores, and hardware for more things. Already Apple has video encoding support. AI seems an obvious idea and it scales much better than most classical computing.
If I may bring up something that may be more of a wish: I wish that we could give up the idea of shared memory and we could have many more cores that communicated by shared messaging. We are already seeing this spread with webworkers - if it became cheap to create a new thread and computers weren't bottlenecked then maybe more games would use it too.
Web workers are basically the worst case, you have to serialise your data to and from JSON when passing it to and from a worker. It’s not built for performance. There have been many cases where people have tried to improve performance of their apps by offloading work to a web worker but the added cost of serialisation ultimately made it slower than running on the main thread.
This article was published on 13 March. It's been known for 5 days (as of the time of this comment) that the difference in weight is due to the Ultra variant's using a copper heat sink, as opposed to an aluminum one. The whole article has this kind of feeling of off-the-cuff, underinformed pontification, and I don't think it's a very good one.
Not this again. They don't mean anything except for being purely marketing designations
it was international women day on that day, i think it was a nice touch from apple
I remember the GUI being responsive and laid out in a way I wished Windows was at the time. I remember reading about the prospects for BeOS 5 menus later. They were going to have a ribbon of color follow your menu selections through drop downs. I forget the look since it’s been so long, but it was a cool idea. Would have made drill downs easier to follow. Notably, modern OSes can be pretty finicky about menu drill downs and outright user hostile. It’s pretty easy to lose an entire drill down by moving the mouse a couple pixels one or another way too far, for instance.
Mobile UI of course is amongst the most limited interfaces. We’ve gone backwards a lot in ways on mobile. It also seems mobile may be steering people away from certain careers by simply being good enough to ignore learning things like touch typing, Linux/foss, or hobbies that lead to tech careers. (Not sure how much sense this last point makes-just spreading to general trends I’ve heard or seen.)
Edit-maybe I’d say BeOS had a certain polish that seems lacking even in todays FOSS GUIs/OSes but especially back then.
The Mac basically solved this problem in 1986 when Apple first introduced hierarchical menus. To make it work, the UI layer has to be able to avoid strict hit testing of the mouse cursor during menu tracking, which I would conjecture is probably difficult in some environments.
There were gaping holes in functionality, but BeOS was a revelation at the time.
Be's all time cumulative net revenues were less than $5 million.
Anyone? Just me? ...ok
And “ultra” means “beyond” [great]. It enters the next realm. :)
M1 Max is ~20x22 mm (~430 mm²), double this, even without some of the interconnect die space, doesn't fit into the reticle anyway.
It's telling that almost the only "bad" thing that you can say about the M1 Ultra is that its single threaded performance is on par with the M1, whose performance is great anyway. Apple pumped up the integration, cache size, pipeline length, branch prediction, power efficiency and what not.
I think that in terms of clock frequency increase that road is closed, and has been for 15 years already.
Realistically the only disadvantage I heard about Apple Silicon is that the GPU performance is not quite as earth-shattering as they claim.
The M1 is truly a great thing. It beats the pants off the Intel 2019 MBP that work gave me while I fixed some M1 problems.
That is, however, comparing Apple Intel to Apple Silicon. The 2019 Intel MBP is, on an absolute scale (vs. my own AMD laptop of the same year), completely and utterly incompetent.
Comparing Apple Silicon to Intel and AMD isn't as straightforward, and there's a lot of good and bad for all three. Apple is now merely competitive.
Possibly, except M1 runs at relatively low clock speeds of around 3.2ghz. This is in no small part how it achieves good power efficiency. It's a bit surprising that a wall powered unit is still capped at this clock speed, although whether that's intentional or just something Apple hasn't gotten around to fixing is TBD. That is, the M1 largely lacks the load-based turbo'ing that modern Intel & AMD CPUs have. So it's "stuck" at whatever it can do on all-cores & max load. This could be intentional, that is Apple may just not ever want to venture into the significantly reduced perf/watt territory of higher clock speeds & turbo complications. Or it could just be an artifact of the mobile heritage, and might be something Apple addresses with the M2, M3, etc...
The Intel/x86 mantra that desktops should be allowed to be massively inefficient just to pump up the clock speed fractionally is what's changing.
I for one, agree with Apple - 450W beasts aren't really needed. Most workflows can be (or already are) parallelized so multiple cores can demolish what a fast single-thread can tackle.
Nonsense. Nobody is going to just give up "free" performance.
> I for one, agree with Apple - 450W beasts aren't really needed.
Nobody makes a 450W consumer CPU, so this is a strawman. The M1 Ultra's 200W already puts it quite far beyond the typical consumer setup of 65-125W anyway.
Regardless if 450W lets your work finish faster than 200W, that's a tradeoff ~everyone makes. Nothing about that is changing.
> Most workflows can be (or already are) parallelized so multiple cores can demolish what a fast single-thread can tackle
If this is truly the case for you then you'd already be on the Threadripper/Epyc train and the M1 Ultra would be kinda boring.
No, they're done with "free" waste. Apple chips idle at far lower and due to efficiency cores, handle moderate workloads with low TDP.
> Nobody makes a 450W consumer CPU
CPU, yes. But the M1 (along with all Apple chips) is a SoC so if you include graphics, storage, memory and motherboard - you can easily eclipse 450W for many enthusiast consumer (gaming) builds. Most gaming builds are 300-500W.
Alder Lake also has efficiency cores. And everyone has aggressive power gating & ultra low idle states, that's not something Apple came up with or innovated.
Apple's inability to turbo isn't some magic advantage. It's just a limitation, and there's no reason to believe they are happy with that limitation nor that others who don't have that limitation would for some reason copy that. Or more accurately, why would AMD & Intel regress back to where they were 5-10 years ago?
> you can easily eclipse 450W for many enthusiast consumer (gaming) builds. Most gaming builds are 300-500W.
If we're talking gaming builds then M1 Ultra is pretty much completely irrelevant as the gaming performance of the M1 Pro & Max were terrible (likely more a software problem than a hardware one, but still). The Max couldn't even keep up with a 3060 mobile. So a 300-500W desktop gaming build will be running laps around an M1 Ultra in games.
Seems just the opposite of what Intel's doing where you might get a fast core, slow core, AVX2, AVX512, etc on a new product. Doubly so when AVX512 worked on Alder lake chips... till it was disabled by Intel.
I did hear that the m1 pro/max didn't do some GPU benchmarks particularly well, although I'm anxious waiting to hear how the M1 ultra does.
I have an old 27" LED Cinema which I used with a PC for many many years, and then with Ubuntu native... and now back to the mac on a Mac Mini.
I'm ithcing to replace it eventually with its "double pixel density" big brother, which is essentially what this new Studio Display is (exactly double of 2560x1440). Personally I love the glass pane, and I really dislike those "anti glare" bubbly/grainy coatings I've seen on PC displays.
Of course, the PC would need an appropriate USB-C connector with support for 5K resolution.
[1]: https://www.theverge.com/2022/3/9/22969789/apple-studio-disp...
> We also tried connecting a 4K monitor with a USB-C to DisplayPort adapter, and that worked well - as expected.
[1] https://www.eurogamer.net/articles/digitalfoundry-2019-02-28...
Now if only I could find a box with a button on it to switch the monitor between two or three computers, at full resolution, retaining Power Delivery and attached USB devices. I'd buy one right now.
This new Studio 27" isn't really a good display on its own. It's an 8 year old panel and is missing a host of modern display upgrades like higher refresh rates, variable refresh rates, HDR, or local dimming
Unfortunately, Apple no longer sells the LG Ultrafine 5K [1], and no one knows if LG is even going to restock them.[2] So, you'll have to find one used, and you'll have to hope that LG continues to service this incredibly flaky series of monitors when you inevitably run into an issue.
On the flip side, if you don't care about the pixel density, you could have bought any of the low res gaming monitors, or 4k 28" monitors, or whatever other ultrawide, low PPI monstrosity the market has coughed up in the past eight years. They've been waiting this long for a reason.
You are stuck choosing between those modern features you listed and a >200ppi display. That is the state of the market right now. Until Apple solves this issue and charges you like $3,000 for the privilege later this year.[3]
----------
[0] https://pixensity.com/list/desktop/
[1] https://www.macrumors.com/2022/03/12/apple-lg-ultrafine-5k-d...
[2] https://www.lg.com/us/monitors/lg-27md5kl-b-5k-uhd-led-monit...
[3] https://www.macrumors.com/2022/03/10/studio-display-pro-laun...
>combining multiple GPUs in a transparent fashion [is] something of a holy grail of multi-GPU design. It’s a problem that multiple companies have been working on for over a decade, and it would seem that Apple is charting new ground by being the first company to pull it off.
https://www.anandtech.com/show/17306/apple-announces-m1-ultr...
Does this tech apply to discrete graphics? You can't really connect separate cards with this, right?
Are you saying this is a blow to Intel's graphics? Or maybe you're implying it's a way for integrated graphics to become dominant?
They would fight Apple in this. They also have AMD to fight for the non-M1 chip market... For a company used to enjoying dominance they have a lot of work to do to remain at the top.
Poking around Apple's 2021 marketshare wasn't anything particularly special, and https://www.statista.com/statistics/576473/united-states-qua... isn't showing any sort of "M1 driven spike", either. Apple's y/y growth was larger than PC's y/y growth (Annual growth for Mac was 28.3%, against 14.6% for the global PC market), but note there both are still growing. Meaning M1 didn't suddenly convert the industry.
Regardless a one year data point certainly doesn't prove a trend, certainly not of the "Apple is going to destroy Intel/AMD/Nvidia" variety.
And fair enough to say many people won’t switch because of inertia / software lock in which is true - but that’s not because of the quality of M1 or MacOS.
Kind of my point though; the low end isn't even worth it for PC manufacturers anymore. And by the time you get to a ~$600 laptop, you can have a cheap M1 device with 10x the performance/watt. The new M1 iPad Air starts at $700, and absolutely blows any existing Intel laptop out of the water, short of a desktop replacement gaming rig with discrete GPU.
Really this dod not come as a surprise to anyone interested in Apple's chips, it was rumored as "Jade-2C" for a long time. And they were not expected to release a "pro" version of M2 before the standard version for the new MacBook Air neither.
Is it impossible to use the 27" iMac as a display monitor for the new Mac Studio?
I use https://astropad.com/product/lunadisplay/ to use my iPad or an old Mac's screen as a secondary screen and it's good.
Do Mac's not have USB ports? It's truly bizarre that you need 1 thunderbolt port per monitor. That's an incredible amount of wasted bandwidth.
It might be possible in theory now, but I suppose that ship has sailed.
I also wonder how well running something like a VNC client fullscreen on the old iMac with universal control might work. (I've also used Jump Desktop's "Fluid" protocol which is the same general idea as VNC, though provided a higher-quality lower lag connection in my case.)
I think there are some decent solutions for a secondary display, but I kind of doubt any of these would be good enough for a primary display for most use cases.
I would guess all of these have tradeoffs in terms of lag, frame rate, quality, reliability, etc. though I'd love to hear different.
DIY display from the hardware should be possible with a display driver from Aliedpress, though.
What I find interesting is that there’s no desktop Mac with an M1 Pro, leaving a gap between the entry-level M1 in the Mac mini, and the Mac Studio with its M1 Max.
For those who might not remember the full lineup: the M1 Pro and M1 Max have the same CPU part, the main difference is the number of GPU cores. For many CPU-bound applications, the Pro is all you need.
I wonder if this is an intentional strategy to sell the more expensive product or if it’s supply related.
It won't launch until 3nm is ramped up.
But that is when it is completely over with Intel, AMD, Nvidia completely.
So presumably $16,000 (or greater) systems?
In what way does this eliminate competitors from the market? Also is Apple doing something that is literally impossible for anyone else to do?
And will the use cases that Apple currently does not cover cease to exist?
AMD is already shipping on 9 tile SoCs, and Intel is doing tile stuff, too.
Unless Apple gets back into the server game, and gets a lot more serious about MacOS, pretty much nothing Apple does makes it "game over" for Intel, AMD, or Nvidia. Especially not Nvidia who is still walking all over every GPU coming out of Apple so far, and is so ridiculously far ahead in the HPC & AI compute games it's not even funny.
Disadvantage of this technique is yields will be worse because the individual component is bigger, and this is limited to combining exactly two chips because it's not a ring-bus design.
The power efficiency of M1 is vastly overstated. Reminder here that the M1 Ultra is almost certainly a 200W TDP SoC (since the M1 Max was ~100W, and this is 2x of those...)
So 10x M1 Max's would be ~1000W. That's possibly to put in a server chassis, but it's of course also not remotely 1/10th the wattage of existing server CPUs, either, which tend to be in the 250-350w range.
And interconnects aren't free, either (or necessarily even cheap). The infinity fabric connecting the dies together on Epyc is like 70w by itself, give or take.
114 Billion transistors are going to draw some power.
An RTX 3090 is 28.3 Billion transistors drawing 350W or so for just the GPU portion of the system.
It's still not ready for everyday usage (though the people working on porting it might already be using it), but it's way moving a lot faster than one would guess. I'm considering a Mac Mini build machine sometime during 2022, it might be feasible for that.
See https://asahilinux.org/2021/12/progress-report-oct-nov-2021/ or follow https://twitter.com/marcan42
Some are yet to receive orders placed since January for the 2021 MBP. Especially for anything beyond the base model. I wonder if they have the capacity to serve the current one going on Sales this week.
Well yes, the event was on March 8, International Women's Day.
It seems that around 90% of developers are men[1]. Therefore, using a binomial distribution, if 6 developers were picked at random, the chance that all of them would be women (or other non men gender) is 0.0001%. (Interestingly, there is a 53% chance that all 6 of would be men).
The reason for such unlikely occurrence is that most likely, as the parent mentioned, Apple wanted to feature women developers for International Women's Day.
[1] https://www.statista.com/statistics/1126823/worldwide-develo...
I'm glad Apple are showcasing the diversity of their workforce (I raised three daughters of my own, wish they would have shown an interest in programming, not really) but I worry that there is a danger of backlash for going too far.