Compared to previous macs and igpus - an nvidia gpu will still run circles arounnd this thing
Nvidia very likely has leading top end performance still, but "running circles around this thing" is probably not a fair description. Apple certainly has a credible claim to destroy Ampere in terms of power per watt - just limiting themselves in the power envelope still. (It's worth noting that AMD's RDNA2 already edges out Ampere in performance per watt - that's not really Nvidia's strong suit in their current lineup).
[1]: https://www.apple.com/v/macbook-pro-14-and-16/a/images/overv... - which in the footnote is shown to compare the M1 Max to this laptop with mobile RTX 3080: https://us-store.msi.com/index.php?route=product/product&pro...
[2]: There's a lot of things wrong with in how vague Apple tends to be about performance, but their unmarked graphs have been okay for general ballpark estimates at least.
Not for loading up models larger than 32GB it wouldn't. (They exist! That's what the "full-detail model of the starship Enterprise" thing in the keynote was about.)
Remember that on any computer without unified memory, you can only load a scene the size of the GPU's VRAM. No matter how much main memory you have to swap against, no matter how many GPUs you throw at the problem, no magic wand is going to let you render a single tile of a single frame if it has more texture-memory as inputs than one of your GPUs has VRAM.
Right now, consumer GPUs top out at 32GB of VRAM. The M1 Max has, in a sense, 64GB (minus OS baseline overhead) of VRAM for its GPU to use.
Of course, there is "an nvidia gpu" that can bench more than the M1 Max: the Nvidia A100 Tensor Core GPU, with 80GB of VRAM... which costs $149,000.
(And even then, I should point out that the leaked Mac Pro M1 variant is apparently 4x larger again — i.e. it's probably available in a configuration with 256GB of unified memory. That's getting close to "doing the training for GPT-3 — a 350GB model before optimization — on a single computer" territory.)
You could throw a TB of memory in something and it won't get any faster or be of any use for 99.99% of use cases.
Large ML architectures don't need more memory, they need distributed processing. Ignoring memory requirements, GPT-3 would take hundreds of years to train on a single high end GPU (on say a desktop 3090 which is >10x faster than m1) which is why they aren't trained that way (and why NVidia has the offerings set up the way they do).
>That's getting close to "doing the training for GPT-3 — a 350GB model before optimization — on a single computer" territory.
Not even close... not by a mile. That isn't how it works. The unified memory is cool but its utility is massively bottlenecked by the single cpu/gpu it is attached to.
It's just that we mostly use GPUs for embarrassingly-parallel problems, because that's mostly what they're good at, and humans aren't clever enough by half to come up with every possible way to map MIMD problems (e.g. graph search) into their SIMD equivalents (e.g. matrix multiplication, ala PageRank's eigenvector calculation.)
The M1 Max isn't the absolute best GPU for doing the things GPUs already do well. But its GPU is a much better "connection machine" than e.g. the Xeon Phi ever was. It's a (weak) TPU in a laptop. (And likely the Mac Pro variant will be a true TPU.)
Having a cheap, fast-ish GPU with that much memory, opens up use-cases for which current GPUs aren't suited. In those use-cases, this chip will "run circles around" current GPUs. (Mostly because current GPUs wouldn't be able to run those workloads at any speed.)
Just one fun example of a use-case that has been obvious for years, yet has been mostly moot until now: there are database engines that run on GPUs. For parallelizable table-scan queries, they're ~100x faster still than even memory databases like memSQL. But guess where all the data needs to be loaded into, for those GPU DB engines to do their work?
You'd never waste $150k on an A100 just to host an 80GB database. For that price, you could rent 100 regular servers and set them up as memSQL shards. But if you could get a GPU-parallel-scannable 64GB DB [without a memory-bandwidth bottleneck] for $4000? Now we're talking. For the cost of one A100, you get a cluster of ~37 64GB M1 Max MBPs — that's 2.3TB of addressable VRAM. That's enough to start doing real-time OLAP aggregations on some Big-Ish Data. (And that's with the ridiculous price overhead of paying for a whole laptop just to use its SoC. If integrators could buy these chips standalone, that'd probably knock the pricing down by another order of magnitude.)
Mindlessly throwing more memory does encompass diminishing returns in 99.99% of use cases because extra memory will inflict a very large number of TLB misses during the page fault processing or during the context switching which will slow memory access down substantially unless:
1) the TLB size in each of the L1/L2/… caches is increased; AND
2) the page size is increased, or the page size can be configured in the CPU.
Earlier versions of MIPS CPU's had a software controlled, very small sized TLB and were notorious for being slow with the memory access. Starting with A14, Apple has increased an already massive TLB, which was on top of the page size having been increased from 4kB to 16kB:
«The L1 TLB has been doubled from 128 pages to 256 pages, and the L2 TLB goes up from 2048 pages to 3072 pages. On today’s iPhones this is an absolutely overkill change as the page size is 16KB, which means that the L2 TLB covers 48MB which is well beyond the cache capacity of even the A14» [0].
It would be interesting to find out whether the TLB size is even larger in M1 Pro/Max CPU's.
[0] https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
We've had AMD APU's for years, you can shove 256GB of RAM in there. But noone cares because a huge chunk of memory attached to a slow GPU is useless.
The ML folks are finding ways to consume everything the HW folks can make and then some.
Does Apple have any ISV certified offerings? I can't find one. I suspect Apple will never win the Engineering crowd with the M1 switch... so many variable go into these systems builds and Apple just doesn't have that business model.
Even with these crazy M1's, I still have doubts about Apple winning the Movie/Creative market. LED walls, Unreal Engine, Unity are being used for SOOOO much more than just games now. The hegemony of US centric content creation is also dwindling... budget rigs are a heck of lot easier to source and pay for than M1's in most parts of the world.
True, but the point here is that M1 is able to achieve outstanding performance per watt numbers compared to Nvidia or Intel.
There are absolutely use-cases where this is going to enable new ways of looking at content and give more control and ability to review stuff in the field.
Having a higher performance per watt numbers also implies less heat from M1's perspective. This means that even if someone isn't doing CPU/GPU heavy tasks, they are still getting better battery life since power isn't being wasted on cooling by spinning up the fans.
For some perspective, My current 2019, 16inch i7 MBP gets warm even if I leave it idling for 20 - 30 mins and I can barely get ~4hrs of battery life. My wife's M1 macbook air stays cool despite being fanless, and lasts the whole day with similar usage.
The point is performance per watt matters a lot in a portable device, regardless of its capabilities.
I am not associated with that guy. In fact I bought one for even my macmini. Get my m1 macmini to avoid all these hot air.
If you run biotcamp windows has registry to disable that and also system setting to limit to 99% (but seem still hot) for my playing with Vr and fs2020 using external egpu.
I worked in the content creators business making videos, photos and music and frankly the need for 15 hours of battery is a (very cool indeed) glamorous IG fantasy.
In reality even when we were really on the move (I used to follow surfers on several isolated beaches in South Europe) the main problem were the phones' batteries - using them in hotspot mode seriously reduce their battery life - and we could always turn on the van's engine and use the generator to recharge electronic devices.
Post processing was done plugged to the generator.
Because it's better to sit comfortably to watch hours of footage or hundreds of pictures.
I can't imagine many other activities that are equally challenging for a mobile setup.
It’s not just about rendering any more.
Now you can get performance off a battery for your entire work day for less money than the competition (if reports are to be believed).
In this scenario, would you render things in a cafe? Why not?
honestly, as a traveler and sometimes digital nomad, the real question is "why yes?"
There is no real reason to work in a cafe, except because it looks cool to some social audience.
Cafes are usually very uncomfortable work places, especially if you have to sit for hours looking at the tiniest of the details as one often does when rendering.
It’s like when the iPad came out and had a camera. “Who is going to lug around an iPad to take pictures!?!?”
But that’s exactly what I started seeing people do. Pull out iPads and take snaps.
Perhaps ML but that’s all proprietarized on CUDA so it’s unlikely.
Perhaps Apple could revive OpenCL from the ashes?
I was offered a large external monitor by my employer, but I turned it down because I didn't want to get used to it, and working in different locations is too critical to my workflow. But I'd love to see how people with more than 2 external displays are actually using them enough to justify the space and cost (not being facetious, I really would).
First - dedicated to personal Chrome profile, Discord, etc.
Second - dedicated to screen share - VS Code, terminal, JIRA, etc.
Third - Work Chrome Profile for email/JIRA/web browsing, note-taking (shoutout to https://obsidian.md), Slack.
I could certainly get by with fewer monitors, and do so when I am mobile, but I really enjoy the screen real estate when at my desk at home.
1. Web Conferencing Content
2. Web Conferencing participants video
3. Screen where I multitask in parallel
4. VDI session to a customer's environment for tests
My friend uses 5 monitors, and I would too if I wasn't mandated to use an iJoke computer at work.
Teams, browser and IDE mandate a minimum of 3 displays.