Apple's M1 Ultra comes with a 32MB TLB bottleneck
twitter.com
twitter.com
You'd almost never say "32MB of TLB". You'd say maybe, 8192 entries of TLB, each with 4kB pages, for instance. (Others here in Hacker News suggest that the M1 uses a 16kB page, so maybe 2048 entries of 16kB each?)
The Twitter posts are talking about temperature at first, tile-memory. I admit that I'm ignorant to tile-memory but... The traditional answer to TLB bottlenecks is to enable large-pages or huge-page support. So I'm already feeling a bit weird that large pages / huge pages haven't been discussed in the twitter thread. (Maybe its not the answer on the GPU side of things? But at least acknowledging how we'd solve it in CPU-land would help show off how features like huge-pages could help)
Like, none of the logic here makes sense to me at all. Maybe something weird is going on in the Apple / M1 world, but I have a suspicious feeling that maybe the person posting this Twitter thread is at best... misstating things in a very confusing manner and using imprecise language.
At worst, they might be fundamentally incorrect on some of these discussion points?
1. General purpose GPU code assumes optimization characteristics where memory constraints are less restrictive for this use case
2. Without optimizing for this platform, typical code will not exercise the full potential of the chip
The thread is discussing it as a design flaw; I don’t know if it is. But if normal “recompile for ____ target” in XCode or whatever isn’t enough for common workloads to be relatively close to optimal, and you have to do and know to do special work for the memory architecture… yeah that’s probably not ideal. And the other takeaway from the thread, that they’re likely to throw more memory hardware at the problem in the future, sounds right to me.
The question is whether it affect other or 3rd party products (that cannot compete, or not shown to the public eg the private api ). And does apple care.
Unless it shows is that limitation important ?
I don't know all the words that are being used in this tweet. But the words I _DO_ understand are being used incorrectly, which is a major red-flag.
1. The temperature issues are non-sequitur, but that's where the discussion starts. There are a whole slew of power-configuration items on modern chips, none of them discussed in this set of twitter posts.
2. TLB is a CPU-issue, traditionally. I understand that Apple has an iGPU that shares the memory controller, but GPU-workloads are quite often linear and regular. It is difficult for me to imagine any step in the shader-pipeline (vertex shader, geometry shader, pixel / fragment shader) that would go through memory so randomly as to mess up the TLB. If you're saying the TLB is the problem, you need to come forth with a plausible use-case why your code is hopping around RAM so much to cause page-walks.
3. In fact, everything they say about GPUs in the thread is blatantly wrong at its face, or possibly some kind of misstatement about the peculiarity of the M1 Apple iGPU.
4. I realize that mobile GPU programmers are always talking about this "tile-based rendering" business (popular on Apple / Android phones), and I've never bothered to learn the details of it. But I can almost assure you, even in my ignorance, that this has nothing to do with TLBs at all.
-----
The argument simply does not track with reality, or my understanding of GPUs in general.
The vague advice I've heard for GPUs on mobile systems is to learn about the peculiarities of tile-based rendering, and to optimize for it specifically, because that's the main difference in GPU-architecture between Desktop GPUs and Phone GPUs. But nothing discussed in this set of twitter posts seems to match what I've often heard about iPhone/Android GPU programming.
I've seen the experts advice on Android/Apple phone architecture / tile-based rendering. I don't know it, but... I think I can recognize it when I see them talking. This set of twitter posts is not an expert talking.
Likely because the person doesn't have any background in understanding the subject.
Twitter profile:
> Jesus Christ is King! Co-host/writer for the Max Tech YouTube channel.
They're a youtuber, that's the limit of their qualifications. And it's not like they're technical either. They review consumer electronics. They probably don't even know anything about CPU architecture.
The page size is indeed 16K. I don't know if 4K is supported at all in the GPU.
I agree though, the article reads like a rambly piece that doesn't follow and really just boils down to "my code runs slow on this machine and I'm going to blame the architecture" without going into proper details.
>transaction lookaside buffer
not
>translation lookaside buffer
I think hishnash made this discovery and this chain of tweets is echoing it.
"when a GPU is waiting for data, it can't just switch to work on something else" - GPU designs specifically use the idea of rotating between threads to not block on memory access.
"the concept that there is a local on-die memory pool you can read/write from with very very low perf impact is unthinkable in the current desktop GPU space" - Optimizing for the characteristics of the specific device's memory hierarchy including local/shared memory is a big part of GPU coding.
If Apple follows the pattern that companies like NVIDIA have used they will keep the overall shape of the GPU similar between generations (in terms of how many threads, memory hierarchy, etc). This will give app devs relatively similar + stable tuning targets, we'll see high dollar apps get aggressively optimized for specific chips, and as that happens power/heat will go up. For an extreme example of this maturation curve look at how console games get more and more out of the same hardware.
All that said the Metal tile shading API is a real thing Apple has been pushing for years on the mobile side, presumably because it fits the hardware and performs better, and now Macs use similar chips. No surprises there. And if an app is built around OpenGL and now needs to port to Metal to improve performance, sure, that could be a lot of work.
You are the computer in "Computer Science," and the science is really just a lot of math. Maybe you wanted to study CE instead? I made the same damn mistake.
Also like compilers, and most compiler engineers are CS grads. That's my dream job!
A computer scientist can solve very hard problems. But a computer engineer can also solve hard problems, yet they can also solve really easy ones, giving them the edge.
A decent CS program will include a some amount of history as well, who first did what and when, why the thing is called what it is, etc. Does that mean historians are computer scientists?! No, a university will round out a curriculum to include germane information that isn't really considered an essential element of that curriculum. In CS, the math is critical. So we can remove the computers and remove programming from the CS curriculum and still have a computer science program. But without the math, there is no CS.
In general, it seems you are making an argument about semantics that doesn't describe the world as it is but rather either the way you want it to be, or the way you want it to be described. Either where all those poisonous influences were purged from CS programs or where "actual" CS programs were rare to non-existent.
You can certainly have a preference that the term be used like that, but that is far from reality on the ground.
First of all, it isn't a pen and paper, or stylus on clay, or finger in the sand, just tools, that the digital computer replaces, it is a mind-boggling amount of time thinking, or figuring, or computing, too. You want to eliminate the thinking, but it is critical that you leave it where it is, because only humans, and possibly other mammals, and perhaps other animals, can think. Computers, like your PS5, can't think, and will not ever be able to. What you want to do, and others that believe falsely as you do, it seems, is dispatch with the person. Computer Science can't do that, and not ever, not until we can make people, and we already can so there's no incentive to find a much much harder and far far more expensive and much much much less fun way.
And people use computers for all kinds of things, but mostly, almost none of it is or is for computer science. And computer science is cheap, it doesn't cost anything but time (but I guess time is money, or the sqrt of evil etc.) So while it is luxurious to have a budget with a lot of cool tools available, there really are plenty of American and other Western universities that teach a decent CS curriculum without 4 computers per student, and their graduates are fully prepared for a CS career. I don't think you meant to insult them by calling them the Third World, but you kind of did that. A big budget helps any department. That doesn't mean money is computer science, although... computer scientists often work in economics positions and finance positions and secretly or officially do really neat computer science sometimes.
Secondly, fundamentally my argument is semantic, and your dismissal of semantics reveals some grave misjudgments. Semantics are of vital importance, and in every single exchange of understood communication, fully half the weight of everything there is in that is semantic in nature. If I don't understand what you're saying, or worse, if you don't understand what you're saying, then you can see there's a problem worth correcting.
The ideal candidate for computer science does not say, "I like computers, I want to work with computers, computers fascinate me. I want to know how they work. I love programming computers. I want to learn every language in use today, and be a programmer." The ideal candidate instead would say, "I love to compute, to figure. I want to know how to solve any problem. Puzzles fascinate me. I like figuring things out. I had exhausted my school districts math classes by the time I got to High School, and got a 4 on the Calculus AP exam in 10th grade. I audited some math classes at the local community college over the summers that I'm transferring in." Maybe that's a little too ideal. I was not, but math wizzes seem kind of common, I mean rare really, but always present. They make great computer scientists.
A computer science grad that wants to be a programmer will do well for themselves, programmers earn, and it is a highly respectable career. If nothing else will convince everyone, the difference in salaries between computer scientists and programmers should, except that some IT positions will inexplicably require a computer science undergraduate degree for $25/hr, so ignore the outliers when researching salaries.
I can see that you needed to rant at someone about your revelation that humans can do computation too, but could you not have found a victim whose comment was at least superficially, tangentially related to what you're banging on about?
The math itself maps to a higher abstraction when it comes to compilers. Focusing more on parsing, grammars and automata theory. CE courses tend to go to a lower level. I enjoy both sides and can see what you're trying to explain. What I got out of university was mostly the feeling that the courses could never go in depth due to the amount of detail and the abstractions that exist. I don't think a CE course would have went to enough depth to cover everything anyways (each course is still just in the realm of 100 hours). Because both the layers on top (Algorithms and computation theory) and the bottom (hardware) both have so much information that it can be daunting to try to learn all of it.
I do agree that it might be a pain point for some people who don't enjoy that part of CS. I'm curious how you came to regret it this much?
I'm not sure regret is the right word, but I am cynical because I was actively recruited by my university's CS dept. and it took me 2 years of study before I realized I was actually studying math, not computers, which was what I thought I was supposed to be studying. But I was mistaken, and it was too late to switch to CE and hope to graduate before the end of the decade. If I had been wise at 17yo, I would have known CE was what I thought CS was. Maybe I would not have worked in the field, but I sure would know a lot more about electronics and engineering, which was really were my young interests lay, I just didn't know any better.
Programming jobs pay pretty well, but if you have your CS degree, you can earn about 20% more starting out by landing a job as a computer scientist, or at least one that advertises for one. Programming is really a part of the IT field, but computer science isn't necessarily. For instance, the FAA uses computer scientists, so does the National Weather Service, so do major automobile manufacturers, the aerospace industry uses computer scientists. I mean, if programming is your thing, carry on, but you're actually limiting yourself if only seeking C++ programmer positions.
Not that it matters for SwordOfMyBone: what matters is what you learn when you enroll in a specific university's Computer Science course.
Btw, why couldn't programming or compilers be a subset of math? Look at eg 'type theory' https://en.wikipedia.org/wiki/Type_theory
Or optimization. It's a thing we do very much in math and in computing.
Applied mathematics is a wide field.
Programming is not mathematics, nor is it computer science. Programming is programming, and it is most like the recipes in cooking, which no one would ever confuse with math. If one wants to be a programmer, the thing to do is learn programming languages, no degree required. OTOH, if one wants to model the weather, or model traffic, or model society, or model the galaxy, or work in informatics, or hunt for weaknesses in the genetic code of some virus, or really, generally, solve complex problems of any ilk, then study computer science, because that is what computer scientists do.
Do not bother with CS if you want to work with computers, or be a computer tech, or systems administrator. Do not bother with CS if you just want to be a programmer, really, it'd be a waste of time, sort of like getting a MD because you want to be a RN, but this metaphor is not apt because there is no hierarchy like that in technology space. The confusion of programming with computer science is such a sensitive one, and a lot of schools now have Software Engineering departments. IMO, calling it engineering is being a little too kind to programming, even if we call a program an engine, it's really not. Also, there is no license available to engineer software, yet all other engineers need a license to work.
That being said, a Computer Science curriculum will entail learning a lot of programming, but that still does not mean that computer science is programming. But before learning program languages, in CS, one will study algorithms, and I believe those do fall under the scope of CS.
Here is another bad metaphor: You'll have to learn a good amount of Trigonometry in order to be able to do much Calculus, yet Calculus is not Trigonometry. Now, one can do good computer science without ever learning any programming language, and without ever touching a computer. But a computer is a tool, and the good analogy is that the computer is to the computer scientist what the telescope is to the astronomer. To be clear, though it was once so, astronomers do not build telescopes, nor are they necessarily expert in the construction of or even function of a telescope. The important part is they look through it, and that is how it is used.
No, I'm afraid most of the computer science that has ever been performed occurred prior to the 19th Century.
It's sort of how like geometry (literally 'the measure of land', originally an applied mathematical field used for surveying) is newer than literally drawing property boundaries. It was the formalization and abstraction that led to computer science as a true mathematical field, which happened after actual programming occurred. We can reproject CS concepts on these earlier programs looking back, but the field of CS is newer than they are.
It has been pointed out to me before that mostly what computer scientists do is reckon. This is why I am onboard with renaming Computer Science to Reckoning Science. It is a subtle thing, but I think it would prevent countless students from wasting the best years of their lives due to incorrect expectations initiated by the confusing name of the discipline. It really is important to understand that the "computer" in Computer Science is not a digital machine, a mainframe, server, a Dell, an Acer, or a Mac, instead it is a living breathing person. I have to envy the Math majors and Math grads. They knew exactly what they were getting into, not a one of them ever asked, "what do you mean, 'it's all math,' ?"
It's sort of like how geometry literally means 'the measurement of land' since it was originally an applied field developed for surveying farm plots for the bronze age super powers after each flood along the Nile and the levant wiped the property boundaries away. I think it's safe to say it has expanded beyond it's original intention, and CS has the same properties.
ML even stands for meta language.
See eg http://dev.stephendiehl.com/fun/ or https://en.wikibooks.org/wiki/Write_Yourself_a_Scheme_in_48_...
If you want an idea of what GPU programming is like you could look at some CUDA intros:
- https://developer.nvidia.com/blog/cuda-refresher-cuda-progra...
- http://cuda.ce.rit.edu/cuda_overview/cuda_overview.htm
If you don't have CUDA hardware you can get various versions of OpenCL that will at least work on your CPU. Lots of interesting compiler stuff happening targeting GPUs.
So it’s exactly as you say: of course one can improve performance by using these features, but it’s not that one has to do it to get good performance in the first place. If there are performance issues with the Ultra, memory hierarchy is an unlikely culprit IMO.
Perhaps there's an interesting discussion to be had about it?
https://www.statista.com/statistics/576473/united-states-qua...
https://www.counterpointresearch.com/us-market-smartphone-sh...
32MB of coverage is the same size for the TLB on a SM for a 1080ti. https://dl.acm.org/doi/10.1145/3491218 I don't think that's gone up a lot since then.
And GPUs in general are totally fine with stalling on memory accesses. That's core to how they work, and why you run into issues treating them like CPUs. They approach the problem of the memory hierarchy at a base level differently by accepting the fact that thread groups will stall on pretty much every memory access, and just barrel scheduling enough thread groups that you can keep ALUs saturated despite the memory dependencies due to massive parallelism. Cache and TLB misses aren't that big of a deal if you have enough work to cycle through. If you don't have massively parallel amounts of work, and can't just keep cycling through contexts, this is the main reason why some algorithms don't work on GPUs, but that's a global problem to how all major GPUs approach the problem of memory that's hundreds of cycles away.
> Hishnash: "What apps should be doing is loading as much data as possible into the tile mem and flushing it out in large chunks when needed. I bet a lot of the writes (over 95%) are for temporary values that could've been stored in tile mem and never needed to be written at all."
I mean, you shouldn't be performing large amounts of writes that aren't a completion of a program on just about any GPU architecture. If you care about the results, heavy random writes screws you just as hard in a traditional GPU architecture too. You really want to be able to look at the data flow and describe it in almost pure functional terms. If you can't look at a program and say "it writes here at this offset or offsets", you're going to have a bad time.
And finally, there's all sorts of ways you can crank up the GPU without doing writes at all, so the inability to turn it into a space heater just because of TLB pressure doesn't make a lot of sense. The xor chains that something like the bogomips calculator uses will eat up every cycle you give it and make their way through the compiler and hardware schedulers all without blocking on memory at all.
Then there's also the GPU tile buffer (which they may also be calling a "TLB" here?) - which I think is what he's referring to here - on a TBDR system there's an advantage if it can calculate multiple passes on the image while keeping all needed info in the on-chip tile buffer (Think stuff like depth and color data when rendering a large scene), as if the intermediate results aren't actually needed (such as a partially drawn scene where more objects will be drawn on top of some of the pixels, or an intermediate output before some post-processing pass) it may never actually touch the main memory so has the ability to save a lot of bandwidth.
I'm pretty sure it's the second the OP is referring to here, as if you blow past this buffer size you can hit a performance cliff, and methods to optimise it's use (such as rendering tiles of multiple passes instead of doing each pass on the full screen view separately) doesn't really make any difference to immediate renderers, so is something "new" developers need to take into account for deferred tiler architectures.
I suspect they've seen someone refer to a TLB (Translation Lookaside Buffer) and TLB (TiLe Buffer) and conflated the two.
So you’re right he may not mean Transaction Lookaside Buffer.
From my simple understanding of it. TBDR GPUs extract performance by tiling and binning primitives and handing that out for rasterization. The multiple passes allow it to work on both at the same time? Kinda like how a CPU pipelines stuff? I thought GPGPU workloads means that it skips the rasterization stage? So what's the problem with treating it like an IMR.
One of the cases where this can cause performance problems is, when you want to read the output of a previous render pass. If you want to be able to read arbitrary parts of the output of the previous render pass, the output buffer probably needs to be copied from the tile memory into a memory that can hold the whole buffer at once. Furthermore, this also means that all tiles of the previous render pass need to execute before the next one can run. This limits how much work can be done in parallel.
Those steps tend to be local, so instead of running step 1 on the whole screen, then step 2 on the whole screen, and so on, you could run all steps on the top-left tile (of e.g. 128x128 pixels), then all steps on the next tile, and so on. The downside is that you'll likely have to compute some data in the boundary regions between tiles multiple times. The upside is that the bulk of intermediate data between post-processing steps never has to be written out to and read back from memory (modern render targets are too large to fit into traditional caches, though that may be different with the huge cache AMD has built for their latest GPUs).
The same principle can be applied to GPGPU algorithms that have similar locality. This tends to be discussed under the label of "kernel fusion".
The page size is 4kb. This means the lower 12 bits of an address is the same between logical and physical addresses. The cache line is 64 bytes. The lower 6 bits of the address are indexing within a cache line. The L1 is 8 way associative so the other 6 bits addresses 8 cache lines in the L1. This makes 64*8 cache lines of 64 bytes => 32k of L1 cache.
The CPU does a lookup in TLB and a lookup in L1 in parallel, and gets 8 cache lines from the L1, which are filtered by the results in the TLB to hopefully get a hit.
Now you'll note that, while most CPU have 32k of L1, the M1 has 128k, which means it needs 2 extra bits to match between physical and logical addresses to pull the same trick. And what do you know, M1 has 16k pages! What a coincidence (not!).
The CPU is said to have 3072 L2TLB entries, so that's not quite a match. Perhaps this "32MB TLB" is a GPU-only TLB; perhaps it does not support variable page sizes.
Which is why I believe they've confused the TLB (Translation Lookaside buffer of the MMU) with the Tile Buffer of a TBDR renderer (IE the on-chip working buffer for in-tile intermediate results of shading).
And it's too bad that apparently the page size can't be changed for larger memory systems to reduce TLB misses (at the expense of fragmentation) which is one traditional solution that doesn't require software changes. 16KB seems like a very small page size on a 64GB system. If we scaled an 8MB VAX with 512B pages up to 64GB it would have a page size of something like 4MB.
This "32MB TLB bottleneck" is a weird thing to say. The TLB size is typically stated in terms of pages, right? with huge pages (or "superpages", as macOS may call them?), that should be a lot more than 32MB of total memory.
I suspect, however, that apple can't really allocate that many (if any at all) 32MB regions of memory at runtime due to fragmentation unless they've substantially changed their contiguous memory allocator since I last looked.
[1] https://developer.arm.com/documentation/den0024/a/The-Memory...
> I suspect, however, that apple can't really allocate that many (if any at all) 32MB regions of memory at runtime due to fragmentation unless they've substantially changed their contiguous memory allocator since I last looked.
That's unfortunate but at least is something they could fix in upcoming software versions. Some folks have put a lot of effort into huge pages on Linux, and it's still ongoing. [1] Not too surprising that macOS could have some room for improvement...
I'm no expert but I suspect the mobile SoCs like Bionic and Snapdragon have the same concept as the M1 Ultra with respect to integrated GPU sharing memory with the apps cores. M1 probably inherited it from Bionic? So some of this work of porting compute software to reflect this environment may have already started. I guess the challenge is that the bar is higher for expectation of GPU performance in a desktop system like the Mac Studio.
But I can say for sure that Intel iGPUs, AMD GCN, AMD RDNA, AMD CDNA, and multiple NVidia-generations of GPUs all have hyperthread-like rescheduling of independent workgroups.
In fact, something like 8x wavefronts / warps run in parallel on modern GPUs. When one wavefront / warp stalls due to a memory read/write (or a PCIe read/write), the GPUs universally "hyperthread-out" and hide the latency.
Its "different" from how CPUs do it, but the fundamental principals are the same. (CPUs have a redundant set of registers tracked in a register file. GPUs on the other hand, have a set of registers and the kernel-scheduler (or whatever handles CUDAstreams) carefully assigns those registers to not conflict with any running wavefronts).
-------
The statement so listed is blatantly false, at least for Intel, AMD, and NVidia GPUs. Maybe Apple iGPUs are built different, but I find that unlikely.
The statement is for Apple GPUs only, that’s the whole point. Software can be easily ported to Metal (in a weekend according to Roblox devs) but until it’s optimised for TBDR it will underperform.
You'd need SMT to do this for memory stalls, and Apple M1 doesn't use SMT - they have the same amount of logical cores (hardware threads) and physical cores.
From the explanation it sounds like an artificial limitation. The chip is technically capable of being saturated with the current design given an optimal program, and fulfilling it's TDP, at which point the current cooling system design would be _entirely necessary_ for the current chip design. This is clear from the second tweet:
> Mac Studio cooling system is OVERKILL in most apps
My emphasis added (not all apps in theory - but practically)
True, but Apple generally do care about either size or quietness. Starting with the G5 cooling system they started to exchange a lot of the former for the latter in their large desktops. I expect that is the case here (i've not actually seen a picture of this thing) - However this doesn't mean they designed the cooling system extra large margins, that's just not their attitude, unlike PC vendors they control all the components they can be more exact (Saying this as quite the opposite of an Apple fan boy).
Point is, the cooling system overcapacity is due to expected (but generally not fulfilled) TDP, rather than adding arbitrary large margins.
Apple for a few decades now has been keeping the same enclosure for at least three product iterations whilst upgrading internals.
So the cooling solution would need to support M2, M3 and maybe even M4.
That shouldn’t produce much heat inside the unit (they claim the power unit can continuously deliver 370W. Guessing 95% efficiency and 200W delivered over those buses it can’t lose more than 10W), but may just push it over some edge.
The tweets refer to Tile-Based Deferred Rendering (TBDR) [1] as compared to Tile-Based Immediate Rendering (TBIR).
Here's another take on it: https://docs.imgtec.com/Architecture_Guides/PowerVR_Architec...
Note that TBIR appears to be a term the author just made up to refer to apps that have not be specifically optimized for the Apple M1 Ultra GPU architecture.
[1] https://developer.apple.com/documentation/metal/resource_fun...
Tile-Based Immediate Rendering, or Tile-based immediate mode rendering (TBIM) is quite well known and not an invented term. [1]
Given this is from MaxTech. I nearly stopped reading when he listed "they started working on the chips 5-7 years ago". But since he quoted hishnash, may be worth a read anyway.
I also doubt the the so called GPU limitation is due to this. The GPU TDP usage depends on a lot of other factors.
This TBDR limitation, if you can call it that has been well known since PowerVR / Dreamcast's days ( Someone please make a low cost Super-H Chip ). Most of the explanation seems very weird from a cache and memory access perspective. And as @lamchester pointed out later [2], a lot of it doesn't make any sense. And remember, M1 Ultra is two M1 Max. Which means there are many some small details missing in how it works together.
[1] https://en.wikipedia.org/wiki/Tiled_rendering
[2] https://twitter.com/lamchester/status/1514346083819786240
I honestly can't make heads or tails of the Twitter thread. It talks about having too many small reads/writes but that's the whole point of tile based rendering. You use the tile as scratch space to do operations that are expensive in system memory.
This approach is well over a decade old at this point and pretty ubiquitous in power constrained rendering scenarios[1]. Best I can tell you need to take architecture into consideration when trying to get the peak GPU usage which has always been true of any platform regardless of tiled or immediate mode rendering.
[1] https://developer.qualcomm.com/sites/default/files/docs/adre...
I think the claim is that the typical working set can exceed what the TLB can map once the GPU scales to 32 or more cores, so the TLB starts thrashing. Optimizing for TBDR would alleviate this because all the tile memory accesses would bypass the TLB, and also likely reduce the working set because the intermediate buffers don't need VM mapping.
Optimizing for TBR involves reducing drawcalls, buffer target switches and operations that are orthogonal to what a TLB historically does.
GPUs don't like to read around in memory(neither do CPUs for that matter but then tend to branch and be a lot less predictable) so optimizing reads/writes are something you will do regardless. The painful parts of TBRs usually(but not universally) have to do with drawcalls setup and render target switching. I remember some of the early iPhones had pretty painfully low drawcalls limits associated with the overhead with dispatching the actual calls as opposed to triangle count or fillrate. It's fairly easy to put together the synthetic benchmarks to shake that out.
(alternatively: it looks like the M1 Pro only has 24 MB L3, so it's likely that actual cache misses become an issue before TLB misses, but the Max/Ultra have more cache so that could invert)
[1] https://twitter.com/VadimYuryev/status/1514295707481501700
However drawcall batching, minimizing render target(and state!) switching are all bread-and-butter performance optimizations you'd make for an immediate mode renderer as well. You can get away with more drawcalls and there are specific things you can do on a per-architecture basis but they don't involve anything having to do with TLBs to the best of my knowledge. If someone wants to spill the beans on how the M1 is somewhat different here I'm all ears but it just doesn't line up with most GPU architectures I've seen.
The big wins between TBR and IMR usually involve getting larger tiles through less usage of rendertargets[1] as it increases the total size of the tiles which drives efficiencies in dispatching the tiles.
[1] https://developer.qualcomm.com/sites/default/files/docs/adre...
Is he popular? I’ve only run across his videos once or twice but I get the feeling our estimations of his technical knowledge match.
TLB size is not where you're cutting costs. If you need to make it bigger so that common workloads can max out the other parts of the chip, it's a no-brainer decision. This isn't a sensible tradeoff.
What I'm confused with is how many silicon engineers are so easily able to make the absolute determination that "This isn't a sensible tradeoff."
The tweets terminology frankly seems weak, I've never even heard of a 32MB TLB - that's insanely huge.
The thread doesn't give any clear indication if their explanation that the GPU sees significant stalls due to waiting on TLB miss has actual data behind it or is pure conjecture based upon observed power usage.
So am I understanding it correctly - if an app is properly optimized - then it's not even an issue?
Seems a bit apocalyptic. Tile based architectures are everywhere in the mobile space. Is it really so inconceivable that macos devs can adjust to it?
It's nothing new.
It is entirely possible that M1 Ultra is bandwidth starved for some workloads, but that can be demonstrated using GPU profiling tools (which Apple readily provides) and not by musing about technical concepts that one barely understands.
> why would a GPU have it's own TLB in the first place?
is because modern GPUs have their own MMUs and have for quite a while. More or less ubiquitous MMUs on GPUs are half the reason why Mantle/Vulkan/DX12/Metal came into being.
As covered in other comments, the thread author has no idea what they're writing about.
If it's a few quick changes and a rebuild, then it's nothing. If it's anything more you'll have an Itanium-like product flop. Maybe. It's also possible that node size and computational headroom relative to the software of the day will make it hard for consumers to notice or care.
Either way this will be interesting to watch.
TBD how hard it is to optimize an existing application to remove this bottleneck. The tweets seem to suggest it's pretty involved. Alternatively, there might be a simple way to use a larger page size (called "huge pages" on Linux, possibly "superpages" on macOS), which would greatly relieve the TLB as a bottleneck, if TLB = translation lookaside buffer. The tweets seem confused enough that I'm not sure.
> If it's a few quick changes and a rebuild, then it's nothing. If it's anything more you'll have an Itanium-like product flop.
Meh, if it means many applications only reach 86W rather than 105W, that's probably not world-ending. Most applications on any platform have some bottleneck that they could avoid with more optimization effort, whether it's the TLB or something else. Users usually don't complain in spite of these bottlenecks. And:
> It's also possible that node size and computational headroom relative to the software of the day will make it hard for consumers to notice or care.
Yeah, people seem pretty happy with these processors.
Maybe the blogger means that the TLB controls a VM working set or footprint of 32 MB?
This honestly makes me want to throw out all my other computers and GPUs that do whatever they do at any performance metric simply because they use so much more energy to do so
Huh. What if it's written in crayon? Is the TLB a problem then?
The M1 Ultra Studio is not what you think it is. It is not Apple's flagship high-performance machine. It isn't. That is still the Intel Mac Pro. The Studio is the Mac Mini update that was rumored 6 months before. The Studio is a Mac Mini, hey just gave it a new name. IOW the Studio is the low-end and entry level machine for professionals.
Also, the current newest generation Apple hardware is not a revolution. Certainly, Apple deserves praise for seamless platform switching, nice work, truly. But these new Macs are only a little bit more performant than the previous generation. Check the benchmarks. Apple's best years were 2010-2012, where each year they doubled the performance of their machines of the previous year. It took another 6 years for performance to double again. The M1, M1 Pro, M1 Max, and M1 Ultra are all marginally faster than the previous generation, the Intel Macs. Except for 2010-2012, which was amazing, this slight increase in performance is entirely typical of new Macs compared to whatever model's previous iteration.
This twitter guy is talking about GPU and games. I realize GPU is important to the computing industry, but not to me. I've been a Mac user since 1989, and even if it was and is included in the machine, I've never once used the GPU. It just sits idle there while I saturate the CPU.
Macs are tools for professionals. Professionals don't play games with their tools. If Apple has somehow offended the gaming community by not catering to their needs, and also boxed the GPU performance in the M1 Ultra, it's really a win-win. Ok, it's a little embarrassing, but regardless, doesn't deserve the expected outrage from the sensationalist reporting. How dare Apple! Right? It's more like, ok, well, that sucks a little, but who gives a shit?
I'd say the overall tone of comments is pretty skeptical.
I expect 3D rendering applications, video editing applications and image manipulation applications to use GPU, as they always have since they first could. Machine Learning uses GPU, but the bigger M1 chips have the Neural Engine for that, so maybe not so much anymore on macOS. I doubt anyone would bother mining crypto on commodity hardware. But when you refer to, "more and more (professional) apps," I don't know what you're referring to. DAWs will only use GPU for displaying the interface, iow, barely perceptibly, not for applying filters or effects on audio. Spreadsheets, accounting software, financial software, municipal software, inventory software won't need GPU. Desktop publishing applications don't need GPU. CAD software doesn't need GPU. Beyond rendering, video editing, accelerated image manipulation, Machine Learning, cryptomining and games, I don't know for what else a GPU can be used.
macOS uses GPU to drive displays and for the GUI, but hardly, and otherwise, not so much. Why would it? How could a GPU make system operating more efficient? Open up your Activity Monitor, under the Window menu open the GPU History window, and you can watch how macOS doesn't really utilize the GPU. It must use it for Quartz screen rendering, but it is such a tiny amount it doesn't even register.
All of those use the GPU since modern UI toolkits are designed to use GPU offload.
> CAD software doesn't need GPU.
Wat.