How true is this? If they're on the money it's an excellent example of a talent retention miss leading to a demonstrable mediocrity in delivery.
How true is this? If they're on the money it's an excellent example of a talent retention miss leading to a demonstrable mediocrity in delivery.
These days I care more about which TSMC process node my chips came from than which company designed them. I need a new computer but I'm waiting until next year because there will be a wave of new CPUs and GPUs coming out with much better performance. Better designs? Maybe a little, but it's really because they're all moving to TSMC N4.
I really hope Pat Gelsinger can save Intel's fab business because we really need another company that can compete in fabs and Samsung isn't doing too hot either.
And even after the mobile revolution shrank the demand for X86 PCs, the cloud revolution further entrched X86 in the cloud.
The fixation on the fab process is bewildering. Yes, it does help, but it is also an optimisation step that is decoupled from and that bears no relevance on the chip design. Yes, the smaller node size also brings the increased density along and an increased number of things that can be whacked into the same sized piece of silicon, but it will not magically improve the overall system performance or result in the linear architecture scalability.
The article is specifically calling out a potentially decreased ROB size in M2 cores, and ARMv9 also potentially not arriving until M3 which are crucial to the speed or software performance. There is absolutely nothing the fab process can do to make SVE2 and matrix instructions automagically appear in lithographic chip designs – those are the «silicon» design time decisions. As we have recently been seeing more and more practical, mainstream use cases of the advanced use of the SIMD instructions at the C/C++/Rust runtime level that bring an order of magnitude level performance gains, having the SVE2 implementation at the ISA level is becoming somewhat critical.
Take OpenSSL as an isolated example. By simply fiddling with the C compiler flags to allow it to use NEON on M1, the sha256 calculation speed-up is 4x for 128 and 256 block sizes, with performance gains quickly tapering off for larger block sizes and resutling in a modest 10% increase only. And that performance increase happens without the involvement of hash functions having been manually optimised for NEON/SVE1.
SVE2 with its variable vector size support could improve performance for larger unit sizes. Perhaps it is the time to spin up a Graviton3 instance and poke around with clang/gcc to see how actually good or faster the SVE2 is.
Because outside of servers where little cores don't exist, 256b ALUs in big cores mean 256b registers in little cores, and Cortex-A510 is way smaller than Gracemont. And then you're giving Samsung another opportunity to screw up big.LITTLE...
And even the server CPUs with SVE are 2x256b except A64FX which is HPC exclusive, so no better than 4x128b.
The purpose of SVE2 is to simplify the writing of the software that exploits the data parallelism, both when that is done manually and when that is done automatically by an autovectorizing compiler.
With SVE2 it should become much easier to deal with data structures where the sizes and the alignments are not multiples of the ALU width and it will also no longer be necessary to write many alternative code paths, to take advantage of any future better CPUs, like when optimizing for Intel SSE/AVX/AVX2/AVX-512.
There are still a majority of programs that do not utilize as frequently as possible the existing SIMD units. With SVE2, their number should diminish.
It is not. A recent paper (https://arxiv.org/pdf/2205.05982.pdf) from Google engineering has compared performance of a vectorised (SIMD) vs non-vectorised implementation of the quick sort in the Highway library as well as the performance difference of the AVX-512 vs NEON/SVE1 implementations. By switching to the SIMD processing alone, the 9-19x speedup has been reported, depending on the SIMD unit size (32/64/128-bit numbers have been sampled and measured up). Even the smallest of the two, the 9x perfomance gain factor, is far from being marginal.
On the SIMD unit size of things, the performance difference between AVX-512 (the average of 1120 Mb/sec has been measured) and NEON implementation (the 478 Mb/sec throughput on average) is 2.4x smaller for NEON/SVE1 largely due to the smaller width of the units of processing. Again, the 2.4x factor is not in the marginal territory.
> What's not marginal is the improvements in power efficiency that come with new process nodes.
And that is an optimisation step, albeit a very important one. However, it will not make a quick sort implementation run 2.4x faster alone.
Even the web browser you are using right now to comment on HN likely makes use of the very same Highway library (Chrome and Firefox certainly do, unsure about Safari) the speedup gains have been reported for. The «overall» browser performance will also improve as the result due to it receiving gains transparently, by simply dropping an optimised implementation into the browser build.
And an optimised QuickSort can also come in handy if one pokes around a large browser history or uses it as a knowledge base, which I do and use it on a regular basis. My browser keeps a uninterrupted record of all visited websites over the last 15+ years and being able to zoom in on a particular time span to find something within that temporal range quickly is important to me. I am almost certain that a sorting of sorts is involved somewhere behind the scenes.
I share your concern about new SIMD instructions not being used. It seems to me we're at an inflection point, though. ISAs such as RISC-V and SVE will enable (properly written) software to benefit from future wider vectors without even recompiling. github.com/google/highway (disclosure: I am the main author) lets you write your code only once, and target newer instructions whenever they are available, with transparent fallback to other codepaths for CPUs.
Given the various physical realities including power efficiency, I believe there will be considerably more SIMD usage within the next few years.
(I suspect any application doing enough quicksort that the 2x speedup is significant, would be even happier going slightly off-core to a coprocessor more specialized in vector processing, like Hwacha. There's plenty of space between "tightly-coupled CPU SIMD" and "GPU" that I think makes more sense than needing to implement 512-bit registers in little cores.)
2.4x difference was, in fact, reported, however I still find it somewhat difficult to interpret the reported results. The processing unit size difference alone and the number of LU's can't account for such a big difference in transfer speeds as the M1 Max that was used in the assessment has a very wide memory bus (256 bit wide for a performance core cluster or 512 bit wide for the entire SoC) as well as unusually large L1-D cache and a large L2 cache, with both caches having deep TLB's. The test set they used could also fully fit into the L2 cache. I have asked the Google engineer a question in a separate thread about what else could influence the observed performance difference but have not received a satisfactory explanation.
The key bottleneck is partitioning. AVX-512 does really well there because it has dedicated compressstore instructions, and it's actually even faster to partition a vector via vperm* (because we only need to do that once, whereas two compressstore are required to partition). So AVX-512 reaches >25 GB/s partition throughput per core; it's instead limited by the memory bandwidth each core can access (around 11 GB/s if a single core is active, less when all are competing for the total "128 GB/s").
By contrast, NEON for example in the M1 has 128-bit vectors. Its "4 vector units" (even if they can actually execute all instructions concurrently, which is not clear to me and unlikely - Intel can also only execute some instructions on certain ports) are definitely not as good as actual 512 bit vectors, because partitioning only has a left and right side, and we don't have enough ILP for each of those to keep 2 vector units busy. Hence NEON reaches 11 GB/s partition throughput. It would seem like this matches Skylake, but no: once a subarray fits into cache, Skylake is freed from the memory bottleneck and is at least twice as fast there (which is a sizable fraction of the total sort time).
Does this help explain the results?
> The test set they used could also fully fit into the L2 cache.
This seems unlikely because we're sorting 8 MB and my understanding is that cores (unless L2==LLC) generally have private, partitioned L2 caches, so 3 MB in the case of M1. Is that incorrect?
It's pretty symmetrical, moreso than Cortex-X2's 4 pipelines; there's analysis that on M1 only some floating point and crypto instructions can't execute on all 4 pipelines. [1] (TP in that table is inverse throughput)
Which means that, for example, byte permutes from tables 256b or less can actually achieve the same throughput on M1 as with Intel's AVX-512, since M1 can sustain 4x 2-register TBL per cycle. And doing the exact equivalent of a 512b vpermb (3 cycle latency, 1/cycle throughput) can be done with 5 cycles latency and 0.33/cycle throughput on M1, via 4x 4-register TBL.
Well, a vpermd in NEON would need an extra MLA to convert indexes, and vpermi2* equivalents fall off a cliff. And Intel still has p01 free, and COMPACT is SVE. But in general, a lot of the parallelism that enables AVX-512 implementations will convert directly into ILP across 128b vectors.
> This seems unlikely because we're sorting 8 MB and my understanding is that cores (unless L2==LLC) generally have private, partitioned L2 caches, so 3 MB in the case of M1. Is that incorrect?
Anandtech [2] measured the same L2 latency up to about 8MB single-core, so regardless of the details, 8MB is a pretty significant cliff on M1. Regardless, RAM bandwidth is ~60GB/s, and unlike Intel, can be just about saturated by a single core.
[1] https://dougallj.github.io/applecpu/firestorm-simd.html
[2] https://www.anandtech.com/show/16252/mac-mini-apple-m1-teste...
huh, that's surprising, that plot indeed looks like a core might be grabbing more than 'its share' of L2, though not all. The 'full random' curve starts creeping up after ~3MB as expected, so the situation seems to be even more complex than "use up to 8MB".
For completeness I'll also measure for 100M elements single core, though on M1 that wouldn't make a difference because as you say, a single core can drive a lot of memory bandwidth, enough that NEON becomes the bottleneck.
Also, there are now several RISC-V CPUs with 512-bit vectors, and it seems fair to call them little cores especially compared to x86 and M1/M2. Perhaps 512-bit is more feasible (and sensible) than is widely believed?
Depends on the applications, I suppose. But did you know that (at least on OoO x86), the energy cost of scheduling an instruction dwarfs that of the actual computation? That is why SIMD, including SVE2, can be so important - it amortizes that cost over several elements. Let's spend (more of) our energy budget on actual work.
Is it really just "very few things that actually start using new SIMD"? I'm not a huge fan of autovectorization, but even that is able to vectorize some fraction of STL algorithms. And there are several widely used libraries, including image/video codecs and encryption, that use SIMD and wouldn't be feasible otherwise.
I would recommend not taking their business conjecture without a giant pinch of salt. Just today they were claiming Apple has lost hundreds of engineers in the chip division. The idea that a single division somehow lost hundreds without the industry noticing is ridiculous.
It's there.
Seems some employees took more than themselves to Rivos. "at least two former Apple engineers took gigabytes of confidential information with them to Rivos."
Apple will eventually be overtaken by another company at some point, but there's a world of journalists and pundits who continue to cry wolf every day.
To quote from that article:
"SemiAnalysis believes that the next generation core was delayed out of 2021 into 2022 due to CPU engineer resource problems. In 2019, Nuvia was founded and later acquired by Qualcomm for $1.4B. Apple’s Chief CPU Architect, Gerard Williams, as well as over a 100 other Apple engineers left to join this firm. More recently, SemiAnalysis broke the news about Rivos Inc, a new high performance RISC V startup which includes many senior Apple engineers. The brain drain continues and impacts will be more apparent as time moves on. As Apple once drained resources out of Intel and others through the industry, the reverse seems to be happening now."
I was very optimistic on Apple on the CPU front until I read this today. Now I'm waiting to see how the A16 pans out for them to see if it's a two generation loss of progress, or just a single generation stumble.
0: https://semianalysis.substack.com/p/apple-cpu-gains-grind-to...
Nuvia started early enough to be a factor here. But Rivos wasn’t even founded until June 2021. To release now, M2 would already have been at finished with design by then.
There is an excellent video on this for anyone interested in Japanese culture and the war against USA via semiconductors:
I think there's always a desire to work at a startup in SV and in a low/zero interest rate environment - VCs could probably fund something in the chip design space.
But now that interest rates are going up, I think that will be a lot tougher and Apple will be a better position due to their direct access to free cashflow - to either compete or acquire them at a later date.
Its also an observation that w.r.t. chip design and consumer electronics, the pay is general lower than say Google, Facebook, Salesforce, Web 2.0 based startups (i.e. AirBnb, Uber, DoorDash), etc.
My presumption is that this is because as a chip designer or embedded software/hardware engineer, the capital costs to do anything interesting on your own as a startup (i.e. tape out a chip, mass production in Asia, etc.) are very very high and very fixed and very up-front. Even fabless semiconductors and factory-less product design companies that outsource manufacturing to Asia would need to go find outside capital for IC masks or HW prototypes. You also need a cadre of supply chain, biz dev, marketing, ad spend, channel sales distribution.
Compare that to AirBnb, Dropbox where you need a good idea, a handful of 10x SW engineers and an AWS account that can scale as you grown and a free tier for onboarding customers. Therefore, Google/FB etc. need to pay more to prevent these folks from going off and starting their disruptor (i.e. Insta, WhatsApp, SalesForce).
The author's argument here about talent leaving after having "gotten Apple off x64" is such an odd take. It's not as if Apple started designing these chips after the M1 launched—the pipeline for even small SoCs is often five or more years. The bit about Rivos is especially bizarre because that company was founded in 2021, well after this chip must have been taped out.
With respect to Rivos, reading the about page - it seems an interesting take on RISC-V.
My take is that this will be rolled back into either Apple or Google at a later date - mostly as a hedge against someone (like Nvidia) acquiring the ARM IP now that its in play - or to provide some realistic alternative that can be used as a counter bid in licensing discussions with ARM.
Two of the founders of Rivos were involved in PA Semi which was acquired by Apple and Agnilux which was acquired by Google ChromeBook team.
Few people are going to upgrade from the M1 to the M2 anyway, so it makes sense to keep powder dry for the M3.
It looks like M2 is neither of those, and it's already 2 year.
And while Apple isn't the max payer in SV, I'm sure they pay fine compared to other big tech. The issue is, chips are big right now and no existing big tech can compete compensation wise with shares in a growth chip startup. With VC drying up, I expect this to change back in Apple (and other big techs) favor.
California has a total ban on non-competes.
A very small handful of other states put restrictions on non-competes, but even those generally allow non-competition agreements if time limited, and the employee makes over ~$100k.
It’s widely accepted that prohibiting non-competes has been a significant factor in the tech industry success in California.
As one example, it is well known that Amazon aggressively enforces non-competes, even against line engineers.
But yes, I wish all states would just ban them outright. Or at least make them require compensation. If an employee is important enough to require a non-compete, then they are important enough to pay during the non-compete time period.
Does CA also ban them as part of an acquisition? I've seen them as part of the sale so everyone doesn't quit the day after the acquisition and start a competitor.
[1], I hope that is not him.
[1] https://www.reuters.com/legal/litigation/apple-lawsuit-says-...