Arm unveils 7nm Cortex-A76 CPU
anandtech.com
anandtech.com
I understand that ARM sells core designs, and that each company assembles them / designs them in the way suits their performance needs, and then gets them manufactured?
But what does that really mean? What is the "user" doing? Are they arranging them like kids' Lego blocks on a die surface? Dragging and dropping blocks of code? Are they tweaking the voltage settings? Are they combining them in ways that do sequential operations special to them? What's the added step here? Like, why doesn't ARM just create the chips that the end OEMs want?
Is it like cooking or chemistry? What's the closest analogy? I would love to get a more intuitive feel about what chip design is all about.
Thanks!
So you can think of ARM as designing the engine, and the vendors as taking the engine design then adding the rest of the cars features (chassis, electronics, navigation, etc) around it, then building and selling it.
"Why don't the vendors just design their own processor to go with their own peripherals?" you might ask. Well designing a processor is MUCH harder to do well than any of the individual peripherals. There is a lot more going on with a processor and it's at least an order of magnitude more complex than any of the peripherals that go with it. But that doesn't mean it's not done; Atmel designed the AVR architecture and this is the processor architecture that dominates the Arduino world. Atmel sells microcontrollers with their AVR architecture processor and their own peripheral designs as well. But AVR is mostly 8-bit and they can't compete with the performance of ARMs 32-bit designs; which is why you see ARM is most embedded applications.
The design for the processor is made in a program that lets you "lay out" the hardware at the transistor level and specify processes for it (like a lithographic mask) but ARM never carries out these processes, these must be taken care of by a fabrication facility. Oftentimes these vendors don't even have fabs themselves, they just design the rest of the "car" and then have it fab'ed by a third party.
~A̶R̶M̶ ̶d̶o̶e̶s̶n̶'̶t̶ ̶w̶a̶n̶t̶ ̶t̶o̶ ̶d̶e̶s̶i̶g̶n̶ ̶p̶e̶r̶i̶p̶h̶e̶r̶a̶l̶s̶ ̶a̶n̶d̶ ̶f̶a̶b̶r̶i̶c̶a̶t̶e̶ ̶c̶h̶i̶p̶s̶ ̶(̶i̶t̶ ̶d̶o̶e̶s̶n̶'̶t̶ ̶e̶v̶e̶n̶ ̶h̶a̶v̶e̶ ̶a̶ ̶f̶a̶b̶,̶ ̶n̶o̶r̶ ̶d̶o̶e̶s̶ ̶i̶t̶ ̶w̶a̶n̶t̶ ̶t̶o̶ ̶d̶e̶a̶l̶ ̶w̶i̶t̶h̶ ̶3̶r̶d̶ ̶p̶a̶r̶t̶y̶ ̶f̶a̶b̶s̶)̶.̶~ (Looks like i'm totally wrong here, also why doesn't HN have strikethrough comments implemented?). They do what they do - design processors - and they do it really well, so well in fact that vendors are more than willing to pay for a license to use their processors in their own chips so that the vendor doesn't have to deal with the headache of making a really good processor.
And they work very very closely with fabs; you can't really design highish performance cores like their's without working with the fabs. You'd have all sorts of weird bottlenecks, and wouldn't hit a competitive frequency (think under 100Mhz if you didn't take into account fab design rules).
What they don't do is sell predesigned SoCs (outside of dev systems). They give you all of the tools you need to integrate your own SoC so that you can take nearly all of the capital risk.
edited my comment, thanks for the info!
> I understand that ARM sells core designs, and that each company assembles them / designs them in the way suits their performance needs, and then gets them manufactured?
Most consumers of ARM cores are simply interested in integrating them into a larger design; usually an SoC. Not so much in tweaking performance, though certainly they'll choose whether to prioritize performance or power in their application.
> But what does that really mean? What is the "user" doing? Are they arranging them like kids' Lego blocks on a die surface? Dragging and dropping blocks of code?
Pretty much like Lego blocks. Most shops just want an SoC that does X, Y, and Z with Foo requirements. So they grab an ARM core, an HDMI 2.0 RX core, and H.265 core, and glue them together.
Depending on what tools they're using this "gluing" is specified in different ways. You can design it somewhat abstract in a block diagram, where you specify all the cores you want, specify what (virtual) pins from each core connect to what others, and maybe tweak a few parameters on some of the cores. It looks kinda like this: https://www.altera.com/content/dam/altera-www/global/en_US/i...
Or there are even higher level tools that let you specify and hook these things together in a GUI specifically designed for building designs like this. They look like this: https://i.stack.imgur.com/yud23.png
But those and other higher level tools all basically just have a compiler that converts specifications into code (Verilog or VHDL).
That is then handed to another compiler which creates a netlist. Think of this like assembly code but for hardware. It specifies the whole design at the level of logical operations. 2-bit AND here, 4-bit Full Adder there, etc. Finally that netlist is thrown through _another_ compiler which does final place and route. Place and route is where the netlist is converted into the transistors and wires for the actual die, and then rendered out into all the layer masks that will be sent to the fab. (I'm glossing over a few details here. E.g. place and route is actually working from a library of transistor designs for each possible logical gate, pre-designed by the silicon fab that your going to send your masks to.). It sounds simple, but until place and route your design is just an abstract spaghetti of logical operations and connections between. Place and route has to solve an NP-Complete optimization to figure out where, on the physical die, all the transistors are going to go, given a set of constraints (transistors need to be close enough to their neighbors to achieve the performance requirements).
Anyway, stepping back, the design phase is where you specify various parameters for the cores. These parameters typically change things like enabling/disabling features of the core. For example on an ARM core you might disable the math co-processor to save space/power if you don't need it; make the pipelines smaller, etc. These are configuration options; you aren't editing the ARM core's code. The core itself specifies how its code changes depending on configuration options.
The place and router phase is where you have some control over power/performance. You can tell the tools to focus on power if your design needs to be low power. It'll then design the transistors in such a way that it'll use less power but also lose some performance. Or the opposite.
Of course how the cores themselves are designed is important to power versus performance, and ARM cores probably have some configuration options that affect their behavior in that regard.
And then the silicon you target is of the utmost importance for design trade offs. Targeting 7nm fabs versus something larger will usually mean a more power efficient and performant design, but will have _much_ higher up-front costs (the physical masks you deliver to the fab cost tens of millions of dollars or more).
To be clear, this is the level at which _most_ consumers of ARM cores are working at. They're just putting Lego blocks together. The ARM core is just a black box in that regard. ARM delivers the core's "code" to you as a pre-processed netlist. You can't see its real code (you just see a chaotic spaghetti of logical operations). But some companies have more special needs and want to tweak the core in specific ways (e.g. Apple). They'll have special deals with ARM that give them access to the core's source code where they can make custom tweaks. But this is rare.
Some companies have their own IP which they might integrate into a design. Their own cores, coded from scratch. These are coded in Verilog or VHDL, for the most part.
> Like, why doesn't ARM just create the chips that the end OEMs want?
ARM cores get used in TONS of custom silicon. Wifi routers, drones, cell phone chips, cell phone coprocessor chips, etc. So part of it is that ARM just couldn't possibly design and build all these different kinds of chips.
It's also that ARM does what ARM does best: design ARM cores. That's a job that takes an entire company all to itself to accomplish. Anything else is just beyond the scope of their company (for now).
I suppose one way to think about this is to imagine old-school computers. I'm talking about the ones built from TTL logic chips; pre-6502/8080/etc.
ARM is basically selling a virtual "board" with their CPU implemented using those logic chips. You, as the designer, can then connect their board up to other boards to have other functionality you want. A graphics board, a sound board, etc.
The difference between those days and today is that these are all virtual. So after you've plugged the boards together a compiler can come through and optimize everything into a final single "board". Which is actually a set of masks used to fab chips on a single piece of silicon.
And these boards are somewhat abstractly specified, letting you enable and disable whole portions of it and have the design adapt accordingly (disabling instructions/functionality/etc).
This analogy isn't far from the truth, since a netlist is really just a list of logical operations, aka just like TTL logic chips, and their connections.
Modern chip development is simply an evolved form of these primordial design techniques. We've replaced manual place and route with "compilers" and optimization algorithms (the original 6502's masks were _hand drawn_. Engineers crawling over a giant plastic sheets making cuts to draw all the transistors and wires). We've replaced manually specifying netlists with higher level languages like Verilog and VHDL. We replaced TTL level CPUs with integrated CPUs. And then eventually replaced whole boards with SoCs.
So if you want to learn chip design; start from the beginning; history is very illuminating. Transistors -> TTL logic chips -> 6502/Z80 designs -> SoCs.
I've always wondered at the efficiency of the Place and Route step of chip design - I've dabbled in PCB design at the hobbyist level, and any autorouter that I've come across has been complete garbage compared to a person manually solving the puzzle of placing parts and routing a PCB. Is this any different at the SoC level? Are the SoC-level autorouters structured differently? How automated is this step really?
Autorouters are often considered garbage because they are typically run underconstrainted. That is, you didn't give enough constraints to the solver so its output, naturally, ends up as garbage.
The reason for that is multifold, but one of the biggest issues is that people often feel the time it would take to codify their constraints would exceed the time it would take to simply route the board themselves.
Place and Route algorithms have a different history. A) It's nigh on impossible to manually place and route modern chip designs. At least, the entire designs. (Sometimes sections or repeatable blocks will be manually laid out.). There's just too much. So it was necessary for P&R to be "not garbage". B) Chip design has always lived at the bleeding edge, thus requiring a rigorous understanding of the physical constraints that designs can work within. C) The stakes for chip design are higher. A failed board costs maybe a thousand bucks max to re-spin and a couple days (expedited). A failed chip costs millions upon millions and months of time. So, again, chip designers have been forced to have a near complete understanding of physical constraints. They had to build rule checkers to ensure that, 99.99% of the time, if the rules pass, their design will work.
So it's no wonder that P&R has had a distinct advantage over autorouters.
That said, another big advantage P&R has is that ... nobody looks at the layout (where all the transistors and wires ended up). You don't really care _how_ P&R solved the problem. You just care that it did, and that all the rules pass. If they did, and you got the performance/power/whatever you wanted, great. Who cares how it did it.
Where as autorouters, you've always got some layout engineering looking it over going "eehhhh, I remember this one time 10 years ago I routed a design like that and we got rejected at the emissions lab." Which, of course, only occurs because nobody bothered to tell the autorouter to optimize for RF radiation.
> How automated is this step really?
Almost completely. Sometimes you do help P&R along a little. There are implicit boundaries defined by the "modules" that you break your code up into. The P&R uses that knowledge to know that certain chunks of knowledge are grouped together. But designers often also manually place "chunks" of logic, when P&R is having a bit of a struggle on its own. That is to say, they tell P&R "put all this logic in this sector of the die". It's not manually routing, but it's enough to give P&R a break so it can focus its time elsewhere.
And repeated logic, like say 32-bit adders, RAM cells, etc, are pseudo-manually routed. P&R is given a suggested routing, but is allowed to tweak as needed.
EDIT: I will caveat all of this by saying any engineer who has dared look at the output of P&R will tell you P&R is "garbage". They do _crazy_ things. But most of the time, nobody cares, and when you do care, those rough placing constraints I talked about solve most practical problems.
Thanks!
You didn't see that kind of leniency when AMD was releasing slow/inefficient CPUs pre Zen.
Additionally, there's far more to compare that raw performance, especially when looking at these types of CPUs. Actual size (others here have mentioned Apple's processors are quite large due to cache), power draw at idle/peak, etc are very important for mobile CPUs.
In the end, it's not a buying guide, it's an info-dump about a new product.
That's in the world where developers make performance at least their third-highest priority, and a couple hundred MIPS can run the vast majority of apps without regularly lagging. Let me know if you have any ideas to make that world become real.
> Quite literally the two devices are incomparable as Apple is a service with proprietary devices.
They have the same form factor and run largely the same apps. It's crazy to say that iphone and android are incomparable.
iPhones still have significantly faster CPUs, especially single threaded.
The end result is that even though web browsing is typically categorized as a basic computing activity (compared with media production, 3D gaming, or simulations) it actually requires regular CPU upgrades to keep up with more complex/bloated websites. There are obviously other factors like whether or not the system is memory pressured, SSD or not, use of GPU acceleration, and if the Internet connection is adequate.
Worth pointing out you can pretty much replace all the default apps with 3rd party apps, it would definitely be an improvement if you could set them as the default apps though.
"Apple releases new X series chip, 40% faster, internal details unknown" is not really a compelling story.
The AnandTech article says[2] that the new Cortex-A76 can run with a power budget of 750 mW/core, which is exactly half of the Celeron's, then.
[1] https://ark.intel.com/products/91830/Intel-Pentium-Processor...
[2] https://www.anandtech.com/show/12785/arm-cortex-a76-cpu-unve...
Vs the projected 3900 Score of a76 at 3ghz.
ooh, probably because apple acquired a company called pa-semi (300m iirc) which had a bunch of cpu dudes doing power-efficient chips for quite a while...
Also principal designers of the DEC Alpha and the StrongARM, FWIW.
They hired some of the best and most experienced in the business. You get what you pay for.
Software will need to take individual core speed in consideration when scheduling threads if it wants to achieve best possible power consumption at a certain performance level. Today, our software assumes all cores are equal. I'm not even sure these asymmetric cores have the same SIMD widths, for instance.
But yeah, the SoC in the iPhone 5S was crazy fast at the time, plus they have other advantages like Secure Enclave, and a smartwatch with ACTUAL all day battery life. I owned a Sony phone with a Snapdragon 810, which was Qualcomm's first high-end 64-bit chip... and it overheated at the drop of a hat. The phone was more or less unusable in Australian summer. Everything I read at the time indicated that Apple really caught the rest of the industry off guard.
Even Wikipedia puts it pretty bluntly:
"The first 64-bit SoCs, the Snapdragon 808 and 810, were rushed to market using generic Cortex-A57 and Cortex-A53 cores and suffered from overheating problems and throttling, particularly the 810, which led Samsung to stop using Snapdragon for its Galaxy S6 flagship phone."
This is in big part due to them having bigger caches than some server CPUs.
From what I can tell, while cache is still a large portion of the area, GPU die space is quite large now. The CPU performance seems to be due to expanded execution ports leading to more dispatched instructions per cycle, which is generally a good idea as long as one can keep the execution ports fed with a good amount of speculation in other stages of the pipeline.
[1] https://developer.arm.com/support/arm-security-updates/specu...