Nvidia Unveils Grace: A High-Performance Arm CPU for Use in Big AI Systems
anandtech.com
anandtech.com
CPU: LPDDR5X with ECC Memory at 500+GB/s Memory Bandwidth. ( Something Apple may dip into. R.I.P for Mac with upgradable Memory )
GPU: HBM2e at 2000 GB/s. Yes, three zeros, this is not a typo.
NVLink: 500GB/s
This will surely further solidify CUDA dominance. Not entirely sure how Intel's XE with OneAPI and AMD's ROCm is going to compete.
It's a good step forward but your average consumer GPU is already around a quarter to a third of that and a Radeon VII had 1000 GB/s two years ago.
- GPU<>CPU/RAM
- GPU<>storage
- GPU<>network
(- GPU<>GPU bandwidth is already insane, as is GPU compute speed)
In the above, they're about cases like logs where there is ~infinite off-GPU data (S3, storage, ...), yet current PCI etc CPU stuff is like a tiny straw clogging it all.
It's now ~easy to do stuff like regex search on GPU, so systems being redesigned to quickly shove 1TB through a python 1-liner is awesome.
To get a feel for where all this is in practice, I did a fun talk w/ the pavilion team for this year's GTC on building graphistry UIs & interactive dashboards on top of this: https://pavilion.io/nvidia/
Edit: A good search term here is 'GPU Direct Storage', which is explicitly about skipping the CPU bandwidth indirection & performance handcuffs. Tapping directly into the network or storage is super exciting for matching what the compute tier can do!
I believe this leaves Apple, ARM, Fujitsu, and Marvell as the only companies currently designing and selling cores that implement the ARM instruction set. That may drop to 3 in the next generation, since it’s not obvious that Marvell’s ThunderX3 cores are really seeing enough traction to be be worth the non-recurring engineering costs of a custom core. Are there any others?
I have no idea what Apple's plans for the M1 chip are, but if they had manufacturing capacity, they could put oodles of these chips into datacenters and workstations the world over and basically eat the x86 high-performance market. The fact that the chip uses so little power (15W) means they can absolutely cram them into servers where CPUs can easily consume 180W. That means 10x the number of chips for the same power, and not all concentrated in one spot. A lot of very interesting server designs are now possible.
With Nvidia, buying Arm and producing their own chip sets, that's no small advantage for companies that are not Nvidia (or Apple who have a perpetual license already). If I were Intel, that's what I'd be looking at right now. Same for perhaps AMD. The clock is ticking on their x86 only strategy and it takes time to develop new architectures; even if you do license somebody else's instruction set.
A counter argument to this would be software compatibility. Most of the porting effort to make linux, windows, and mac os run on Arm has already happened years ago. It's a mature software ecosystem. Software is actually the hardest part of shipping new hardware architectures. Without that, hardware has no value.
And a counter argument to that is that Apple is showing instruction set emulation actually works reasonably well: it is able to run x86 software at reasonable performance on the M1. So, running natively matters less these days. If you look at Qemu, they have some interesting work going on around e.g. emulated GPU where the goal is not to emulate some existing GPU but to create a virtual only GPU device called Virgil 3D that can run efficiently on just about anything that supports opengl. Don't expect to set fps records of course. The argument here is that the software ecosystem is increasingly easy to adapt to new chip architectures as a lot of stuff does not require access to bare metal. Google uses this strategy with Android: native compilation happens (mostly) just in time after you ship your app to the app store.
But a million (est) new general purpose ARM computers hitting the population certainly affects the prioritizing of ARM issues in a bug tracker.
Despite the fact that 64-bit Linux had been running successfully on DEC Alpha systems for years, we ran into no end of difficulty because pointers were truncated all over the place, which apparently hadn't mattered on Alpha systems.
It seems like it must have been an Endian issue, but after 20 years my memories are basically toast. I just know nearly every bug we found was pointer truncation.
How many compilers didn't support ARM?
While there are using some future ARM core, and I've read rumors that future designs might try to emulate what has made Apple cores successful; we cannot say whether Apple designs will stagnate or continue to improve at current rate.
There is potential for competition from Qualcomm after their Nuvia acquisition though.
The 40 core Xeon also costs around 10k.
There's rumors that the new iMac will have a 20 core M1 (16+4). I imagine that will be faster than even the top line $10k Xeon.
I have absolutely no doubt apple could put together a server based on the M1 which would wipe the floor with Intel if they wanted to. But I very much doubt they will since it is so far out of their core competencies these days.
I have absolutely no doubt apple could produce a ridiculously good server CPU from the M1. I doubt they will actually do it though.
Not really, part of why the ARM chips are so good is that the memory bandwidth is so fast. With 40+40 cores you're going to have at least NUMA to contend with, which always hampers multithreaded performance.
No the Apple Silicon chips use the arm _instruction set_ but they do not use their core design. Apple designs their core in house, much like Qualcomm does with snapdragon. Both of these companies have an architectural license which allows them to do this.
That will probably change with their Nuvia acquisition.
It may be they don't want to detract from focus on the GPUs for vector computation so prefer a CPU without much vector muscle.
Also interesting that they're picking up an arm core rather than continuing with their own design. Something to do with the potential takeover (the merged company would only want to support so many micro-architectural lines)?
It's all greenfield and growing so far, they'll win more by having the very best products they can make on both sides.
There was no information whether it will have any good SVE2 implementation. On the contrary they insisted only on the integer performance and on the high-speed memory interface.
I'd suspect NVidia would be using the V1 here as it's the higher performing core, but not way to be certain.
"E" is efficiency, N is standard, V is high-speed. IIRC, N is the overall winner in performance/watt. Efficiency cores have the lowest clock speed (overall use the least amount of watts/power). V purposefully goes beyond the performance/watt curve for higher per-core compute capabilities
One of the reasons why M1 is good is pure and simple that it has a pretty enormous transistor budget, not solely because it's ARM.
It's also very hard to achieve more than 4X parallelism (though I think Ice Lake got 6X at some additional cost) in decode, making instruction level parallelism harder. X86's hack to get around this is SMT/hyperthreading to keep the core fed with 2X instruction streams, but that adds a lot more complexity and is a security minefield.
Last but not least: ARM's looser default memory model allows for more read/write reordering and a simpler cache.
ARM has a distinct simplicity and low-overhead advantage over X86/X64.
Furthermore, the high-performance ARM designs, starting with the Cortex-A77, started using the same trick---the 6-wide execution happens only when instructions are being fed from the decoded macro-op cache.
I have vTune installed so I guess I could investigate this if I dig out the right PMCs
0 lsd_uops
1,092,318,746 idq_dsb_uops ( +- 0.49% )
4,045,959,682 idq_mite_uops ( +- 0.06% )
The LSD is disabled in this chip (Skylake) due to errata, but we can see only 1/5th of the uops come from the uops cache. However, the more relevant experiment in terms of power is how many cycles is the cache active instead of the decoders: 0 lsd_cycles_active
378,993,057 idq_dsb_cycles ( +- 0.18% )
1,616,999,501 idq_mite_cycles ( +- 0.07% )
The ratio is similar: the regular decoders are not active only around 1/5th of the time.In comparison, gzipping a 20M file looks a lot better:
0 lsd_cycles_active
2,900,847,992 idq_dsb_cycles ( +- 0.07% )
407,705,985 idq_mite_cycles ( +- 0.33% )Forget Bitcoin mining... how many tons of CO2 are released annually decoding the X86 instruction set?
I’d say ARM has a big advantage for instruction level parallelism with 32 registers.
And it seems to me that ARM has an advantage here. If you want execute 8 instructions in parallel, you gotta actually have 8 independent things that need to get executed. I guess you could have a giant out of order buffer, and include stack locations in your register renaming scheme, but it seems much easier to find parallelism if a bunch of adjacent instructions are explicitly independent. Which is much easier if you have more registers - the compiler can then help the cpu keeping all those instruction units fed.
> include stack locations in your register renaming scheme
Registers aren't related to the stack. "The" stack is just RAM being accessed in a specific cache friendly pattern, with additional optimizations (if you use specific registers) from the hardware in the form of the stack engine. The compiler explicitly loads and stores to and from the registers named by the ISA. Register renaming has absolutely nothing to do with the stack.
When the CPU can tell that a later instruction doesn't depend on the previous value of a register, it's free to rename it. The result is that two independent registers get used even though only one was ever directly referenced. In reality, there are a _huge_ number of registers available on modern processors. Estimates place Skylake, Zen, and Cortex-X1 at 200+, with the M1 at 600+. The ISA just doesn't provide a way to access them directly. (If you want to read about this, the term to look up is reorder buffer.)
Also, there is a giant out of order buffer for stores waiting to be written back to L1. That buffer does indeed have to keep track of cache locations, which directly map to memory addresses, which sometimes happen to refer to stack locations. So in a sense, what you suggested already exists. (If you want to read about this, the term to look up is store buffer.)
> it seems much easier to find parallelism if a bunch of adjacent instructions are explicitly independent
That would indeed make things simpler in some cases. However, many operations such as loading a value into a register (ex mov, [addr]) or zeroing it (ex xor eax, eax) explicitly break the dependency chain by definition. Cases where the CPU fails to properly account for this are documented as false dependencies.
> the compiler can then help the cpu keeping all those instruction units fed
The "compiler handles ordering" thing was tried with Itanium. It seems it didn't go so well.
The CPU is free to simultaneously load two different pieces of data into the "same" register and execute two independent instruction streams on that "single" register thanks to renaming. Speculative execution helps when the CPU can't be completely certain that there isn't a dependency.
For particularly complicated sequences, the compiler spilling due to running out of named registers could indeed pose an issue. However, the CPU is free to elide a store followed by a load if it determines that the address is the same. (If you want to read about this, terms to look up include store-to-load forwarding and load-hit-store.)
I know Itanium didn’t work - but that’s because here the compiler is supposed to do all the reordering work. That’s different from allowing the compiler to explicitly define that instructions are independent by having more registers.
Although apparently Zen 2 changed this and can pull off zero latency. (https://www.agner.org/forum/viewtopic.php?t=41)
Some general background: (https://travisdowns.github.io/blog/2019/06/11/speed-limits.h...)
a = m[i+1] + b
c = m[i+3] + c
e = m[i+7] + d
assume you only have 3 registers, in a RISKy architecture. Every statement becomes something like r1 = *pb // load c
r2 = r0[1] // m[m+1]
r1 = r1 + r2 // a = ...
*pa = r1
Since all registers are used, and all but two instructions are dependent, in the assembly the blocks have to follow one another. There`s also spilling of the b,c,d variables, they have to be read from registers (which could be elided). Assuming no re-order buffer, these instructions runs in three cycles (the first two are independent) - even though the top level instructions are independent.If you want them to run all statements with 4 instructions at a time, you need to have a reorder buffer that covers the whole sequence (12 instructions). (Imagine if b,c,d get modified inside the inner loop and spilled into memory, you have to track memory locations in order to do register renaming.)
Now lets assume you have 6 registers. Now all variables fit in registers and the compiler can easily interleave the code giving a sequence of 3 or 4 independent instructions at a time. If you want to run 4 instructions at the same time, you need no reorder buffer.
This is a kind of specific example, but it shows that if you have more registers (i.e. ARM vs x86), the compiler can more easily interleave instructions, which can help reduce the number of instructions that need to be in the reorder buffer. Or with the same size re-order buffer, its easier to find more independent instructions and keep all the execution units fed. Or, when jumping to some code thats not in pipeline or icache, it allows to sooner run more instructions in parallel, when only a small number of instructions are decoded and in the re-order buffer.
In practice, x86_64 works just fine for HPC number crunching code. Outside of some serious number crunching, when are you going to have more live values than named registers, have instruction streams whose output depends on _all_ of those values (which is why they would be live), and also those streams complete so quickly that you stall on the next set of loads? And you have absolutely no other useful work to do? Honestly I think you're being silly.
Historically, I understand that the 32 bit version of x86 did have scheduling challenges surrounding function calls. The 64 bit version of the ISA expanded the number of named registers and (as far as I understand things) it largely resolved the issue.
Also note that typical hardware can sustain a surprisingly large number of loads per clock. You just need to find something useful to do while you wait for the load to complete. In case you really can't there's also SMT. Really though, the PRF and ROB are only so large.
> If you want to run 4 instructions at the same time, you need no reorder buffer.
You always need a reorder buffer if you want to achieve good performance. Among other issues, the compiler can't predict the latency for each load in advance due to caching behavior depending on the runtime state of the full computer system. I previously mentioned Itanium. It's directly relevant here.
> Imagine if b,c,d get modified inside the inner loop and spilled into memory, you have to track memory locations in order to do register renaming.
No. You can't just rename registers any longer. A store to memory means the memory model for the ISA gets involved. Things become significantly more complicated. The store buffer exists specifically to deal with such issues efficiently on an OoO core. Seriously, go read about it. It's astoundingly complicated for any OoO core regardless of the ISA.
> the compiler can more easily interleave instructions, which can help reduce the number of instructions that need to be in the reorder buffer
Unless I have a serious misunderstanding (I don't design hardware, so I might) everything passes through the reorder buffer. Every instruction is speculative until all previous instructions have retired. (https://news.ycombinator.com/item?id=20165289)
What percent of the die is an ARM instruction decoder?
ARM A32/A64 instruction decoding is dramatically simpler -- all instructions are 32 bits wide and word-aligned, so decoding them in parallel is trivial. T32 ("Thumb") is a bit more complex, but still easier than x86.
What nearly everyone uses is a 16 byte buffer aligned to the program counter being fed into the first stage decode. This first stage, yes has to look at each byte offset as if it could be a new instruction, but doesn't have to do full decode. It only finds instruction length information. From there you feed this length information in and do full decode on the byte offsets that represent actual instruction boundaries. That's how you end up with x86 cores with '4 wide decode' despite needing to initially look at each byte.
Now for the efficiencies. Each length decoder for each byte offset isn't symmetric. Only the length decoder at offset 0 in the buffer has to handle everything, and the other length decoders can simply flag "I can't handle this", and the buffer won't be shifted down past where they were on the next cycle and the byte 0 decoder can fix up any goofiness. Because of this, they can
* be stripped out of instructions that aren't really used much anymore if that helps them
* can be stripped of weird cases like handling crazy usages of prefix bytes
* don't have to handle instructions bigger than their portion of the decode buffer. For instance a length decoder starting at byte 12 can't handle more than a 4 byte instruction anyway, so that can simplify it's logic considerably. That means that the simpler length decoders end up feeding into the higher stack up full decoder selection, so some of the overhead cancels out in a nice way.
On top of that, I think that 5% includes pieces like the microcode ROMs. Modern ARM cores almost certainly have (albeit much smaller) microcode ROMs as well to handle the more complex state transitions.
Once again, totally agreed with your main point, but it's closer than what the general public consensus says.
IMO, you would either go towards bitaligned instructions like the iAPX 432 or the Mill, or 16-bit aligned variable width instructions like the s360 and m68k on the CISC side, and ARM Thumb and RV-C on the RISC side.
That being said, you're definitely thinking about it the right way. Modern Istream bandwidth conscious ISAs absolutely (and perhaps unsurprisingly) look at the problem from a constrained, poor man's huffman encoding perspective similar to how UTF-8 was conceived.
Maybe one could come up with an instruction encoding that encodes some number of instructions per cache line. Every time the cpu jumps to a new instruction (at cache line address + index), the whole cache line needs to be loaded into icache anyway, and could get decoded then -> internally they get represented in microcode anyway.
This is also not a good security property since it means you can hide secret instructions in a program by jumping into the middle of innocuous ones.
> ARM A32/A64 instruction decoding is dramatically simpler -- all instructions are 32 bits wide and word-aligned, so decoding them in parallel is trivial. T32 ("Thumb") is a bit more complex, but still easier than x86.
A64 doesn't have a Thumb equivalent, also, and supporting A32/T32 is optional.
I'm not familiar with how ARM's memory model effects the cache design - Source?
There's a lot of brute force, yes, but it's not the only reason. There are lots of smart design decisions as well.
Plus, most of the last decade software is software that runs on some sort of VM or another (be it JVM, CLR, a Javascript engine or even LLVM).
Soon (in years), x86 will only be needed by professionals that are tied to really old software. And those particular needs will probably be satisfied by decent emulation.
There's also that PC & console gaming markets, which are not small and have not made any movements of any kind towards ARM so far.
For example, the M1 has 128 bit wide memory. This has been standard for decades on the desktop(dual channel), but unheard of in cellphones. The M1 also has similar amounts of cache to the new AMD and Intel chips, but thats several times more than the latest snapdragon. Qualcomm also doesn't just design for the latest node. Most of their volume is on cheaper, less dense nodes.
The M1 is one "node" ahead. Apple forked out the cash to get all their chips on TSMC's 5nm process. This is about 2 years of advancement over the 7nm process AMD pays TSMC for. Intel's latest 10nm node is similarly behind TSMC 5nm.
Semiconductors are tricky. Small performance gains take large increases in power. If you play with overclocking, you'll learn power increases quadratically or even cubically with clocks. The mere "2nm" shrink may seem inconsequential, but for these iso-perfomance comparisons(performance@constant-thermals), it is key.
All this to say, you get what you pay for. Chips can get the same performance on TSMC's 5nm node while using 70% of the power as chips on the 7nm node.[1] Compared to TSMC's 10nm (similar to Intel's popular 14nm still in production), 5nm chips can be expected to use ~45% of the power.
Hopefully that shed some light on the M1's biggest advantage for you.
[1] https://images.anandtech.com/doci/15219/wikichip_tsmc_logic_...
The M1 isn't necessarily a win for Arm in general. Other manufacturers weren't competing before and its yet to be seen if they will.
By going on-package there's almost certainly latency advantages in addition to the much-vaunted bandwidth gains.
That's going to pan out to better perf, and likely better power usage as well.
Currently Apple is the only company making performance-competitive ARM cores that can make a reasonable justification for an architecture switch.
Otherwise AMD's CPUs are still ahead of everyone else, including all other ARM CPU cores not made by Apple. And even Intel is still faster in places where performance matters more than power efficiency (eg, desktop & PC gaming)
upd: oh also in the HPC world, Fujitsu with the A64FX seems to be like the best thing ever now
But it can't just be competitive it needs to be significantly better in order for the consumer space to care. Nobody is going to run Windows on ARM just to get equivalent performance to Windows on X86, especially not when that means most apps will be worse. That's what's really impressive about the M1, and so far is very unique to Apple's ARM cpus.
> oh also in the HPC world, Fujitsu with the A64FX seems to be like the best thing ever now
A64FX doesn't appear to be a particularly good CPU core, rather it's a SIMD powerhouse. It's the AVX-512 problem - when you can use it, it can be great. But you mostly can't, so it's mostly dead weight. Obviously in HPC space this is different scenario entirely, but that's not going to translate to consumer space at all (and it's not an ARM advantage, either - 512bit SIMD hit consumer space via x86 first with Intel's Rocket Lake).
If x64 ISA had major advantages over Arm then that would be significant, but I've not heard anyone make that case: instead it's a debate about how big the Arm advantage is.
Can x64 remain competitive in some segments: probably and inertia will work in its favour. I do think it's inevitable that we will see a major shift to Arm though.
but one factor that you can replicate is colocating memory, CPU, and GPU, the system-on-chip architecture. that's what Nvidia looks to be going after with Grace, and I'm sure they've learned lessons from their integrated designs e.g. Jetson. very excited to see how this plays out!
Not really, they are still just using the same ARM ISA as everyone else. The only hardware/software integration magic of the M1 so far seems to be the x86 memory model emulation mode, which others could definitely replicate.
> but one factor that you can replicate is colocating memory, CPU, and GPU, the system-on-chip architecture.
AMD introduced that in the x86 world back in 2013 with their Kavari APU ( https://www.zdnet.com/article/a-closer-look-at-amds-heteroge... ), and it's been fairly typical since then for on-die integrated GPUs on all ISAs.
The current fastest supercomputer uses ARM.
More broadly, as to why the ISA doesn't make a big difference: The major differences are at the microarchitecture level since OoO processors have such flexible dataflow machinery in them that you can kind of view the frontend as compiler technology. x86 and ARM are decades-old ISAs that have seen a many many rounds of iteration in form of added instructions and even backwards incompatible reboots at the 64-bit transition points so most hinderances have been fixed.
In the olden days ISAs were important because processors were orders of magniture simpler, and instructions were processed as-is very statically (to the point that microarchitectural artifacts like branch delay slots were enshrined in some ISAs). This meant that eg the complexity of individual instructions could a bottleneck to how fast a chip could be clocked. Or in CISC land your ISA might have been so complex that the CPU was a microcoded implementation of the ISA and didn't have any hardwired fast instructions...
The near future. A few years out, RISC-V is gonna change everything.
The magic of Apple's M1 comes from the engineers who worked on the CPU implementation and the TSMC process.
The architecture has some impact on performance but I think it is simplicity and and ease of implementation that factors most into how well it can perform (as per the RISC idea). In that sense Intel lags for small, fast and efficient processors because their legacy architecture pays a penalty for decoding and translation (into simpler ops) overhead. Eventually designs will abandon ARM for RISC-V for similar reasons as well as financial ones.
Really, today it's a question of who has the best implementation of any given architecture.
Wait, Nvidia's been making ARM CPUs for years now; most memorably Project Denver.
But another reason they won't do it is that TSMC has a finite amount of 5nm fab capacity. They can't make more of the chips than they already do.
They have interconnects from Mellanox, GPUs and their own CPUs now.
I suspect the supercomputing lists will be dominated by NVidia now.
Speaking in general terms, data rate and transaction rate don't necessarily match because a transaction might require the transmitter to wait for the receiver to check packet integrity and then issue acknowledgement to the transmitter before a new packet can be sent.
Yet another case, again, speaking in general terms, would be the case of having to insert wait states to deal with memory access or other processor architecture issues.
Simple example, on the STM32 processor you cannot toggle I/O in software at anywhere close to the CPU clock rate due to architectural constraints (to include the instruction set). On a processor running at 48 MHz you can only do a max toggle rate of about 3 MHz (toggle rate = number of state transitions per second).
PCIe has the optional "relaxed ordering" feature, allowing sending new packets before the ACK has been received from preceeding ones. Not sure precisely how this works, if there is some TCP-like window scaling algorithm in play or not..
NVIDIA is buying ARM.
Multiple competition investigations permitting.
Will they auto-detect workloads and cripple performance (like the mining stuff recently)? Only work through special drivers with extra licensing feeds depending on the name of the building it is in (data center vs office)?
Still, every company does it differently.
For example, both NVIDIA and AMD compute GPUs are necessarily more expensive than gamer GPUs because of hardware costs (e.g. HBM).
However, NVIDIA gamer GPUs can do CUDA, while AMD gamer GPUs can't do ROCm.
The reason is that NVIDIA has 1 architecture for gaming and compute (Ampere), while AMD has two different architectures (RDNA and CDNA).
I don't think they even have display ports.
Not sure what good would a "gaming driver" do you on those cards.
Same for the opposite. Do the RDNA GFX cards have even hardware for compute? They don't even have tensor cores, so why would AMD invest money into creating a compute driver for hardware that's bad at compute?
> I'm only talking about driver/license locks,
Not "locked" is a big understatement. A driver release for some hardware needs at least some QA, so the assumption that doing this is just "free" because its software is incorrect.
Nvidia detects mining workloads in software based on heuristics and disables them. Probably causes more support burden than less and took extra engineering time to implement, not less.
Given how well that turn out, I have a hard time believing they put any effort into this.
Apple is also I think going to soldered on / close in RAM. Nvidia looks to be doing this two CPU / GPU / Ram all close together and it doesn't look like any upgrade options. Some thinking was that Apple was continuing to increase durability / reliability etc with their RAM move.
Does anyone know requirements for the LPDDR5X type of ram mentioned here. Does this require soldering things (you obviously get lots more control if you spec chips yourself and solder on)?
ARM is more of a tool kit to build different purpose built computers (you even see them show up in usb sticks). While x86 is particular ISA that has a long history behind it. So you may see something like 'Amazon builds its own ARM computers'. That means they spun their own boards, built their own toolchains (more likely recompiled existing ones), and probably have their own OS distro to match. Each one of those is a fairly large endeavor to do. When you see something like 'Amazon builds its own x86 boards', they have shaved out the other two parts of that and are focusing on hardware. That they are building their own means they see the value in owning the whole stack. Also if you have your own distro means you usually have to 'own' building the whole thing. So I can go grab an x86 gcc stack from my repo provider. They will need to act as the repo owner and build it themselves and keep up with the patches. Depending on what has been added that can be quite the task all by itself.
I understand the here-and-now AI applications. But this is smelling more like Big AI Hype than Big AI need.
Nvidia +4.68%,
Intel -4.65%
AMD -4.47%There is bottled demand because Intel's failure to deliver was not fully anticipated by anyone.
I doubt anyone really deliberately sets out to be like "haha yessss today I shall elide this woman's credentials", but this is one of those unconscious gender-bias things that is commonplace in our society and is probably best to try and make a point of avoiding.
https://news.cornell.edu/stories/2018/07/when-last-comes-fir...
https://metro.co.uk/2018/03/04/referring-to-women-by-their-f...
(etc etc)
I'd prefer they used "Hopper" instead, in the same way they have chosen to refer to previous architectures by the last names of their namesakes (Maxwell, Pascal, Ampere, Volta, Kepler, Fermi, etc). I'd see that as being more professionally respectful for her contributions.
But yes I very much like the idea of naming it after Hopper.
Vaguely related: J. K. Rowling's "real" full name is Joanne Rowling. The publisher "thought a book by an obviously female author might not appeal to the target audience of young boys".
There's another famous (in the UK at least) computer scientist called Hopper: Andy Hopper. So "G.B.M. Hopper", perhaps? That would have more gravitas than "Andy"!
I guess I'm not sure if "Hopper" refers to the product as a whole (like Tegra) and early leakers misunderstood that, or whether Hopper is the name of the microarchitecture and "Grace" is the product, or if it's changed from Hopper to Grace because they didn't like the name, or what.
Otherwise it's a little awkward to have products named both "grace" and "hopper"...
Unfortunately, at least in most Western societies, using the first names is the only way to refer unambiguously to women.
According to the tradition, in most Western countries the women do not have their own family names, but use either the family name of their father until marriage, or the family name of their husband after that.
So while Grace is the computer scientist, Hopper is her husband and Murray is her father. Using the name Grace makes clear who is honored.
Nowadays, in many places there are laws that allow women to choose their family names or to combine the family names.
Nevertheless, the old tradition is still entrenched, so searching for a certain woman, when the last information about her is many years old, can be difficult due to unpredictable family name changes.
Ideally, a human should keep forever the family name used at birth and the parents should choose one of their family names for the children.
Ideally, a human should keep forever the family name used at birth and the parents should choose one of their family names for the children.
I prefer the Spanish way, have two family names. We have been doing it for centuries, it baffles me that other countries find it so difficult to adopt a similar system.If other companies don't make genuine investments in ARM for the desktop there's a real chance that Apple will get a huge an difficult to assail application performance advantage as application developers begin to focus on making Mac apps first, and port to x86 as an afterthought.
Something similar happened back in the day when Intel was the de facto king, and everything on other platforms was a handicapped afterthought.
I wouldn't want to have my desktops be 15 to 30% slower than Macs running the same software, simply because of emulation or lack of local optimizations.
So I'm really looking forward to ARM competition on the desktop.