Arm Announces Armv9 Architecture: SVE2, Security, and the Next Decade
anandtech.com
anandtech.com
https://www.servethehome.com/amd-psb-vendor-locks-epyc-cpus-...
https://blog.cloudflare.com/anchoring-trust-a-hardware-secur...
I don't intend this as a criticism, just a fact. Newer cores cost more so RPi has to use older technology to meet its price target.
I ultimately think you're right, and we won't see a V9 in an RPi for a while, but it's a more complex situation than most SoC integrators that has a slight chance of working out in RPi's favor. Does ARM care enough to make a tiny core in their gate count niche? Does ARM want to give the free new cores to increase market share and get V9 features in the hands of tinkerers? etc.
edit - the microcontroller was called Pico, not Nano
If there's any truth to this rumor it's probably ARM discounting the license fee for the fraction of devices that Broadcom sells to the Raspberry Pi foundation.
ARM knows that RPi is the go-to ARM SBC, and they benefit tremendously when developers treat it as a first-class platform.
Apple pays ARM for an architectural license, just like any other company.
ARM was created as a joint venture between Apple, Acorn and VLSI.
At that point, yes Broadcom is ultimately paying the fees to ARM, but separating RPi from that negotiation is an oversimplification. RPi absolutely has a seat at that table.
And ARM probably doesn't care that much about it being the go to SBC (some cheap chinese board would take on that mantle without RPi leading the charge), but instead that it's the practical successor to what ARM was founded to do. Put passable performance, hackable computers designed by Brits in front of British school children as cheaply as possible.
The chance they'll have stuff using RISC-V in 2027 is definitely above zero. It's easy to forget that the Raspberry Pi Foundation have been RISC-V Foundation members for years.
The sad reality is that the faster ARM cores don’t trickle their way into mainstream ARM SoCs very quickly. The fastest ARM chips go into cellular phones, set top boxes, and other mass produced electronic goods from vendors who can afford to implement them.
I hope this changes in the coming years as more chipmakers embrace mainline Linux and open source drivers instead of relying on the old models.
Aren't there any Intel Celeron or some such super cheap x86 small boards? I imagine you'll get a ton more oomph, probably much better software support and they shouldn't cost that much more.
It's a good target if you want to provide something very close to a "turn-key appliance" without actually selling your own hardware: You can provide an image that a user writes to a microSD card, plugs in, and it just works.
For example: HomeAssistant [1] for home automation, OctoPi [2] 3D printer software, OpenElec [3] media center, RetroPi [4] game console emulator, Volumio [4] music player, and lots more [5].
[1] https://www.home-assistant.io/installation
Better yet - they boot from USB now (and even NVMe if you're using a CM4).
The biggest downside is your PXE server going down is going to kill all the Pis, so if you are doing something important you need a highly-available PXE setup. And if you have that, it's probably a better place to host a lot of the services people use Pis for in the first place.
I've setup a small office with 5 lines + 10 extensions all running on one of those (+ an analog line adapter), and it's been trouble free for the 3-4 year uptime so far.
Sadly, as our usage of the IP phones dwindled (cell phones, texting, and now conference software taking over), my significant other began evicting them from around the house ("too big and ugly"). A couple years ago, my instance was at the point I needed to do some major OS upgrades to it, and considering it was basically my office phone and a cordless left, I ended up retiring the Pi and PBX. Now I just register my last couple phones directly to my VoIP provider.
It looks like a Pi 4 can handle "dozens" of simultaneous channels (though depends on codecs used), and that would be my starting point if I got back into doing PBX work or deploying a system today -- though I'd probably run it from an SSD instead of SD card.
This is a great example of the benefits of the Pi hardware, too: People don't tolerate PBX downtime. With the Pi, it's incredibly cheap to have hardware for a failover server (try pricing that out on a commercial PBX!), and in an emergency you can run out and buy one locally or have it shipped overnight.
It's not mysterious. Making and supporting Raspberry Pis is the Raspberry Pi Foundation's thing. Wandering around in a confused state is Intel's thing.
The only one I'm aware of that's in the same ballpark is the Atomic Pi, which was apparently a limited production run using heavily discounted surplus components. It's not as popular as you'd expect given the $40 price point, which I assume is because there isn't anything like the same level of community support that the Raspberry Pi has.
Used Celeron NUCs are not more expensive, and they are way more powerful.
To me, they aren't always a better fit than an Rpi, but they often are.
Well, it's a Pentium 4 with 2GB DDR2 RAM... You may wish to search around a bit.
https://www.ebay.com/itm/Dell-Optiplex-745-PC-SFF-Desktop-In...
HP T620 workstation ~$45 plus shipping?
https://www.ebay.com/itm/HP-T620-Dual-Core-AMD-Gx217Ga-1-65G...
[1] https://hejdom.pl/blog/22-home-assistant/208-home-assistant-...
We can easily compare that on Linux with benchmarks. I'd say at the same price point for used x86, x86 is still much faster.
Cheap, robust, solid support and tools. Definitely hit the sweet spot for what we needed.
Raspberry Pi 3 was a Cortex-A53 back ported to 40nm outright.
Raspberry Pi 4 is on 28nm. (for comparison, Apple A8 and Snapdragon 810 were on 20nm already, in 2014-15)
It runs with a Cortex-A72 at a low clock (1.5GHz, phones even are at twice that clock nowadays) and quite high power consumption because of the process node.
The memory interface is narrow (32-bit bus, phones ship with a 64-bit bus and laptops/desktops with a 128b one) at a low data rate (LPDDR4-3200).
In addition to that, the CPU can only use 5GB/sec of it, with the remainder being reserved for the GPU only.
Those were some of the sacrifices needed to reach this price point. (and why it isn't representative at all of the performance of higher-end ARMs)
At the same time Intel was heavily subsidizing their x86 mobile SoCs to tablet manufacturers and for a while it was actually quite competitive price wise but it ultimately went bothered.
I have a _very old_ PN40 from Asus -- this is a full PC with an intel Celeron, a SATA SSD and about 8GB of RAM that idles at 1.7W (serving websites via GB Ethernet) as measured _at the wall_. This is just standard PC Linux distro with zero customization (other than `powertop --auto-tune`).
For comparison, the latest RPI4B without the SSD and with 1/8 RAM idles at 3.5ishW on the wall. Even the older RPI3 I could never get below 2W, and that is without any peripherals (no display, no Wifi, no eth, min USB) and significant tuning.
On the other hand the PN40 + components cost significantly more (probably close to 10x), and the CPU performance itself is not that good these days.
https://raspi.tv/2019/how-much-power-does-the-pi4b-use-power...
- The Odroid H2+ x86 board claims about ~4W idle power: https://www.hardkernel.com/shop/odroid-h2plus
- This measurement of idle power on the Raspberry Pi 4 says it uses ~3.4W: https://www.tomshardware.com/reviews/raspberry-pi-4
in the past as a relatively cheap x86 CPU. ($9.62 per unit)
However, it's a 400MHz 486 with some backports (1st generation Pentium, no MMX even) and broken LOCK prefix so that software specifically needed to be recompiled for it.
And at 2.2 W too, it never ended up succeeding anywhere, with no successor.
The right era pentium had the same LOCK prefix issue (we used to call it the F00F bug)
Edit: Thanks for link; I must have anchored on the instruction set and ignored that it was using the 486 pipeline.
"The new x86 CPU uses an old-school 486 scalar pipeline and the original Pentium instruction set with some modern enhancements"
Not quite direct competition, but pretty close. Maybe more directly similar to the Pi 4CM with a daughter card for NVMe, network, sata, etc.
Any possibility to transition from enterprise/cloud app dev, but with no comp Eng background?
Hardware-adjacent looks more interesting to me - would you say that I’m right or I’m wrong? :-)
The Pis really can do a lot too. Mine hosts a vpn, pihole, calibre-web, and a couple of other basic things. Really the only downside to the ARM based SBCs is they mostly use MicroSD cards, which probably are the greatest bottleneck in the system and mostly likely thing to fail.
As a dedicated, small, cheap computer for the “brains” of various hobbies and certain professional projects (at least where reliability isn’t the top concern)? Absolutely.
It’s by far the popular $35 Linux box out there. People create distros for all sorts of use-cases. Off the top of my head:
Pi-hole, RetroPi, Volumio, OctoPrint, Home Assistant, piAware, Plex/Kodi, many more.
Raspberry Pi needs to do two things to become really useful.
1) ship chips with native-AES/hardware crypto
2) get rid of SD and have onboard NAND, like a phone
I’ve raised this on their forums but just get flamed for some reason by senior engineers suggesting that these ideas are ridiculous and it’s for education in Africa, I was even banned for suggesting they were shortsighted.
Maybe education’s where it started, I don’t see why they are blind to the reality that most current users are tinkerers and Linux hobbyists.
Oh dear. That will involve the process of flashing OS onto the Pi and will increase the risk of bricking it. That is the reason for using SD cards instead of a NAND chip.
Secondly, you can't upgrade the storage on it either and would have to choose a RPi with fixed storage space on it. Might as well get a M1 Mac Mini.
I cannot imagine having to choose a future RPi 5 having either a 8GB, 16GB, 32GB NAND' and being unable to upgrade the space on it. So no thanks and no deal to (2).
The SD cards are nice for most use cases, but on some boards I'd really like to see a SATA port or M.2 slot which can be booted from and a x4 PCIe port (even if through an optional board or a HAT).
I see the SD cards on the Raspis as modern floppy discs. They are good to have, but not ideal for some cases.
People have already achieved it by hardware hacking
My hope is the next Pi model B might include an M.2 slot (maybe 42mm long) on the bottom.
Which ones are you seeing with M.2 slots? I've clicked a good 5-8 of the top listings from that list for which it would make sense to have such slot and they're missing it.
But that being said, I've burned out plenty an SD card. Haven't managed to do that to an eMMC module yet.
Also, on devices like the Beaglebone, you can use eMMC and an SD card at the same time.
I disagree. Having onboard NAND seems like a disadvantage, for like when it goes bad from all the usage.
Should have one M.2 slot.
Normal thing might be to boot from on-board NAND but put real workloads on M.2 if present.
I have to also vote for crypto extensions. The Pi 4 is the only ARM64 I have ever seen that lacks AES and CLMUL. Makes it pretty lame for web serving, router/firewall/VPN gateway, etc.
I would love to know why they chose to exclude this, it couldn’t make a big difference in price could it?
Good.
The last thing we need is profit mad crypto miners buying up every single Pi and we end up with shortages and insane prices just like with gaming GPUs.
Also a free market means people can do with the stuff they buy as they wish.
Whether I mine on it or stick it up my arse is just as valid as you gaming.
The main advantage will be general network ops.
I now understand why the Pi Zero comes standard without the GPIO header.
But it does appear the the GPU has perhaps 2x to 5x the FLOPs of the CPU, albeit with not much in the way of APIs to actually use it.
> 2) get rid of SD and have onboard NAND, like a phone
Why are these important?
But hey, maybe they just say 'screw it' and remove it anyway.
Only a portion of the workloads that are commonly used can be profitably vectorized using SIMD. The curiously perverse nature of SIMD is that the wider the vectors, the smaller the proportion of time used by these portions is, and therefore the less you gain from further vector width increases. AMD is currently spanking Intel in almost all practical vector workloads, despite having half the vector width. Apple isn't far off either, despite a quarter of the vector width.
It wouldn't really ever significantly hurt Apple if they just literally never implemented any flavor of SVE. Spending all that engineering effort on improving scalar throughput probably has much better real-world payoff.
It is because AMD and Apple have wider architectures with more vector ports, they can execute 3 or 4 of these instructions per cycle while Intel can only execute 2(or even 1 in some cases with avx512).
AMD has already said they will be adding AVX512 to the next zen, so they apparently think SIMD matters.
Apple will almost certainly implement SVE, they would be stupid to not do so, and they aren't stupid.
I don't think keeping NEON capability in hardware is going to hold back their chips much, so they probably won't be under any pressure to break compatibility with NEON-using apps anytime soon.
And the iPhone 12 won’t be out of support in 3 years.
It used to be the case if you used NumPy on a Mac it would use Accelerate but it no longer does because the BLAS/LAPACK API is so outdated.
Does Apple now have enormous input into the ARM spec process? (Or maybe they did already, because iPhones?)
For ARMv9 this potentially means they can pick and choose what they want to implement. I'm sure Apple has had plenty of input into the specification (along with other partners).
As an exception, an ARMv8.x-A chip can have some ARMv8.x+1-A extensions, but _never_ 8.x+2-A extensions (forbidden by Arm).
You just have one ISA minor revision of wiggle room.
And yes, Arm architectures are co-designed with plenty of output from partners.
The not well kept rumor is that Apple has a much looser than architectural license due to their very close relationship with ARM, both generally since the early 90s, and that Apple contributed very heavily to early AArch64 design and it's arguably theirs as much as it is ARM's.
There's also the fact that only pc and sp are retained on WFI, without the other GPRs.
Those quirks only affect bare metal kernel-mode (EL2), and do not affect user-mode or virtualised machines at EL1 in any way.
https://medium.com/swlh/apples-m1-secret-coprocessor-6599492...
>All licensees are not equal however, the first few are called lead licensees and companies pay an added fee for this honor. ARM picks 2-3 lead licensees for each market segment and works closely with them.
https://semiaccurate.com/2013/08/07/a-long-look-at-how-arm-l...
Even though Apple and Arm are possibly "at-arms" organizations these days (no one outside of these organizations really will know, except the lawyers, and the contractual obligations between Apple and Arm are likely locked up and highly secretive)
I'd say the other way around, that Arm is beholden to where Apple wants to take the Arm ecosystem, if only by way of showing other participants in the arm ecosystem what is possible. IE. Amazon is able to use the M1 as a gauge to how far it might be able to take its own Graviton cores.
Apple designs their own chips using ARM ISA (instruction set architecture). They design completely different chips separate from ARM.
Handful of companies like Apple, Broadcom, Marvell, Intel, Qualcomm, Samsung, have bought ARM architecture license.
>"SVE2 was announced back in April 2019, and looked to solve this issue by complementing the new scalable SIMD instruction set with the needed instructions to serve more varied DSP-like workloads that currently still use NEON."
Could someone say what "DSP-like workloads" workloads would be? I understand what a DSP is but I'm wondering what type of workloads that are not signal processing share similar characteristics.
But you can do SIMD-in-GPR tricks, or dedicated hardware, or GPGPU to replace that, so it's not a big problem if it's missing.
Frankly, having worked on consoles for 20 years and been through multiple architecture changes, I'd be perfectly happy to have another generation or two on x64 no matter how crusty it is -- devil you know and all.
Which generation are you referring to with 'last-gen'? If you mean PS4 / Xbox One then what else was on the high-performance CPU market other than x86 in 2012-2014? The POWER family had moved deep into HPC specialization. POWER7 was a bit old by 2013 and POWER8 wasn't quite there, but neither would make for a good gaming CPU with the heavy multithreading focus (4-way SMT on POWER7 & 8-way SMT on POWER8), and both would require substantial design changes to be scaled down to what a console would want. The ARM CPU of the time was the Cortex-A15, which was a fine mobile CPU but wasn't pushing any boundaries to make laptop or desktop CPUs nervous (and you'd have to substantially invest it its IO capabilities - there weren't really any ARM CPUs + PCI-E x16 + SATA SoCs laying around at the time)
If you mean the PS5/XSX generation then I think it's simply a why not Zen 2 & keep backwards compatibility? It's not like there's anything else you can buy that's a clearly better CPU offering anyway.
Qualcomm is likely going to want to focus their talent on the mobile market where they've historically been running only slightly customized ARM designs. I think Qualcomm would like to close the performance gap between it and Apple and fend off MediaTek who is taking an increasing amount of marketshare. As the US weans off CDMA, it's likely that Samsung might end up using their own chips more. So Qualcomm might want to focus Nuvia on mobile.
Qualcomm had tried its hand at Intel competitors, both on the consumer and server side. They seem to have given up on that for now.
Ampere is trying to get into the server space, and Anandtech notes that they're competitive with the AMD EPYC Rome series (https://www.anandtech.com/show/16315/the-ampere-altra-review...). Oracle said they'd launch some in 2021, but who knows how limited that will be or whether their plans will change.
Amazon will want to use their own chips rather than pay a third-party. More and more datacenter operations seem concentrated in the big three providers who might not want to let a new third-party get margin there (specifically, Amazon, Google, and Microsoft). I can't imagine Google not going with an in-house design if they wanted to launch an ARM platform.
Consumer devices are difficult. macOS won't work on non-Apple hardware. The Windows ARM experience will be sub-par because there's no dictator to force ARM on everyone (like Apple). Apple can say, "we're moving to ARM" and developers either get on board or are left behind. If Microsoft says, "we're moving to ARM" it's more like, "we're going to add ARM support, but we'll always be a first-class experience on Intel and you can still expect to run Win16 apps from 1990 on your new ARM computer and if developers and consumers don't show interest, we're flexible and we'll pivot away from ARM...so maybe don't buy an ARM machine right now because we haven't been able to convince developers...and since you won't buy the ARM machines, we'll probably just think it's a flop in two years and put fewer resources towards ARM...so there's no real reason for an ARM CPU company to want to make good CPUs...which reinforces why consumers shouldn't buy them..."
We are seeing movement in the space, but it's hard. I don't think we'll see a lot of consumer stuff for Windows and Linux comparable to Intel. I think it's just hard to break into that space. With Linux, the market is small already. With Windows, trying to convince consumers on a less-compatible experience or an experience that Microsoft is less committed to is hard. I think it's easier to compete against AMD and Intel in the server market where so much software is already CPU-independent and doesn't have the same reliance on consumer software and compatibility. I think if you're making consumer processors, you want to target Android and Chromebook where you won't be dealing with convincing consumers to select a lesser-compatible, lesser-supported alternative to Wintel.
Graviton2 is basically just an implementation of ARM's N1 design which uses a Cortex A76 CPU core. The single thread performance is better than like a Kirin 990, likely thanks to 32mb l3 cache and 8-channel memory controller, but it's not going to win any gaming performance crowns either.
I doubt either Microsoft or Sony are going to change tack and try to fight Nintendo (and probably lose) on Nintendo's home turf.
Or, could we see Nvidia (since it has acquired ARM) jumping into the console space?
A suite of three consoles ranging from mobile/portable, to 1080p/TV, to 4K/Desktop quality could be very appealing, especially if the top-end model could also use a mouse and keyboard and dual-boot into ARM Windows.
Nvidia would just need to design a Linux game OS and storefront. If they offered a 10% developer commission and paid for some big exclusives, they could have a very compelling product.
The rumor is that the Switch contract only came because Nvidia had a firesale on those older chips that they expected to make their way into flag ship Android devices, but instead sat in inventory for years.
Maybe after the ARM acquisition goes through (if it does), they'll start looking down that line again.
(You can take a peek at LinkedIn for example, of former Nintendo engineers, which makes the timeline more clear)
I find it very difficult to believe that contrary to the rumors, Nintendo had been sitting on a SoC for many years without releasing a product or even pushing for a die shrink. Like, the Tegra X1 was announced in 2014, and released into products by 2015, and the Switch didn't come out until 2017.
Turning off an entire core complex also points to it not being designed for them. Nintendo isn't known for paying for gates they aren't using.
https://www.linkedin.com/in/eyhchen from the NVIDIA side
"Gave a power consumption related demo to Nintendo team during sales process" (for his Jul 2013-Dec 2014 period of employment)
https://www.linkedin.com/in/gyferic from the Nintendo side
"Benchmark parallel processing - OpenMP stress test on SoC Nvidia Tegra X1" (for a Sept 2014-Mar 2015 period of employment)
The Nvidia Arm acquisition is very far from being a done deal, and may not happen at all (e.g. https://www.cnbc.com/2020/10/01/tech-investors-predict-nvidi... and https://www.cnbc.com/2021/02/12/qualcomm-objects-to-nvidias-...).
SVE is pretty much the evolution of the older Fujitsu vector extensions for SPARC64.
For client, stay tuned.
Which means you should be able to pick one up in some flagships phones early next year.
Last year, the A78 was announced in July and we had first Snapdragon 888 phone launch on Jan 1st.
So not a first of its kind, and the current version is still too low level, with an higher level API expect for some later version.