Exciting Days for ARM Processors
smist08.wordpress.com
smist08.wordpress.com
With all due respect, I love my raspberry pi, but Apple just needs to thank whoever was in charge of acquiring PA Semi’s know how in 2008, along with the chip mastermind that is Johni Srouji. Or themselves from 1990, when they actually founded ARM as a joint venture with Acorn computers to make chips for the Newton.
The first ARM-based device for Linux that I remember getting big was the Sheevaplug. It was an ARM-based "plug computer" with an SD card slot, a USB port IIRC, serial/JTAG port, Ethernet, Wifi, and no display.
I had its next generation, the Guruplug, which had 2 Ethernet ports, Bluetooth, and a couple USB ports. Used it as a router for a few years.
I don't know how long Debian's arm distro was around before then but that was my goto for the Guruplug. Worked great.
Later on the Linksys NSLU2 (2004, a network USB hard drive adapter) came out and was the first ARM based "computer" to really catch on. The custom firmware and optware offered a pretty good posix experience. It wasn't until 2008 when debian support came out, give the world a fully modern Linux experience on a little embedded device.
After that companies started seeing a niche market and programmers accepted the challenge of installing linux on everything and we got things like the Sheeva Plug and BeagleBoard.
A few years later the Raspberry Pi came out with an unbeatable price and community support.
And then a decade later we're getting hints of the possibility of ARM computers with true workstation level of computing power.
And we had countless applications in mobile phones and automotive infotainment systems for Cortex A series CPUs and other embedded systems for Cortex M series CPUs before.
But of course it's still a bit weird. For example, XHCI (due to being attached through a bad PCIe host controller) can only DMA to the lower 3GB of memory. In FreeBSD, we've never had to implement DMA limits for ACPI devices, because no real system had this kind of limitation before, and now I've had to write this: https://reviews.freebsd.org/D25219
Amazon are betting on this on the server side. They've developed a number of ARM-powered EC2 instance types: A1, M6g, C6g, and R6g.
They haven't yet made a burstable ARM instance type.
As you both rightly say, the ARM CPUs Amazon are using are not the same as Apple's. Instead, they are custom built for Amazon. [0][1]
[0] https://aws.amazon.com/blogs/aws/new-ec2-instances-a1-powere...
[1] https://aws.amazon.com/about-aws/whats-new/2019/12/announcin...
So I don't think it would have that much of an impact.
Jim Keller as much as anyone I would think.
Really? AMD's recent moves with Zen/Zen2/Zen3/EPYC look like a big step forward. Zen2 chiplets are the biggest change in years. Zen3 IPC is supposed to be significantly better than Zen2 (17%). The previous gen Ryzen was like 15w TDP, where as the A12Z in the Mac Mini ARM is 15w TDP, but the Zen crushes the A12Z Bionic on benchmarks.
It's difficult for people to remember, but ARM came from really terrible performance to a spot where it is getting in the ballpark of x86, so the advancements look impressive, but that's like saying if I go from $100 to $200, the gains look impressive, whereas you only went from $10000 to $11000, it looks like you are standing relatively still in comparison.
But x86 architectures are decades of maturity, so someone getting a 17% IPC lift (Zen2->Zen3) or a massive reduction in TDP on such a complex chip, isn't stagnation, it's actually MORE impressive IMHO.
ARM is going to reach marginal returns, and soon the yearly perf boosts won't look as impressive anymore.
In the end, I think we'll see convergence of performance. ARM will still have a power advantage on mobile, because they don't have so much backwards compatibility legacy that x86 has to support. However, are Macbooks and Mac Desktops going to be performance and price competitive with Linux x86? I doubt it.
First of all, ARM vendors haven't even come close to the GPU performance of NVidia or AMD's discrete GPUs. An ARM A13 or Mali is not going to compete with a laptop with an RTX 2060, 3060, or RDNA1/2. And secondly, just looking at AMD, it's possible to lift performance / watt still in x86.
I think the laptop space in the x86 realm is still exciting, because of what AMD and NVidia are doing.
Further, the average laptop consumer is happy with their intel integrated graphics. All they do is browse the web and punch some numbers in excel.
The news is based on leaked benchmarks from upcoming tiger lake mobile socs with markedly higher graphics scores. I suppose it is possible that these tests were done with a discrete GPU hooked up. Hard to say for now. But it seems like a logical step for intel to upgrade their anemic integrated graphics with their next generation stuff.
Apple's view of what "discrete" performance is, is probably 2-generation old cut-down mobile variants. For example, their top-of-the-line uber-expensive Mac Pro ships with Radeon Vega II architecture from 2017. Meanwhile, AMD RDNA2/Navi and NVidia Ampere are about to ship.
From my view, they're at least 2-3 generations behind in performance, and their focus on mobile means they're always going to be behind given thermals. NVidia and AMD are focused on maximum performance, and they assume gaming takes place plugged into a wall socket. And because they're focused on maximum performance, especially for triple-A title games and DCC, they're putting in features like DXR (DirectX Raytracing) with support for hardware accelerated ray-intersection. How long until I can run Minecraft with ray-tracered shaders at 2k or 4k on Apple's GPU? (https://www.youtube.com/watch?v=AdTxrggo8e8)
I mean, Apple's engineers on ARM are good and they did a fantastic job improving the PA Semi IP, but I don't think they're going to take what's essentially a PowerVR architecture which was never competitive with discrete, and leapfrog NVidia and AMD.
They do less well in compute, but even there they do well, and a 16 core A14 GPU is likely to bat with a 1060.
[1] https://images.anandtech.com/graphs/graph13661/103805.png (not a perfect benchmark because of CPU bottlenecking, but there are no good benchmarks to choose)
Of course it's impressive. Even if that was somehow all down to the better node, their next generation SoCs will still use a newer node than NVIDIA's next generation GPUs.
There's no reason NVIDIA should use a process behind Apple forever because both uses TSMC (and Samsung) and process improves slowing down.
I don't really get your argument. Apple customers buying Apple Silicon Macs—which, again, will probably have a GPU over three times as fast as the A12X—aren't going to let hypotheticals detract from their powerful and power-efficient GPUs. ‘But NVIDIA didn't optimize for efficiency’ and ‘but NVIDIA hypothetically could have used a newer node than they did’ don't count for squat.
The key is consistent performance and that is something they can manage with full control over API and silicon end to end. They are already doing this well on cores that are two generations old. I expect to see a good compromise on performance and power which is what we need for general purpose computing.
To be clear I have actually canned my gaming pc recently which was a ryzen 3700x and GTX 1660. I haven’t missed it, the big titles or the pain in the arse getting it stable and built to start with.
What did you have, crashing games? It has been like 7 years since I had any stability issues on windows for gaming.
Every PC build takes a lot of risk on. Not one I’m willing to take any more. Commercial desktops are either junk or too expensive. I’d rather use a single sourced well integrated appliance.
Edit: also the dubious nature of some parts is a worry as demonstrated here https://www.reddit.com/r/WatchPeopleDieInside/comments/g0420...
Of course by "integrated" Apple could mean they just solder a GPU chip to the board. Or they built a surprisingly performant GPU core for their SoC, though I wouldn't expect it to be much faster than what AMD integrates in their CPUs (30% at most? Still much slower than discrete).
Source for 15W A12Z TDP? And the benchmarks?
Whether it's accurate or not, the reality is, the Ryzen Mobile is pretty efficient for its performance, and Intel LakeField is basically going with a "BIG.litte" architecture as well (x86+atom). AMD in particular is a fierce fighter, and they're not just going to sit still and let ARM capture the laptop and desktop markets, especially since a huge existing ecosystem gives them an advantage, and many people with work or business laptops aren't necessarily going to throw away everything to get an extra mm of thinness, or slightly longer battery life. After all, if battery life was the end all and be all, mobile phones could have doubled battery life a long time ago by simply doing less. Putting super high pixel density displays and constantly running background services is one reason why battery size keeps going up, but battery life hasn't improved as much.
Current Cinebench R15 benchmarks [0] of real world Lakefield based hardware put this CPU 10+% lower than a 2013 AMD A6-1450 8W 'Temash' CPU [1]. Let's say it's the same ballpark. Unless it's a benchmarking fluke or early hardware issues Lakefield is nothing to write home about.
[0] https://www.notebookcheck.net/Exclusive-First-benchmarks-of-...
Four extra, eight total. The A12X is the one with a core disabled.
Hmm? The A12X beats the 3700U in Geekbench 5's single-thread, multi-thread and compute benchmarks, and not by trivial margins.
Your power comparisons are also unfair; if the A12X ever draws 15W, it would be way at the edge of its power curve on all cores[1], whereas 15W on the 3700U is a tepid clock for the Ryzen. The A12X is much more reasonably considered an ~8W part.
There is no chance the A12z in the DTK is a 15w chip.
Do you have a source?
The Zen3 rumoured impressive IPC improvement will still be below Willow Cove or Tiger Lake. An uArch that was supposed to be out in 2018. Intel's 7nm ( In between TSMC 5nm and 3nm Node ) Golden Cove was supposed to be launch this year in Intel's original roadmap.
From roughly 2.5 years ahead of the industry to now lagging behind 1.5 years. That is 4 years of difference.
That is why The Intel world has stagnated in recent years. 4 years is very long in tech industry.
So when Apple releases a new chip, it can just re-compile the LLVM-IR to it, to make use of newer features and compiler optimizations for that chip.
Basically, the only applications that have to do anything to transition to Arm on Apple are those that are not using the AppStore... which from the platform's perspective, is kind of their own fault.
In general, you are correct. LLVM-IR is architecture dependent. Things like the size of a pointer, the size of an `int`, parts of the calling convention, or "architecture-specific defines in C code" have already been "hardcoded"/expanded into the generated IR.
In practice, you can easily re-compile x64 IR to arm64, as long as the IR does not use, e.g., "arch-specific" LLVM intrinsics (like explicitly using, e.g., NEON intrinsics). Apple does not let you ship this kind of bitcode to the AppStore, so if you want your software to use NEON on iOS, you need to call the opaque Apple math libraries, which they can just provide for a different architecture.
The other things you need to worry about is, e.g., calls to system libraries (e.g. you can't call a Windows API, generate IR for it, and then try to re-compile that IR for Linux, because that API will fail to link). In practice, if you provide the symbol, everything will work, and the system APIs for MacOSX, iPadOS, and iOS are quite similar.
The x86-macos IR won't be recompilable to Linux or windows, but can be made recompilable to arm-macos.
Note also that this is not the first time Apple does this with the AppStore. They silently migrated all apps from 32-bit ARM to 64-bit ARM, recompiling the software for you. This hints that they have additional capabilities to changing some architecture details in the AppStore's IR, like pointer-sizes (these Apps are not running in 32-bit compatibility mode in 64-bit CPUs, but are running as native 64-bit apps).
>Additionally, Mac AppStore apps are shipped without Bitcode.
The Apps themselves are shipped to users as binary blobs, but developers ship bitcode to the AppStore, which Apple has been recompiling for each new arm processor in their iphones (so on an iPhone 11 you get a different binary blob than on an iphone 6, but the developer didn't ship two blobs, nor they recompiled their App).
It's currently driving 2x 1080p screens over USB C + DisplayPort chaining. It can drive 3 of these chained in this way. Keyboard and mouse are connected to the USB hub on the display, so it's a single cable connection from my laptop to start working in the morning.
I hope to have a phone at some point that runs Windows 10 Pro, that I can plug into a USB C plug and get straight to work, that would be amazing!
MSFT Edge and Windows Terminal both have ARM64 builds (and so does VS Code Insiders), and WSL works great. I work with tmux+nvim on Ubuntu ARM, which runs almost everything I need.
A lot of stuff _doesn't_ work though, but the stuff that does works well. I hope that we see some inexpensive (200-300USD?) NUC style devices that can run Windows 10, it would make for a great little computer.
So plenty of stuff is available.
Besides it is not like one can blindly run .NET on other platforms, because plenty of applications make use of COM or DLLs written in C++.
Win32 is supported with MSVC just fine.
You can also use MinGW with LLVM (not MinGW with GCC), available at: https://github.com/mstorsjo/llvm-mingw
(this also is probably part of the reason apple deprecated opengl, I imagine)
Apple Removing Support for AMD GPUs in macOS Arm64
Quite likely they'll support eGPUs, and maaaaaybe big MacBook Pros with Apple SoC + AMD GPU could eventually happen??
I suspect the real reason is people don't use it much anymore, there were some figures published that showed the number of bootcamp users had fallen from 15% to 2% over its lifetime.
I still use it although I will probably be retired by the time my current machine needs replacing.
These definitely are exciting times. Most Brits born in the 80's and 90's will have used Acorn computers in school, no one predicted that what would grow out of it (they weren't particularly fast).
Yes, that's a very tiny number compared to the billions of CPUs made by Intel, Apple, AMD, Qualcomm and Samsung, but the point is, they are already doing ARM CPU integration and have carved out a decent niche for themselves.
That's 7 major generations behind the curve.
For normal PC workload, such bandwidth is not useful and GDDR increases latency.
IMO Interesting thing like HBM is happening in Server/HPC world.
https://screenrant.com/ps5-io-ssd-speed-tech-specs-playstati...
Also. It might be much longer than "soon". Though Sony is targeting the next gen SSD specs (unlike Microsoft), there will need to be a new generation of motherboards and CPU upgrades before PCs will be able to come close to the PS5 bandwidth, even with the same exact next gen SSD.
But I want the PS4 to enforce a fast SSD as a requirement for PC game, like you say. I think game consoles have a power to standardize great techs, but no longer makes innovating techs for computing.
"a new generation of motherboards and CPU" was available in 2019.
I remember them being amazingly fast. I was comparing them to Amigas in 1987 though. The Motorola 68000 in the Amiga 500 was less than 1.0 dhrystone MIPS, while the Archemedes A3000 was about 4.0. For reference, the machines cost about £499 and £599 respectively in 1987.
As a result the frame rate, poly count and screen size in Zarch on the Archemedes was higher than Starglider II on the Amiga :-). Look https://youtu.be/MNXypBxNGMo?t=36, https://youtu.be/edDLqlG4quw?t=132
Apps opened so fast you couldn’t perceive a delay between double clicking and then appearing. New windows for apps also instant.
The only times I noticed waiting was when doing something that legitimately seemed like I should have to wait. A big operation in a paint program. Starting up a game.
I think this consistent ‘zero delay’ came through firstly the fact that lots of the OS was in ROM, then the tight design and coding of the OS, then perhaps that the ARM CPU was pretty fast.
Running RISC OS on a Raspberry PI natively it’s still just as snappy for normal operations. Things that take significant CPU are also now fast thanks to the several hundred fold increase in MHz.
I’d love to see another ‘personal computer’ OS appear which focuses on making the UX feel like it’s all working hard real time (or as close as it can) so that we can get back to the joy of feeling that the machine is consistently predictable in how it responds. Today’s ‘personal’ OSes feel more like driving a car with automatic transmission, steer-by-wire and several mechanical problems that cause it to randomly fail to respond in the expected time or just do something entirely different.
Right now, getting Apps ported to ARM seems a lot more likely than AI / ML porting their work off CUDA. And I am thinking if CUDA will end up like Excel for Microsoft. It is not that we cant move off Excel, but the code and value running on CUDA / Excel will not be worth the effort porting away from it.
No, you can turn off secure boot.
When SIP (rootless v1) was introduced, macOS came with a way to disable it, for the sake of kernel developers. When APFS was introduced by default, macOS retained two workarounds for kernel developers: both the ability to boot from alternate APFS volumes within the APFS container; and also the ability to install to/boot from HFS+. And when the system volume became truly read-only in Catalina (rootless v2)—that's right, even that restriction could be overridden via csrutil, specifically to facilitate kernel development.
Apple is never going to lock themselves out of booting custom OSes that are almost, but not quite, macOS; because that's how new macOS gets made.
And, of course, as long as you can boot something that's not literally macOS; then that same mechanism can always be abused to boot something that is much less macOS, or in fact not macOS at all (but may still seem to be macOS, from the bootloader's perspective.)
So actually, going forward they only need to make it easier for their own kernel teams, not third parties.
Given how they keep stressing the security and kernel stability during those sessions.
They are turning macOS into a proper micro-kernel, by changing the engines while keeping the plane flying.
Edit: this one, about 19 minutes in: https://developer.apple.com/videos/play/wwdc2020/10686/
"Enable/Disable Secure Boot. On non-ARM systems, it is required to implement the ability to disable Secure Boot via firmware setup. A physically present user must be allowed to disable Secure Boot via firmware setup without possession of PKpriv. A Windows Server may also disable Secure Boot remotely using a strongly authenticated (preferably public-key based) out-of-band management connection, such as to a baseboard management controller or service processor. Programmatic disabling of Secure Boot either during Boot Services or after exiting EFI Boot Services MUST NOT be possible. Disabling Secure Boot must not be possible on ARM systems."[1]
This requirement is no longer needed for Windows 10 logo certified x86_64 computer,[2] but I've yet to see a vendor actually take it out.
For SOC systems, at least if you wish to have Windows 10 Logo, it makes it optional if I'm reading their spec sheet correctly:
"Requirement 10: OPTIONAL. An OEM may implement the ability for a physically present user to turn off Secure Boot either with access to the PKpriv or with Physical Presence through the firmware setup. Access to the firmware setup may be protected by platform specific means (administrator password, smart card, static configuration, etc.)"[3]
That said, that's only for Windows logo certified devices. I'd assume if you don't intend to make you system Windows compatible then you can do whatever you want.
------
[1]https://docs.microsoft.com/en-us/previous-versions/windows/h...
[2]https://www.pcworld.com/article/2901262/microsoft-tightens-w...
[3]https://docs.microsoft.com/en-us/windows-hardware/drivers/br...
This is where I think Apple can do it right. They can tune drivers and fully optimize system performance, since they basically own the whole stack, and won't be stuck with proprietary broken blobs or reverse engineering a 3rd party design.
Is there anything else I should be looking out for?
Raspberry Pi 4 is now OpenGL ES 3.0 certified, so I’m expecting it to only get faster with time. My main concern is a Pi5 being released and my “old” hardware becoming obsolete.
Just as a helpful data point.
No, it didn’t really work out for you.
Yes, it’s working out great for me!
Konsole with Terminus is an aesthetically pleasing environment in which to work, 8 hours a day.
i3 in 4k, even at 30Hz, looks great. Vim renders fast (66-80ms keyboard latency.)
eog, evince, and gimp let me fiddle with materials. gnome-screenshot copies to the clipboard for rapidly sharing notes with others, based on what I have on screen.
davmail transparently connects dumb IMAP clients to the corporate email, address book, and calendaring systems. I have an XBiff alerting me to incoming Outlook emails!
REPL.it is slow in Firefox but fast in Chromium. It’s going to work just fine for managing pupils’ class work.
My home directory is encrypted and mounted on login. Mutt and offlineimap let me handle last term’s mail in bulk.
Python and git and tmux let me prepare and mark work from the students, or at least they should do come next term. I have all my little tools working to give me an excellent markdown environment for note taking on classes and pupils and meetings. Asciidoctor for the more typographically demanding class materials.
Absolutely all of this was a breeze to set up because all of these technologies just work in Raspbian/Debian on armhf. It has, so far, been unbelievably clean and pleasant.
I even have LXC working locally, for running Unix and IPv6 experiments with multiple little Alpine “VMs”. χ!
For reference, my day is spent:
• editing code, markdown, and asciidoc for classes;
• managing pupil work in repl.it, GitHub, and Google Classroom;
• managing pupil reports and records of work in custom school IT web applications; and
• emailing with colleagues and pupils with outlook’s webmail.
If I can get a snappy and correctly scaled terminal and browser — which I think you’ve managed after some effort — I hope it will work out well.
I have a 4k LG display at home which might be the biggest pain point. Rest assured though it’s good ole 1080p all the way, at school: HDMI1 for my private desktop and HDMI2 for the 80” class display.
I am optimistic this is a workload that an ARM desktop can handle.
While Apple has open sourced a few things, I doubt they will open source their graphics drivers. Though one can hope.
I also want something weird. Kids kind of look down on PCs and especially Linux. It’s seen as kind of crappy compared to the shiny of iPhones and Windows gaming rigs. I hope to show them that there’s a third kind of computer with which one can do a lot, that also happens to be pocketable and under $100.
(And also: ANSI keyboard Pinebook Pro has been sold out for weeks.)
Works fine on my Pi4 with official 32-bit Debian Buster OS: https://github.com/Const-me/Vrmac/tree/master/VrmacVideo
Based on CPU usage, the included VLC player uses same API.
(It literally did just work — shipping arrived first thing this morning, armhf Raspberry Pi OS installed and running within half an hour.)
I’d love to try an Atom / Intel compute stick if you had a recommendation though.
Pine’s RockPro64 and Libre Renegade Elite are two lovely RK3399 devices that I’ve tried. However, each has sufficient quirks that something more mainstream like the Raspberry Pi 4 (with a wide user base, more Google hits for common problems, etc) will hopefully let me focus on using the device as opposed to getting it to function.
Non-RPi SBCs also seem to require a lot of off-brand community patches to get them working. I’m pretty wary of bringing an OS build which I downloaded from a third party github page onto a campus full of children. I don’t have any reason to believe the community builds are nefarious, but it’s less of a risk to stick with Raspberry Pi, IMHO.
The Pinebook pro is really nice (I've had one since the beginning of the year) but I completely share your sentiment. I really like mine and spend almost no time fiddling with it now. It made a really nice, reliable daily driver for my son this spring during the sudden and unplanned distance learning the pandemic brought us.
But I would not recommend it to anyone whose computer I couldn't/wouldn't personally fix right now. It's close, but not there yet IMO. Bootstrapping, power management and sound all have more rough edges than running Linux on a normal x86 system does right now. I say this fondly: it's still a bit of a project right now. (That's part of the fun for me, but doesn't match your stated focus at all.)
Neither of which they executed on.
The A13 bionic wouldn't physically share any components built by ARM?
If that's the case, it's like saying AMD use Intel chips, but really they just licence the instruction set
I appreciate some commenters saying the raspberry helped, but really it's apple who are the taking a big step forward with popularising a third player in the CPU space. Unless they are willing to sell their chipsets, I doubt it will affect consumer products much outside of their own ecosystem.
Again nothing exciting for ARM, but more the future of apple chips
Ampere will, though — Altra's going to be good. Also there's Marvell (ex-Cavium), but that's custom cores, and rather HPC-focused.
More choice is more better though, and since Jon Masters works for Nuvia, we know Nuvia is going to have the most compliant and least quirky hardware, especially in the PCIe area :) Though to be fair, Ampere's first generation is already very decent in terms of this, and Ampere even promised to eventually contribute support for their hardware to EDK2 TianoCore open firmware..
https://twitter.com/jonmasters/status/1237879913245368321?s=...
Jon Masters works at Nuvia.
You can probably make a good case for the optimal base page size being larger than 4kB today, but probably not by very much. Maybe 16 kB or so. But then it's not a huge advantage over 4 kB which has the benefit of compatibility, so, meh..
Extraordinary claim, but any evidence to support that Fugaku is super useful apart from making headlines?
When they did this with PPC they said the same thing, buying a PPC Mac Mini ended up being a bad choice.
Proof will be in the pudding of course.
Now my comment has been down voted into oblivion.