Banana Pi to Launch 24-Core Arm Server
cnx-software.com
cnx-software.com
Sorry to be the one to wake you from your apparently very pleasant daydream, but I just picked three random chips from their website and the only links available are "request info" and "contact sales." I don't see any evidence to support a claim that they've changed at all. They're still the worst.
This seems a bit uncharitable. Here's another way to look at it: perhaps the commenter you replied to was familiar with Broadcom's policies from some years ago, but had not kept up to date with them and didn't want to say one way or the other about Broadcom's current behavior.
Right now if I want to cross-compile Gentoo for ARM and ARM64. I can build "most" packages with the included emerge-wrapper. Which is great as it runs basically at 1:1 speed as doing x86_64 compile jobs
However a ton of packages still fail, many of their upstream developers also refuse to incorporate changes that might fix that https://dev.gnupg.org/T2370
For those packages, I have to use a QEMU usermode chroot. Which on my i7-4790 build host, is slower than native compilation on the Raspberry Pi itself
I'd love to be able to do direct, native builds to sidestep these flaws. But every "consumer" ARM board (Raspberry Pi, ODROID, ROCK64 etc) are flawed in some way as to make them unusable for development
- Raspberry Pi lacks enough RAM and will hang even when doing single-threaded builds (heavy packages like GCC and Rust usually hit this limit)
- The ODROID-C2 hits a similar limitation due to only having 2GB of memory itself
- The ROCK64 "can" complete full self-hosting builds (slowly) but has a staggering amount of kernel bugs relating to its Ethernet and USB 3, which frequently cause the system to hang
Here with the rock64 it was just a matter of time till hardware support settled in mainline, and it seems like we're there!
Can someone ELI5 what exactly the compiler is doing when needs far greater than 2 Gigs of an RPI's RAM to compile itself/Rust?
Also, how is it a "flaw" of the RPI that GCC must exceed 2 Gigs of RAM to compile itself/Rust?
Warning-- I may use the explanation to cudgel future Electron-so-fat threads...
And how long does it take?
Cross-compilation gives you performance, but a lot of other headaches... and currently using a quad or hex-core ARM takes quite a long time... so this would give a significant performance boost, especially with the 32GB of ram it seems to have.
For linux distros, they download the source, apply some patches, then let their build server/farm have at it.
Managing a cross compilation toolchain in addition to cross compilation flags and what-not for joe-random-package with quirks quickly becomes a nightmare. Version upgrades will break everything, or pull in build-system linked libraries when they're not supposed to, etc.
This sort of high-core count arm server alleviates a lot of headaches - it gives you faster build times as well as native builds. It also does away with a lot of complexity in your build pipeline, since you don't have to carefully isolate the build environment to prevent rouge linking x86 binaries.
Most linux distros that offer ARM versions already build natively... so this just gives them better build times, or more concurrent builds. That's a win.
Besides, with Cortex-A53, any single-threaded step in a build (linking, etc), is going to be very slow.
It's a sort-of cool product, especially if it comes at a good price point. But it's unlikely to make a huge difference in the overall market.
You want to keep it plugged into a monitor for a while and watch out for the undervolt warning (either a lightning bolt or a rainbow in the corner of the screen). Improving the cooling, the quality of the USB cable, and the power supply are also really good ideas.
What have you found so far to be the best options?
In my experience the quality of cabling and power supplies are the factors that most frequently affect stability of electronics hardware.
Make sure the USB port can pump out enough current. Make sure you use a short, thick USB cable. Once you do that, stability issues are gone.
I also put a heat sink on the CPU, but it probably has no effect on stability. It reduces the amount of time the CPU spends in thermal throttling.
The long cables where? I think I know what you're getting at but I would really love some more explanation if possible.
All wires have resistance, and all else being equal, a 2x longer wire will have 2x more resistance. Due to Ohm's law, there is a voltage drop in the cable when you are powering a load (such as the RPi), and this drop will eventually go below the minimum for the load.
Example: the RPi needs 5 volts at 2 amps.
If your cable is 0.1 ohms, then the voltage at the Pi will be 5 - (2 * 0.1) == 4.8 volts.
If your cable is 1 ohms, the voltage at the Pi will be 5 - (2 * 1) == 3 volts.
But because the Pi doesn't always need the full 2 amps, that 1 ohm cable will work sometimes, making debugging harder. Only when you try to do something compute and/or RAM intensive will you notice the board randomly die.
The LED light does go steady red which is an indication that it has conked off and needs a hard reboot
My Pi currently has an uptime of 388 days. I use it as a file-server/backup-machine in my home network. It runs hourly backups for multiple boxes in the background (I use vhpi [0] for this). I have a usb battery attached to it for uninterruptible power supply. It works like a charm. I use NFS over Samba, though.
But I power them off dedicated wall warts.
... and what effect are you implying that this has in this particular context?
That's a lot of ARM64 power in one platform.
It depends on the type of ARM core and the manufacturing process. If it has 24 A53 cores then it wouldn't beat a 4 core desktop CPU from 2012. If it has 24 A72 cores but a manufacturing process older than 28nm then it still wouldn't be competitive with anything on 14nm or better.
Same goes with most long-running VPS'... bursts of activity here and there or once a day or whatever.
Being able to run those workloads on cheaper cores that consume far less power would be a huge win. Not all workloads would be a good fit, but those which are, this can be big.
Might not make a difference if you attach them to spinning disks or something, but you can really turn the thing into a wall-plug sort of thing and never worry about it again.
if it's active 24x7, then hard-drive overshadows both ATOM and ARM CPU in terms of drawing currents.
ATOM supports multiple native SATA ports and it is still hard for ARM to catch up. Marvell is the only ARM vendor does nice ARM+SATA but its chip is hard to find, plus it is used in low-end NAS devices while ATOM is in the mid-range, it's more for the cost-saving instead of power-saving though.
I was attracted by ARM lowest-power in the last, until I realized hard-drive is the one consumes most power and ARM does not do well with non-multimedia IO.
WD Red 10TB maxes out around 5.7W [1] under heavy read-write activity and can use much less when idling.
Intel Atom is typically in the 20W range for the parts commonly picked for NAS solutions, but can go as low as 8W too [2] and again when idling would use less.
I don’t think they have anything <5W which would directly support your statement the HDD uses more power than the CPU.
[1] https://www.wd.com/content/dam/wdc/website/downloadable_asse... [2] https://www.anandtech.com/show/11144/intel-launches-16core-a...
Big memory for zfs arc would help too.
Amazon isn't just providing ARM CPU servers because of customer demand, and nor is it just to get leverage with Intel. That's not the way they tend to work. They're providing them so that they can leverage them in-house as well. Amazon just likes to take the position that what is good for the goose is good for the gander. If it's good for them, there's bound to be customers that will find it good for them also, and as you sell it to customers, you get to those advantages of scale much quicker.
Medium term, I expect Intel to look much more innovative than current. Long term, we’re either all dead because the robots won or uploaded to the cloud hive mind once our contributions to the universe outweigh our costs.
Instead what we got are big chips with ARM ISA that are essentially scaled up to roughly Xeon sizes... which in hindsight I suppose is obvious coz that's probably where the money is.
Nevertheless I still think there is a niche for smaller servers with smaller granulation of quanta of compute capacity.
Have not kept myself up to date in case AMD offers decent alternatives in this segment.
Though my hope would be for ARM to lower the price point for these kind of servers.
I wonder what the TDP is on the SoC, the SC2A11 is only 5W (though they get a lot out of that!). I'd like to see what they could accomplish with a higher TDP target.
NEON is implemented on a per-core basis, not as a separate coprocessor. Every A53 (and your handful of A72s) supports NEON.
I'll give you the GPUs... if you really want to put in all the effort for that sweet sweet OpenCL 1.2 action on an out of date vendor kernel.
An RK3399 cluster, and a single many-core system is a nonsensical apples to oranges comparison.
... well... and largely because most neural network software only supports CUDA and/or has rudimentary OpenCL support.
Nvidia hardware isn't really cheaper per FLOP in all cases, AMD and Nvidia still leapfrog each other back and forth.
AVX instruction set, for instance, seems to only be on select Intel and AMD cpu's, even modernly manufactured ones. TensorFlow changed to requiring AVX since version 1.4 (Nov 2017), which has caused grief for many users... even those with recent/performant systems.