Ampere Altra Dev Kit has launched 32, 64, 80 core arm64 processor
ipi.wiki
ipi.wiki
Why? In colo you’re usually not as concerned about density. We’re renting a chunk of rack space for a bottom basement price of $200/month and we can easily stick a router and compute / storage nodes with 64 TB of disk and boat loads of memory in that rented space; plenty for our needs without coming anywhere close to the power allocation.
When we were purchasing a few servers for our colo, we ultimately just went with AMD because the price per core was there. It ends up being less performant than an m2, maybe m1 ish.
Ultimately we’re a small market though. The ‘sexy’ thing to do is waive your AWS invoice size around at other hipsters during cocktail parties; not to own your own hosting even though you’ve paid for it serval times over. I don’t see this changing anytime soon.
See https://news.ycombinator.com/item?id=26058061 about the M1. If ECC were added between M1 and M2, Apple would almost certainly have announced it as a feature.
https://www.tomshardware.com/news/ampere-64-core-arm-worksta...
For those keeping score the Ampere Altra 128 core 3GHz Supermicro telco edge server being used in 5G base stations benchmarks equal to 100 RPi4 and 22% more energy efficient.
And of course, Nvidia has their 72x2 core Grace chip coming soon. It will be very interesting to see how that stacks up against the ARM incumbents.
I'm curious about these cocktail parties you speak of - over here, in-person tech meetups have been pretty rare since covid
There is a difference in the type of workload and the ultimate total size (i.e. if you have a single 2-rack workload then the M1 and M2 aren't an option at all -- those only make sense in parallel instantiated workloads and container-based loads).
After experimenting with a shipment of M1's (before they had to be commissioned for an office that was still in the process of being built and furnished) I'm pretty confident that similar JVM performance isn't currently available on anything else if your max allocation per instance is 8-core or lower. As soon as your cores-per-workload is higher the cost/density swings back in favour of Intel again because the Ultra in the studio doesn't fit in 1U and energy/cooling COP becomes on par with classic servers again. Maybe on M2 that's not the case, but when I ran the experiment the M1 was all that was available.
Maybe once the FPGAs for Intel get public bitstreams it might be as efficient on Intel again, but right now, it's not (and also not on AMD), and thus also not on anything Oracle.
Note: didn't test Ampere, but would love it.
Sort of like the ideal machine for the Qubes OS.
GPUs are massively parallel processors with 1000s of cores. While CUDA isn’t the easiest to work with, some Python libraries such as JAX and Tai-chi are attemting to remove the bar between CPU and GPU computing completely.
A program using either cab transparently switch between a CPU or GPU backend.
I think something along the lines of a Pentium Pro or an ARM core. Pentium Pro had 5.5 million transistors. A modern CPU has about 1000x more, so about a thousand Pentium Pro-grade processors would fit in die like a modern 7770X.
I'd take that over my GPU any day.
The hard and expensive part is, obviously, memory, cache, and interconnect. The even harder part is software. And the less hard part I'm intentionally oversimplifying is power consumption.
If someone could wave a magic wand, and there were OS, app, compiler, video game, etc. support for both MIMD and current architectures, I think MIMD would take over overnight.
Most of what computers do is ridiculously parallel. From each browser tab getting an isolated CPU, to having a spreadsheet spread out among cores, to rendering fonts in a document.
However, given a universe with trillions of dollars invested in the status quo, a disruption would need some sort of rather complex pathway, with some niche markets, some growth strategy, etc. As someone pointed out, Intel tried with Phi and failed.
An example of MIMD system is Intel Xeon Phi, descended from Larrabee microarchitecture.
https://en.wikipedia.org/wiki/Larrabee_(microarchitecture)
Its x86 cores were based on the much simpler P54C Pentium
https://en.wikipedia.org/wiki/Xeon_Phi
That is the core from the generation before the Pentium Pro. Larrabee was supposed to be a GPU, wasn't good enough to compete at that, then they rebranded it as Xeon Phi but cancelled it a few years ago.
The fact that programs can switch between those two execution contexts seamlessly is a nice benefit but still isn't the same as having many ordinary CPUs. I've used GPUs extensively and I'm really happy to see their capabilities more integrated in everyday languages but they are at best for now a co-processor like device, a chunk of hardware that you offload specific parts of your workload (typically: the numerical chunk of it that is massively parallel).
To elaborate on this: OS creates processes, assigns ids to them, assigns other physical resources to them s.a. association with namespaces (which later gives them user permissions, network access, virtual memory access, filesystem access etc.) and then these processes are associated with some GPU resource.
If and when OS will start creating processes entirely on GPU, then it will make sense to talk about how GPUs are solving the problems of threads per core etc. For now it's a moot point.
This is not said to discourage though. I really feel like CPU-centric model of what we call "computers" is not a good one going forward. The "periphery" is growing smarter with each generation, and wants to do its own computing, and spread its load somehow, and we keep coming up with ad hoc solutions that don't mix well with CPU-based concurrency, s.a. async I/O or CUDA. We really need a different concept of concurrency that would be more uniform and at the same time more flexible across different devices that can do work concurrently, and this, interface if you will, must come from the operating system, not as a user-space library to be truly effective.
That is how IBM has handled their mainframes for quite a long time. You can hot-plug them in/out.
Here's another offering from Gigabyte, dual socket and supports the 128 core processors. EATX formfactor.
edit Comparing apples and oranges — this is probably closer to AWS Graviton systems, but in a desktop form factor.
I might try it on Scaleway, but I wouldn't be surprised if the Hetzner machine has different performance characteristics.
(also, Oracle cloud has them in their free tier).
https://www.linkedin.com/feed/update/urn:li:activity:6977673...
They were willing to spend whatever it took to provide reliable WiFi throughout the ~20 acre property. Equipment cost was concern #4 or #5.
We still went with gigabit fixtures and not 2.5G. The ISPs were only providing a 200mbps pipe (symmetrical), intra-LAN communication is not super common in this type of deployment, and no client devices currently on-property supported 2.5G.
So, yeah, for us, it still made sense to go with gigabit. Even when we look out over the next 5 years (expected lifetime of this new equipment) we didn't see the need for 2.5G growing _that much_.
I’ve been under the impression that there are patents (expiring soon) keeping the cost of 10G hardware high enough to prevent it from becoming the default choice. I don’t have a concrete source for that info though.
I wish these machines (or Apple Silicon with Asahi Linux) had existed at the time I worked on GHC and aarch64 support was starting to take shape. Would have been a game changer in terms of productivity...
I did have a 32 core Threadripper.
I spend a lot of time trying to write parallel code. On my 12 core intel NUC (a mobile CPU chip) My bank simulation shards money between threads and can handle 700 million transactions per second. I am sure you could with more optimisation increase this and it scales per number of threads due to the sharding.
Adding persistence would drastically slow things down.
Assuming there is no powercut to multiple availability zones, you could keep data in memory on one machine in a circular buffer and then persist every second on the other availability zone. Depends how long it takes to restore power.
If you don't use a mutex or a lockfree algorithm, you can end up with more money (money creation) or money destruction (money missing) when you transfer money between accounts and there is another transaction to the same account in a similar time frame. This is due to a race hazard where addition and subtraction are each individually 3 operations and if they interleave then they cause the writeback of an incorrect result and omission of another result:
thread 1 thread 2
read 1000
read 1000
calculate 1000 - 25
calculate 1000 - 50
write 950
write 975
This simulation has money creation, because the two transactions don't see eachother.So I add the money up at the end of the simulation to see if it is equal to the money that the simulation started with.
When I say money sharding, I am not sharding by account. If an account has 12,000 in it, and I have 12 threads, each thread stores 1000. The fast path is a transaction goes to that thread and that thread handles the transaction if the amount of money being deducted is less than or equal to 1000. If the transaction amount is greater than what is available in one thread, then it has to be routed to other threads.
I never generate a transaction greater than what is available than all threads, so that part always works, for routing to accounts that have enough money.
An extremely slow path would have to transact a bit from multiple threads until the transaction is complete.
I've never worked on fintech or banking software, so I don't know how it works in practice but I did try implementing an order matching engine (which from my perspective is just a sort ascending + sort descending of participants bids or asks) I haven't worked on parallelising that yet.
In this design, it appears that the memory slots are perpendicular to the PCIe slots, and thus perpendicular to typical server front->back airflow. Is this a common design element in all COM-HPC CPU boards? If so, it seems like a flawed design..