Hyperscale in your Homelab: The Compute Blade arrives
jeffgeerling.com
jeffgeerling.com
Then the next semester I am denied the ability to take a parallel computing class because it was for graduate students only and the prof. would not accept a waiver even though the class was being taught on the cluster me and a buddy built.
That I still had root on.
So I added a script that would renice the prof.'s jobs to be as slow as possible.
BOFH moment :)
* I frequently took computer science graduate courses and received only undergrad. credit because they could not offer the undergrad course
* Other majors were default prohibited from taking computer science courses under the guise of a shortage of places in classes. Even when those majors required a computer science course to graduate
I would like to point out that 300 and 400 level courses in the CS program usually had no more than 8 students. I distinctly remember meeting in a closet for one of my classes, because we had so few students they couldn't justify giving us a classroom.
Contrast that with the math department where I wanted to take some courses in parallel rather than serial. After a short conversation with the professor he said "ok sure, seems alright to me".
Math departments also tend to be a lot more lax in my experience. Case in point: I got sign-off to take a 300-level pure math class without its 300-level prereq. As well as just replace one required course for another to graduate.
Other majors were default prohibited from taking computer science courses under the guise of a shortage of places in classes. Even when those majors required a computer science course to graduate
I went to an institution that did the opposite; seats were reserved for non-cs majors despite a shortage of sections. This resulted in CS undergrads waiting for courses just so they could graduate. It was frustrating because it felt like the department was taking care of outsiders over its own.[1] https://en.wikipedia.org/wiki/Single_system_image [2] https://en.wikipedia.org/wiki/Distributed_operating_system
I'm quite grown, and I wonder about the ownership/control of the cluster and why he didn't simply lock the professor out entirely, contingent on the approval of his waiver.
If anything, doing something as small as lowering the priority of his jobs instead of brazenly stonewalling him might be the sign of immaturity.
It's not really much of a prank war if you do a small prank and the other guy tries like hell to pretend he didn't notice and also tries like hell to pretend he didn't see you derping around in the building... escalating would have been evil on my part :)
>He's running forty Blades in 2U. That's:
>
> 160 ARM cores
> 320 GB of RAM
> (up to) 320 terabytes of flash storage
>
>...in 2U of rackspace.Yay that's like... almost as much as normal 1U server can do
Edit: I give up, HN formatting is idiotic
(1) https://hardware.slashdot.org/story/01/07/14/0748215/can-you...
Distributed shared memory is another intriguing possibility, particularly since large address spaces are now basically ubiquitous. It would allow users to seamlessly extend multi-threaded workloads to run on a cluster; the OS would essentially have to implement memory-coherence protocols over the network.
https://web.archive.org/web/20010715201416/http://www.scient...
Sterling and his Goddard colleague Donald J. Becker connected 16 PCs, each containing an Intel 486 microprocessor, using Linux and a standard Ethernet network. For scientific applications, the PC cluster delivered sustained performance of 70 megaflops--that is, 70 million floating-point operations per second. Though modest by today's standards, this speed was not much lower than that of some smaller commercial supercomputers available at the time. And the cluster was built for only $40,000, or about one tenth the price of a comparable commercial machine in 1994.
NASA researchers named their cluster Beowulf, after the lean, mean hero of medieval legend who defeated the giant monster Grendel by ripping off one of the creature's arms. Since then, the name has been widely adopted to refer to any low-cost cluster constructed from commercially available PCs.
But then 5 years later I was working on them for a living in HPC, but they were no longer called Beowulf Clusters then.
While you are at it, also setup Longhorn for storage. With that solved, you might as well start hosting Gitea and DroneCI on the cluster, plus an extra helm- and docker repo for good measure. And in no time you will have a full modern CI/CD setup to do nothing but updates on! :-)
Seriously, though, you will learn a lot of things in the process and get a bottom up view of current stacks, which is definitely helpful.
Step 2: add advertising
Step 3: make more money than God.
Installing MPICH from source instead of from your distribution is best if you can't have all your cluster members running the same version of the same distro and/or have multiple architectures to contend with. But it takes forever to compile, even on a fast machine.
<https://www.oreilly.com/library/view/high-performance-linux/...>
A cluster of workstations (COW) is usually opportunistic exploiting existing systems, and lower density than a dedicated (usually rack-based or datacentre-based) cluster.
In practice, COWs usually turn out to be not especially useful, though there are exceptions.
I don't think there's any theoretical reason someone couldn't build a fairly realistic highly-complex "brain" using, say, 100,000,000 simplified neural units (I've heard of a guy in Japan who is doing such a thing), but I don't really know what it would do, or if it would teach us anything that is interesting.On the graybearding of the cohort, here’s a weird one to me. These days, I mention slashdot and get more of a response from peers than mentioning digg!
In 2005, I totally thought digg would be around forever as the slashdot successor, but it’s almost like it never happened (to software professionals… er, graybeards)
We were hanging out in the garage of a mutual friend, chatting. Got to the "what do you do" section of the conversation, and he says he works in massively parallel stuff at XYZ corp. Something something, GPUs.
I make the obvious "can you make a Beowulf cluster?" joke, to which he responds (after a pregnant pause), "you... do know who I am?"
Yep. Donald Becker. A slightly awkward moment, I'll cherish forever.
Q: Why would a 1U server need more than 200W if you're doing nothing more than basic network services?
I have mini tower servers that draw a fraction of that at idle.
Umm, I'm not sure I can afford the electricity to run kit like that :)
I'm currently awaiting delivery of an Asus PN41 (w/ Celeron N5100) to use as yet another home server, after a recommendation from a friend. Be interesting to see how much it draws at idle!
like
this
and asterisk for italics (I don't think there is a 'quote' available, and I'm not sure how they play together.* does this work? * Edit: No! Haha
*how*
*about*
*this*
Edit: No, no joy there either.I agree, it's not the most intuitive formatting syntax I've come across :)
I guess we're stuck with BEGIN_QUOTE and END_QUOTE blocks!
I wanted to do this
> * list element 1
> * list element 2
without the indent.I don't get why it just doesn't use common mark or something, it's just some inept, half-assed clone of it anyway.
What about cost, and other metrics around cost (power usage, reliability)? If space is the only factor we care about then it seems like a loss.
Having some affordable low power device with ECC would be a game changer for me.
I added affordable to exclude expensive (and noisy) workstation class laptops with ECC RAM.
Sorry if I am missing the obvious here, but why would ECC consume so much power?
I think it was idling at something like 30-40W with four HDDs and a UPS. I didn't have an especially efficient PSU and the UPS must have taken some power too. The motherboard alone would draw as little as 15W, I suppose.
>Corsair SF450 PSU
>ASRock Rack X570D4U w/BMC
>AMD Ryzen 7 Pro 5750GE (8C 3.2/4.6 GHz)
>128GB DDR4-2666 ECC
>Intel XL710-DA1 (40Gbps)
>LSI/Broadcom 9500-8i HBA
>64GB SuperMicro SATA DOM
>2 SK Hynix Gold P31, 2TB NVMe SSD
>8 Hitachi 7200rpm, 16TB HDD
>3 80mm fans, 2 40mm fans, CPU cooler
That was an at the time modern “Zen 3” (using Zen 2 cores) system on an X570 chipset. The CPU mostly goes in 1L ultra SFF systems. TDP is 35W, and under stress testing the CPU tops out around around 38.8-39W. The onboard BMC is about 3.2-3.3W of power consumption itself.Most data ingest and reads comes from the SSD cache, with that being more around 60W for high throughput. Under very high loads (saturating the 40Gbps link) with all disks going, only hits about 110-120W.
By comparison, a 6-bay Synology was over double that idle power consumption, and couldn’t come close to that throughput.
I could drop a few more watts if ASRock could put together a decent BIOS where disabling things actually disables things.
SuperMicro costs what it does for a reason.
—- ————-
If you’re looking for a chassis, I’m using a SilverStone RM21-308, with a Noctua NH-L9a-AM4 cooler, and cut some SilverStone sound deadening foam for the top panel of the 2U chassis.
Aside from disks clicking, it’s silent, runs hilariously cool (I 3D printed chipset and HBA fan mounts at a local library) and it’s more usable storage, higher performance (saturates 40Gbps trivially) and lower power consumption than anything any YouTuber has come remotely close to. That server basically lets me have everything else in my rack not care much about storage, because the storage server handles it like a champ. I really considered doing a video series on it, but I’m too old to want to deal with the peanut gallery of YouTube comments.
Do you know how to make the BMC not a laggy mess when using the “H5Viewer”? I’m getting basically unusable latency when the system is two yards away compared to a RDP server 1,000 miles away.
You don't need a cluster for that, even a 1st gen Pi can run those services without any problem.
I have multiple services running on it (including pihole, qbittorrent, vpn) and it's at about 40% mem usage right now.
Latency and bandwidth is atrocious in comparison, and you're going to run into problems like no individual memory allocation being able to exceed 8 Gb.
Like for running a hundred truly independent jobs then sure, maybe you'll get equivalent performance, but that's a very unique scenario that is rare in the real world.
That's not always a viable option because of hardware costs, and sometimes you want redundancy, but those concerns are on an orthogonal axis to performance.
They're currently putting Mx chips in every device they have, even the monitors. It'll be the base system for any electric device. I'm sure we'll see more specialized devices for different applications, because at this point, the hardware is compact, fast, and secure enough for anything, as well as the software stack.
Hello Apple Fridge
I don't think anybody in the HPC business really pursued mega-SMP after SGI because it was not cost-effective for the gains.
The really important thing is that the big ‘single machine’ you’re talking about already has numa latency problems. Sharing a chassis doesn’t actually save you from needing to tackle latency at scale.
It's just that there's a range of very well paying problems that scale quite well in message passing systems, and this means that even if your problem scales very badly on them, you might have easier time brute forcing the task on larger but inefficient supercomputer rather than getting funding for smaller more efficient one that fits your problems better.
It’s a fun toy for learning (and clicks, let’s be honest).
It’s not a serious attempt at a high performance cluster or an exercise in building an optimal computing platform.
Enjoy the experiment and the uniqueness of it. Nobody is going to be choosing this as their serious compute platform.
As usual this is either done for entertainment value or to simulate physical networks (not clusters).
However, while the high-core-count CPUs have excellent performance per occupied volume and per watt, they all have extremely low performance per dollar, unless you are able to negotiate huge discounts, when buying them by the thousands.
Using multiple servers with Ryzen 9 7950X can provide a performance per dollar many times higher than that of any current server CPU, i.e. six 16-core 7950X with a total of 384 GB of unbuffered ECC DDR5-4800 will be both much faster and much cheaper than one 96-core Genoa with 384 GB of buffered ECC DDR5-4800.
Nevertheless, the variant with multiple 7950X is limited for many applications by either the relatively low amount of memory per node or by the higher communication latency between nodes.
Still, for a small business it can provide much more bang for the buck, when the applications are suitable for being distributed over multiple nodes (e.g. code compilation).
...but the normal server is much cheaper.
Hyperscale in your Homelab. Something to hack on, learn, host things like Jellyfin, and have fun with.
Hard to see to an advantage given obvious limitations, although it may make it more fun to work within latency and memory constrictions, I guess.
> Yay that's like... almost as much as normal 1U server can do
It’s a fun toy. Obviously it isn’t the best or most efficient way to get any job done. That’s not the point.
Enjoy it for the fun experiment that it is.
If you are building full racks, it probably makes more sense to use ordinary systems, but if you want to have a lot of actual hardware isolation at a smaller scale, it could make sense.
In some colos, they don't give you enough power to fill up your racks, so the low energy density wouldn't be such a bummer there.
Can you expand on this please?
Of course, then you have to ask if you need the density. There are lots of ways to put Rpi in a rack.. and this approach gives up Hat compatibility for density.
For example, I’m considering a rack of Rpi with hifi berry DACs for a multi-zone audio system. This wouldn’t help me there.
https://store.planetcom.co.uk/products/gemini-pda-1
I absolutely can't imagine what I'd use it for, and yet, my finger has hovered over "buy" many many times over the last few years
There's a 5% chance that I fall madly in love with this thing and go tinker on some project in a coffee shop every weekend... but it's much more likely that I end up almost never using it :|
As for if that's a good use-case is a whole another thing.
The big win here would be that all of the network wiring is "built in" and compact. Blade replacement it trivial.
Have your fans blow up from the bottom and stagger "slots" on each row and if you do 32 slots per row, you probably build a kilocore cluster in a 6U box.
Ah the fun I would have with a lab with an nice budget.
Then I'd take a shared nothing cluster (typical network attached Linux cluster) and refactor a couple of algorithms that can "only" run on super computers and have them run faster on a complex that costs 1/10th as much. That would be based on an idea that was generated by listening to IBM and Google talk about their quantum computers and explaining how they were going to be so great. Imagine replacing every branch in a program with an assert that aborts the program on fail. You send 10,000 copies of the program to 10,000 cores with the asserts set uniquely on each copy. The core that completes kicks off the next round.
That stuck out to me too, they are making custom boards and custom chassis, surely it would be cleaner to route the networking and power through backplane instead of having gazillion tiny patch cables and random switch just hanging in there. Could also avoid the need for PoE by just having power buses in the backplane.
Overall imho the point of blades is that some stuff gets offloaded to the chassis, but here the chassis doesn't seem to be doing much at all.
What do you have in mind? I couldn't find this part. Really am asking.
The Chinese made ones are even cheaper, open up a TP-Link "desktop 8 port Gigabit Switch" and you will find the current "leader" in that market. Those datasheets though will be in Chinese so it helps to be able to read Chinese. (various translate apps are not well suited to datasheets in my experience)
Yeah, I found the ones on mouser and digikey. $20 is a bit much (not for a one off, but if you are aggregating low end processors you will need a lot of them).
I'd love something like a 12-20 port 1Ge with a 10Ge uplink. If you find a super cheap 1Ge switch chip and docs (I suppose you could just reverse engineer the pcb from a tp-link switch), please post it.
I really like Graviton from AWS, and Apple Silicon is great, I really hope we move towards ARM64 more. ArchLinux has https://archlinuxarm.org , I would love to use these to build and test arm64 packages (without needing to use qemu hackery, awesome though that it is).
* 4 Ampere Altra Max processors (in 2 or 4 servers), so about 512 cores, and much faster than anything those Raspberry Pi have.
* lots of RAM, probably about 4TB ?
* ~92TB of flash storage (or more ?)
Edit : I didn't want to disparage the compute blade, it looks like a very fun project. It's not even the same use case as the server hardware (and probably the best solution if you need actual raspberry pis), the only common thread is the 2U and rack use.
I'm keeping my eyes open though.
But of course the config I talked about is maxed-out and would probably be more expensive than 20k. It would be interesting to compare the TCO with an equivalent config, and I wouldn't be surprised to see the server hardware still win.
More efficient use of space compared to my current silent mini-home lab -- also about 2U worth of space, but stacked semi-vertically [1].
That's 4 servers each with AMD 5950x, 128GB ECC, 2TB NVMe, 2x8TB SSD (64c/512GB/72TB total).
The rest of the components are:
Board: AsRock Rack X570D4I-2T (2x 10GBe and IPMI!)
NVMe: 2TB Transcend TS2TMTE220S TLC
SSD: 2x 8TB Samsung 870 QVO
PSU: Seasonic SSP-300SUB (overkill, went for longevity)
CPU Cooling: Thermalright AXP-100 Series All-Copper Heatsink with Noctua NF-A12x15 PWM
Exhaust fans: 2x INEX AK-FN076 Slimfan 80mm PWM
On the air intake side, there's a filter sheet that I replace (or vacuum) once in a blue moon - the insides are still pristine after running for over a year now.
Interesting thing about cooling: one of those cases has a PSU with custom made cabling (reduced cables by about 90%). I was hoping it will reduce the temperatures a bit. Surprisingly there was basically no change. At full load all keep running at around 70 celsius.
Important: in such a small case, if you want silence you'd better disable AMD's "Core Performance Boost". This will make the CPU run at its nominal frequency, 3.4GHz for 5950x, otherwise it'll keep on jumping to it's max potential, 4.9GHz for 5950x, which will result in more heat, and more fan noise.
Like the Pine64 SOQUARTZ
The killer feature for their go fund would be is if they sourced a batch of pi compute modules ...
Took awhile to land on it though. Before that I tried all of the other distros on Pine64's "SOQuartz Software Releases"[2] page without any luck. The only one on that page that booted was the linked "Armbian Ubuntu Jammy with kernel 5.19.7" but it failed to boot again after an apt upgrade.
So there's at least one working OS, as of last week. But its definitely quite finicky and would probably need some work to build a proper device tree for any carrier board that's not the RPi CM4 Carrier Board.
Sure, it's hulking, obsolete, and very loud beast, but it's hard to beat the price to performance ratio there... just make sure you don't put anything super valuable on it because HP's old proliant firmware likely has a ton of unpatched critical vulnerabilities (and you'd need an HP support plan to download patches even if they exist)
I picked up a HP 705 G4 mini on backmarket for $80 shipped the other day to run Home Assistant and some other small local containers. 500gb ram, Ryzen 5 2400GE, 8gb ddr4 w/ a valid windows license.
Sure it's not as small or silent, but there's no way to beat the prices of these few-years old enterprise mini-pc's
At this rate I have so little hope in other vendors that we'll probably just have to wait for the RPi5.
There is still near zero availability in mass market for CPUs you can stick into motherboards from one of the top ten taiwanese vendors of serious server class motherboards.
And don't even get me started on the lack of ability to actually buy raspberry pi of your desired configuration at a reasonable price and in stock to hit add to cart.
Modifying the services I'm working on to build multi-arch container images was not as straightforward as I imagined, but now I can take advantage of both ARM and AMD64 nodes on my cluster (plus I learned to do that, which is priceless)
https://www.theverge.com/2014/6/4/5779468/twitter-engineer-b...
I'm noticing that our JVM workloads execute _significantly_ faster on ARM. Just looking at the execution times on our lowly first-gen M1s Macbooks is significantly better than some of our best Intel or AMD hardware we have racked. I'm guessing it all has to do with Memory bandwidth.
Will need to deal with NUMA issues on the software side.
They're not going to get maximum performance from a nvme disk, the cpus are too slow, and gigabit isn't going to cut it for high throughput applications.
Until manufacturers start shipping boards with ~32 cores clocked faster than 2ghz and multiple 10gbit connections, they're nothing more than a fun nerd toy.
I would, however, say that while I'm in the general target audience, I won't do crowdfunded hardware. If it isn't actually being produced, I won't buy it. The road between prototype and production is a long one for hardware.
(Still waiting for a very cool bit of hardware, 3+ years later - suspecting that project is just *dead*)
> 320 GB of RAM
Depending how you feel about hyperthreading, there are commodity dual-CPU Xeon setups than can do this as well.
- running self-hosted services for 3.5 users doesn't take many resources and Pi can often handle multiple services
- Compilation is CPU heavy and I/O heavy operation, more memory you have on a single machine is better.
I use Pi as my on-the-go computer in places where I can't ssh to my home-server. Sometimes I can't even get projects indexed without language server being killed by OOM 20 minutes later (on my PC it takes <20 seconds to index).
Probably with one of these in the middle: http://www.mercotac.com/html/830.html
And efficiency is much better than the Intel or AMD chips you could get in a used system around the same price.
Hopefully the next generation Pi has more PCIe lanes or at least a faster lane.