Raspberry Pi Compute Module 4 on sale now from $25
raspberrypi.org
raspberrypi.org
See the announcement post for all the details, but I was lucky to be able to work with the CM4 for a few weeks, and I posted my impressions here: https://www.jeffgeerling.com/blog/2020/raspberry-pi-compute-...
(Oh and BTW I had completely forgotten but cheers for the follow-up on TRIM/UASP for USB the other week :))
For my Turing Pi, when it's delivered, I'm hoping to start with a small Kubernetes install for testing. Then when the TPi 2 is released next year, build a second one with CM4 modules for more involved Kubernetes testing/development while putting Ceph on the old one.
(To head off inevitable comment: no I'm not expecting this to perform well, and it's not for production work, it's purely for experimentation with clusters)
One note about your article:
Due to the nature of PCIe, ANY GPU will happily run on the 1 lane presented here. The limitation will be data transfer, which primarily shows up in texture loading. Tom's Hardware did an article [0] a decade ago comparing lanes to performance. This is also why some x16 connectors will be listed as "electrically" x8; they only have 8 lanes exposed.
However, you will have issues with the power delivery; PCIe GPU's are allowed to pull up to 75 watts from the slot itself before going to direct connectors. There are ways to fix this; I ran a mPCIe (WiFi) to external PCIe adapter for a couple years on an old laptop with a GPU just fine.
Another use case (and likely more relevant) would be plugging in an HBA for a NAS. Upgrading the NIC to something multi-gig sounds interesting, but that PCIe x1 is not going to be any faster than the built-in GBe except in edge cases. I'm actually really curious to see what people do with this.
[0] https://www.tomshardware.com/reviews/pcie-geforce-gtx-480-x1...
For the NIC, it's cool to see we'd be able to get two or maybe three gigabit interfaces on the same Pi, could be interesting for a variety of different networking devices, coupled with the faster SoC in the Pi 4.
That changes the use case a bit; the Aquantia multi-gig cards have drivers for ESXi. You could then run an eMMC boot for the hypervisor and run everything via the higher speed adapter, saving the GbE for management/migration.
32 flavours/models seems like choice-saturation. If they still make all 32 flavours a year from now, I'll be supprised.
They guarantee 7 years of availability for Compute Modules, including this one, so... prepare to be surprised.
They had plenty of time over the last 16 months since the launch of Pi 4 to decide how many SKUs to make for the Compute Module. This certainly wasn't done on a whim by some inexperienced team, which is what your post seems to imply. They've been doing this for years, and they realize the commitment they're making.
It is perfectly viable for them to shift to socketed eMMC and still accommodate all available flavours. But then how they sell will set that and if out say 99% of sales are for only 5 models and a socketed eMMC consolidate the offerings whilst still providing exact compatibility with the initial guaranteed offerings, then I still would not rule out such a change a year on, once they see what flavours do sell.
That was the basis of my comment - nothing more, nothing less.
If they change the product at all, downstream companies may have to get their products entirely re-certified, which can be an expensive and time consuming process. Those companies rely on having a stable supply of exactly the compute module that they designed their product around. Every detail matters.
These SKUs are fixed in stone. If they discontinue or modify them at all before the specified 7 year date, that would substantially damage their relationship with downstream companies. They have upheld their previous commitments, so I don't expect them to start breaking commitments now.
Compute Module 4 + IO Board looks like a winner setup for a small multimedia box
Really want something similar to the Ubiquiti CloudKey form factor (though ports could be on same edge)
Or do I need to write code with a specific SDK to take advantage?
Faster than what? The previous Compute Module 3+? Absolutely.
Other examples of raspberry pi compute module devices, custom drones, ad-hoc ventilator brains, custom network hardware, custom "IoT" devices etc etc. Without all the inconveniently large consumer-sized ports on the compute module, these devices are well suited for all sorts of low to mid volume B2B products. Raspberry Pi Foundation has built quite a reputation of being a very solid SBC supplier in the commercial realm.
It's too bad there isn't something like a game console (for profit) wrapped around this thing, that would really kickstart hardcore optimization.
Much like the game consoles of old (ps-one, nes, snes, genesis), the devs would squeeze every last bit of performance out of those platforms, but there wasn't a non-gaming computing userbase to take advantage of those optimizations.
I had heard the cell in the PS2 was used in some supercomputers, but that might have been the Cell whitepaper presenting unrealistic flops and someone on the Folding.Home taking the hype at face value.
They are used in some NEC commercial displays because the cool part is you get to add a small connector to the base design, and bolt on a linux computer, because the CM is just that, a full computer, designed to be integrated into cool things
Dreaming up a bit of a glass palace, I'd imagine you could fit a cluster of 3x4 CM4s without much extra and 1x4 CM4s with NVMe connectors into a 1U server chassis after some angle grinding. Possibly even two layers above eachother. That would be quite a computer cluster to have to play with at the homelab.
It's very absurd to break it out for each of them modules to get an independent PCI-e slot. This could raise the layout of the board significantly.
Yet the speed they currently offer (PCI-e x1 Gen2) isn't fast enough for RDMA to effectively "chain" the compute modules.
So I would see that they would have some sacrifice (e.g. fallback to the smaller and more cliche mini-PCIe like most recent fresh IOT boards out there), or even remove PCI-e expansion but offer something else (e.g. embedded SATA host/multi-NICs where only the master node could control, then the rest of the children will have to rely on RDMA, despite it will be slow and painful).
It's not quite economical for Turing Pi to implement NUMA-like architecture so I would rule this out.
From my perspective its 4*7 = 28 weak ARM cores vs 8 strong x86 cores. You can see that ARM actually had more cores, giving it an advantage plus to parallelized workload compared to x86.
Hell, maybe we can mix them in a bunch so that x86 runs powerful applications like GitLab, Prometheus and Postgres while ARM runs massively parallel workload that GPUs can't handle: Function as a Service (in AWS terms, Lambda; in CNCF's term, OpenFaaS), Linkerd handler (service mesh needs some kind of scheduling though), microservice replicas.
In the end CPU are all going to have a designated purposes, despite it should have had been "general purpose".