Even a laptop can run RAM externally thanks to CXL
techradar.com
techradar.com
https://www.computeexpresslink.org/about-cxl https://en.wikipedia.org/wiki/Compute_Express_Link
It might work for laptops, but it is not the main goal and it might be only Framework who will do something in that direction.
(I do embedded and not server stuff mostly these days anyways, but generally interested in systems stuff, so)
Obviously if you're just interested then go for it. It's just not practical generally speaking.
Just today, I replaced the whole thing with a single mini PC with a laptop Ryzen chip (6900HX). Pulls like 60W at the wall under load.
This is a circuitous way of saying that, for non-mission critical stuff, you can self-host and save a ton of money over the equivalent hosted option. You don't even need server hardware necessarily, depending on what you're trying to learn from the experience. Standard caveats about securing your home network apply.
(Note: I'm not actually seriously saying a 16x blade server and a mini PC are equivalent. I got the blade server to learn things (more networking, mostly), it was very, very overpowered for my workloads)
The thing I'm building/hosting can benefit from plenty of cores, but for the start it's just prototype with a limited audience, so I don't need to go hardcore.
But I have a friend who has rackspace in a DC, and my thought was if I bought something in a rack form I could move it there and pay him to host it later. But I suspect it doesn't make a lot of sense, really.
Rack stuff can be efficient (<100W idling), especially recent model year stuff, but it gets pricey quickly. You could spec a used 1U single socket Epyc Rome or Milan for under $2,000 if you're willing to shop around, but if you don't have other rackmount equipment, I'd stay away as it's such a price (and power) premium
The laptop I'm typing on is a Ryzen 6850H, so not too far off from the 6900 you're talking about and it has plenty of oomph. I might consider such a thing.
That's pretty awesome! I would have loved that at one point but they're extremely loud and I remember them needing 16 amp IEC leads
And indeed, the power supplies run on ~200V, my house serendipitously is wired with a dedicated 240V/50A circuit because the previous owner did glasswork and had a massive electric kiln.
The purpose of CXL is to allow for memory coherency between different CXL devices. To quote the spec on type 2 devices:
> CXL Type 2 devices, in addition to fully coherent cache, also have memory, for example DDR, High-Bandwidth Memory (HBM), etc., attached to the device. These devices execute against memory, but their performance comes from having massive bandwidth between the accelerator and device-attached memory. The main goal for CXL is to provide a means for the Host to push operands into device-attached memory and for the Host to pull results out of device-attached memory such that it does not add software and hardware cost that offsets the benefit of the accelerator.
There's some cool things you can do with CXL, like resurrecting the whole persistent memory idea with low-latency flash, making hardware offload devices more capable since you now get free cache coherency, and a whole bunch of other stuff.
But yes, it's really not for consumer use-cases. The applications I've seen colleagues work on are mostly enterprise stuff like cool RDMA integrations, cache-coherent flash, and more I can't talk about here.
https://www.youtube.com/watch?v=x1NWdoXmVsE
Persistent Memory databases are such a neat idea, it's a shame that the hardware for it isn't commonplace.
Alas, optane's dead now. I do know people actively working on resurrecting a lot of pmem work on low-latency flash, however, and it seems like this is one area with a low of momentum behind it.
If you treat it like memory it's always going to be slow memory. If you treat it like storage, it can be very fast storage. If that makes sense.
Software that "treats storage like memory" would end up looking a lot like javacard imo. Or something like Samsung's in-memory key-value database stuff. But it wouldn't really look like a linux kernel allocating memory inside an all-pmem partition.
Difference, in my head, being that essentially the filesystem is an unnecessary layer/abstraction in the middle. You want something that either looks like a garbage-collected runtime, or an in-memory database that vacuums, or LISP collections of object-trees, etc.
There is no point to having a separation between memory and disk anymore, that is the point - "resource allocation is disk persistence" so to speak.
It would have required a big-bang rewrite/second-system that is just not possible with the dominance of the existing RAM/storage dichotomy. Or at least a couple killer apps from large vendors/etc that really outperformed what was possible with RAM.
SNIA has some interesting talks regarding using non volatile memory and its applications (DBs). Where using pmem and DAX to directly to store logging operations [1].
[0] What is Direct-Access (DAX)? https://kb.pmem.io/faq/100000008-What-is-DAX/
[1] SDC2020: How can Persistent Memory Make Databases Faster, and How Could we go Ahead? https://www.youtube.com/watch?v=rTgITrVhpQM
"How to Build a Non-Volatile Memory Database Management System"
Sometimes it becomes useful to appreciate the hardware and what it is doing, reminds me of the quote.
* You don't have to be an engineer to be a racing driver, but you do have to have Mechanical Sympathy *. Jackie Stewart
wasn't Optane (hardware) trying to use CXL?
It does kind of upset me a bit that CXL 3.0 still seems purely host-to-switch-to-device oriented. If you have your formerly PCIe slots on your cores speaking CXL and doing directory memory over fabric, I'd really really love to be able to talk to other hosts. Maybe that happens & is possible in 3.0, but it feels like CXL isnt paving that cowpath, isn't making is obvious, and that there will be a bunch of proprietary nasty ways to bridge computers & chat over CXL that are all non-standard, & I wish CXL had been more direct about making themselves & their upcoming switched fabric viable & interesting for host-to-host.
I'm wondering when we'll see the first CXL DPU style device that uses some of the fancy new 800Gbps or even 1.6Tbps networking stuff that's being developed right now. That'd be enough bandwidth to put things on the other end of a data-center with very little other than the latency penalty.
Maybe switch to the source article? https://www.servethehome.com/fadu-cxl-2-0-switch-and-pcie-ge...
What are you going to do with the other 999,999,360KB?
Anything to do with data processing, machine learning, llms- hundreds of gigabytes of ram can be incredibly nice.
Especially as pandas recommends ram equal to 5-10x the size of the dataset.
When you're an individual- not having to think about managing a cluster...
Edit: One of the lucky 10000- famous bill gates joke- got it. I walk away cultured
OK? What laptops even support it? How much capacity is there ("over 1TB" is kind of vague)? So many questions...
[0] Emulating CXL Shared Memory Devices in QEMU https://memverge.com/cxl-qemuemulating-cxl-shared-memory-dev...
[1] CXL support in QEMU https://www.qemu.org/docs/master/system/devices/cxl.html
But with Genoa for example, socket to socket latency has climbed up to 220ns, and going across nodes on a socket is 110ns. I feel like CXL will be less than 2x a hit, if only because cores themselves are having higher and higher latencies. https://chipsandcheese.com/2023/07/17/genoa-x-server-v-cache...
I have hopes. I can't help it. Cool shit is inspiring. I, however, am not holding my breath.
If CXL is not being used by servers, will it ever become an offering for laptops?
It is entirely possible to have many bits in transit on a single channel, wired or wireless.
Electrical Engineers since at least DDR2, probably earlier, need to ensure that all lines are delay matched to about 100picoseconds.
That is, in DDR2, if DataBit#1 takes 1.1nanoseconds from start-of-wire to end-of-wire, then all other bits must be somewhere between 1.0ns to 1.2ns in length.
This requires impedance controlled pcbs from the manufacturer, and length tracking software for PCB design. (Note: advanced PCB CAD software will even recalculate the speed of light across different lengths, as "inner" tracks have more dielectric surrounding them, slowing down electricity's speed, while "outer" tracks are a bit faster)
-------
The dielectric of FR4 (the glass/resin used to make PCBs) is what determines the speed of light of the copper actually. And it can be 3.6 or 4.2 or whatever, but the PCB manufacturer will tell you in the PCB specsheets.
Copper guides the wave, but the actual wave travels in the dielectric / insulation between the wires (FR4 in modern PCBs, but if you had open-air wire the dielectric would be the open-air surrounding the wires that form the ground-loop return path). There's been a huge amount of improvements to the understanding of electricity in the past 3 decades, and the old circuit models (sometimes still taught today) are kind of obsolete btw.
-------
I think I heard that DDR4 is now 10picosecomds (10x more accuracy in delay matching), which is doable with modern CAD and PCB software. I dunno the delay requirements of DDR5.
------------------
https://en.wikipedia.org/wiki/Transmission_line
Anyway, look through this Wikipedia on Transmission Line theory, which is a closer model to how electricity "actually" works (but still isn't perfect). But its the level you need to think of electricity to understand modern CPU-to-RAM connections.
In particular, DDR2 / DDR3 / DDR4 / DDR5 connections are almost certainly either Microstrip or Stripline connections, with a fair amount of PCB / Electrical Engineering going into the design to make sure everything works as expected.
I'm personally now imagining a specialized database appliance which takes the role of the whole of the pager and buffer pool management from a DB (or KV store or whatever); a physical box which ties secondary storage arrays + large quantities of RAM + buffer pool mgmt firmware together on a box, then connect to host system via CXL. Host system does query planning end execution and everything else...
Is anybody doing this? Does anybody want to found a startup with me to do this? <sips more and more coffee...>
[1] https://hpi.de/rabl/teaching/master-theses/ongoing-masters-t...
For some reason the whole issue of buffer mgmt is something I nerd out on a bit.
https://en.wikipedia.org/wiki/CAS_latency#Memory_timing_exam...
Laptop SODIMMs are already severely limiting DDR5 bandwidth, server/desktop DIMMs are getting close.
Users aren't gonna like it, but unswappable packaged RAM is coming to most CPUs.
Depends. Apple uses LPDDR5X on a cheap substrate, which is basically just RAM chips soldered to a tiny motherboard, and that is already very fast. But it could also mean HBM on an expensive interposer or cheaper Intel EMIB, or something like Samsung's proposed Wide I/O, or even something new like stacked RAM with TSVs.
CPU/RAM packaging is getting increasingly complicated.
> what protocol does it use that would be faster than DDR5
Depends, but it can just be straight up DDR5. SODIMMS are terrible because they need 1.35V (vs 1.1V stock DDR5) for really slow speeds and terrible ram timings, while soldered DDR5 and LPDDR5X do not. The closer the DDR5 gets to the CPU, the faster, more power efficient and lower latency it gets.
i would like to see whether the high memory use cases (like running LLMs) can benefit more from direct memory access to GPU that efforts such as DirectStorage is working on (https://news.ycombinator.com/item?id=33265666) and just faster, low-latency drives can bridge this gap faster.
From an operations point of view (Cook's forte), it makes a lot of sense.
Jony Ive would be proud as well.
If you don't have the need for all the extra CPUs, just being able to attach more memory to a single CPU through CXL may be cheaper.
Even more so on a laptop.