There still isn't a clear path to profitability for any of these AI products and the capital expenditure has been enormous.
There still isn't a clear path to profitability for any of these AI products and the capital expenditure has been enormous.
Their inventories are not what consumers use.
Consumer DDR5 motherboards normally take UDIMMs. Server DDR5 motherboards normally take RDIMMs. They're mechanically incompatible, and the voltages are different. And the memory for GPUs is normally soldered directly to the board (and of the GDDRn family, instead of the DDRn or LPDDRn families used by most CPUs).
As for GPUs, they're also different. Most consumer GPUs are PCIe x16 cards with DP and HDMI ports; most hyperscaler GPUs are going to have more exotic form factors like OAM, and not have any DP or HDMI ports (since they have no need for graphics output).
So no, unfortunately hyperscalers dumping their inventories would be of little use to consumers. We'll have to wait for the factories to switch their production to consumer-targeted products.
Edit: even their NVMe drives are going to have different form factors like E1.S and different connectors like U.2, making them hard for normal consumers to use.
I put together a small server with mostly commodity parts.
Wrong. It is still just NVMe over PCIe like every other modern SSD form factor.
All you need is a fixed-latency, dumb translator bridge where the adapter forces everything into a simplified JEDEC-compliant mode.
CA/CK Line Translator with a Fixed Retimer as the biggest mismatch between RDIMM/UDIMM is the command/address path.
RDIMMs route CA/CK to RCD to DRAM, and the UDIMMs route CA/CK to DRAM directly, take the UDIMM CA/CK, delay + buffer + level shift it, feed it into a "RCD" like input using a delay locked loops (DLL).
Throw in a SPD translator, PMIC and voltage correction, DQ line conditioning and some other stuff into a 10–12-layer PCB, retimer chips, vrm, and level shifters.
It would cost about $40 million to fab and about $100 per adapter but would make bank with all the spare UDIMMs when the bubble bursts.
HBM/GDDR is not necessarily as useful to the average person as DDR4/DDR5
Will just have to settle for insanely cheap second hand DDR5 and NVMe drives I guess.
But anyway, the trick is to run it in the winter and keep your house warm.
And while data centers might sign favorable contracts, I don't think they are getting electricity that far below retail.
I could go on, but I think you get the point: the dollar cost for me to run a hypothetical version of Gemini at home far exceeds the price Google pays to deliver the same service to me. I'm effectively arbitraging the price of electricity by subscribing to an LLM service, because I could never run it for anywhere near the same price in my house -- and to provide the service at the speed and reliability of Gemini, I doubt I could even fit the equipment in my house-sized house and still have room for me to live in it.
Anyway, I stand by my original comment - it's neither easy nor cheap to run a frontier model at home.
AI GPUs are stripped away of most things display-related to make room for more compute cores. So in theory, they could "work", but there are bottlenecks making that compute power irrelevant for gaming, even if they had a display output.