I definitely don't think compute is anything like railroads and fibre, but I'm not so sure compute will continue it's efficiency gains of the past. Power consumption for these chips is climbing fast, lots of gains are from better hardware support for 8bit/4bit precision, I believe yields are getting harder to achieve as things get much smaller.
Betting against compute getting better/cheaper/faster is probably a bad idea, but fundamental improvements I think will be a lot slower over the next decade as shrinking gets a lot harder.
> I definitely don't think compute is anything like railroads and fibre, but I'm not so sure compute will continue it's efficiency gains of the past. Power consumption for these chips is climbing fast, lots of gains are from better hardware support for 8bit/4bit precision, I believe yields are getting harder to achieve as things get much smaller.
I'm no expert, buy my understanding is that as feature sizes shrink, semiconductors become more prone to failure over time. Those GPUs probably aren't going to all fry themselves in two years, but even if GPUs stagnate, chip longevity may limit the medium/long term value of the (massive) investment.
Could you show me?
Early turbines didn't last that long. Even modern ones are only rated for a few decades.
For comparison, Moore’s law (at 2 years per doubling) scales 4 orders of magnitude in about 27 years. That’s roughly the lifetime of a modern steam turbine [2]. In actuality, Parsons lived 77 years [3], implying a 13% growth rate, so doubling every 6 versus 2 years. But within the same order of magnitude.
[1] https://en.m.wikipedia.org/wiki/Steam_turbine
[2] https://alliedpg.com/latest-articles/life-extension-strategi... 30 years
[3] https://en.m.wikipedia.org/wiki/Charles_Algernon_Parsons
There is an absolute glut of cheap compute available right now due to VC and other funds dumping into the industry (take advantage of it while it exists!) but I'm pretty sure Wall St. will balk when they realize the continued costs of maintaining that compute and look at the revenue that expenditure is generating. People think of chips as a piece of infrastructure - you buy a personal computer and it'll keep chugging for a decade without issue in most case - but GPUs are essentially consumables - they're an input to producing the compute a data center sells that needs constant restocking - rather than a one-time investment.
- Most big tech companies are investing in data centers using operating cash flow, not levering it
- The hyperscalers have in recent years been tweaking the depreciation schedules of regular cloud compute assets (extending them), so there's a push and a pull going on for CPU vs GPU depreciation
- I don't think anyone who knows how to do fundamental analysis expects any asset to "keep chugging for a decade without issue" unless it's explicitly rated to do so (like e.g. a solar panel). All assets have depreciation schedules, GPUs are just shorter than average, and I don't think this is a big mystery to big money on Wall St
If we're talking about the whole compute system like a gb200, is there a particular component that breaks first? How hard are they to refurbish, if that particular component breaks? I'm guessing they didn't have repairability in mind, but I also know these "chips" are much more than chips now so there's probably some modularity if it's not the chip itself failing.
* memory IC failure
* power delivery component failure
* dead core
* cracked BGA solder joints on core
* damaged PCB due to sag
These issues are compounded by
* huge power consumption and heat output of core and memory, compared to system CPU/memory
* physical size of core leads to more potential for solder joint fracture due to thermal expansion/contraction
* everything needs to fit in PCIe card form factor
* memory and core not socketed, if one fails (or supporting circuitry on the PCB fails) then either expensive repair or the card becomes scrap
* some vendors have cards with design flaws which lead to early failure
* sometimes poor application of thermal paste/pads at factory (eg, only half of core making contact
* and, in my experience in aquiring 4-5 year old GPUs to build gaming PCs with (to sell), almost without fail the thermal paste has dried up and the card is thermal throttling
Consuker gpu you have no idea if they've shoved it into a hotbox of a case or not
Since they were run 24/7, there was rarely the kind of heat stress that kills cards (heating and cooling cycles).
I could imagine scenarios where someone wants a relatively prompt response but is okay with waiting in exchange for a small discount and bids close to the standard rate, where someone wants an overnight response and bids even less, and where someone is okay with waiting much longer (e.g. a month) and bids whatever the minimum is (which could be $0, or some very small rate that matches the expected value from mining).
Number of cycles that goes through silicon matters, but what matters most really are temperature and electrical shocks.
If the GPUs are stable, at low temperature they can be at full load for years. There are servers out there up from decades and decades.
Effectively every single H100 in existence now will be e-waste in 5 years or less. Not exactly railroad infrastructure here, or even dark fiber.
That which survived, at least. A whole lot of rail infrastructure was not viable and soon became waste of its own. There was, at one time, ten rail lines around my parts, operated by six different railway companies. Only one of them remains fully intact to this day. One other line retained a short section that is still standing, which is now being used for car storage, but was mostly dismantled. The rest are completely gone.
When we look back in 100 years, the total amortization cost for the "winner" won't look so bad. The “picks and axes” (i.e. H100s) that soon wore down, but were needed to build the grander vision won't even be a second thought in hindsight.
How long did it take for 9 out of 10 of those rail lines to become nonviable? If they lasted (say) 50 years instead of 100, because that much rail capacity was (say) obsoleted by the advent of cars and trucks, that's still pretty good.
Records from the time are few and far between, but, from what I can tell, it looks like they likely weren't ever actually viable.
The records do show that the railways were profitable for a short while, but it seems only because the government paid for the infrastructure. If they had to incur the capital expenditure themselves, the math doesn't look like it would math.
Imagine where the LLM businesses would be if the government paid for all the R&D and training costs!
What killed them was the same thing that killed marine shipping — the government put the thumb on the scale for trucking and cars to drive postwar employment and growth of suburbs, accelerate housing development, and other purposes.
The age of postwar suburb growth would be more commonly attributed to WWII, but the records show these railroads were already losing money hand over fist by the WWI era. The final death knell, if there ever was one, was almost certainly the Great Depression.
But profitable and viable are not one and the same, especially given the immense subsidies at play. You can make anything profitable when someone else is covering the cost.
National infrastructure is always subsidized and is never profitable on it's own. UPS is the largest trucking company, but their balance sheet doesn't reflect the costs of enabling their business. The area I grew up in had tarred gravel roads exclusively until the early 1980s -- they have asphalt today because the Federal government subsidizes the expense. The regulatory and fiscal scale tipped to automotive and to a lesser extent aircraft. It's arguable whether that was good or bad, but it is.
State-level...? You're starting to sound like the other commenter. It's a big world out there.
> National infrastructure is always subsidized
Much of the network was only local, and mostly subsidized by municipal governments.
Actually, governments in the US rarely actually provided any capital to the railroads. (Some state governments did provide some of the initial capital for the earliest railroads). Most of federal largess to the railroads came in the form of land grants, but even the land grant system for the railroads was remarkably limited in scope. Only about 7-8% of the railroad mileage attracted land grants.
Did I, uh, miss a big news announcement today or something? Yesterday "around my parts" wasn't located in the US. It most definitely wasn't located in the US when said rail lines were built. Which you even picked up on when you recognized that the story about those lines couldn't have reasonably been about somewhere in the US. You ended on a pretty fun story so I guess there is that, but the segue into it wins the strangest thing ever posted to HN award. Congrats?
H100s are effectively consumables used in the construction of the metaphorical rail. The actual rail lines had their own fare share of necessary tools that retained little to no residual value after use as well. This isn't anything unique.
This remains to be seen. H100 is 3 years old now, and is still the workhorse of all the major AI shops. When there's something that is obviously better for training, these are still going to be used for inference.
If what you say is true, you could find a A100 for cheap/free right now. But check out the prices.
Edit: https://getdeploying.com/reference/cloud-gpu/nvidia-a100
Bulk pricing per KWH is about 8-9 cents industrial. We're over an order of magnitude off here.
At 20k per card all in price (MSRSP + datacenter costs) for the 80GB version, with a 4 year payoff schedule the card costs 57 cents per hour (20,000/24/365/4) assuming 100% utilization.
This is definitely not true, the A100 came out just over 5 years ago and still goes for low five figures used on eBay.
Are we? I was under the impression that the tracks degraded due to stresses like heat/rain/etc. and had to be replaced periodically.
I am an avid rail-to-trail cycler and more recently a student of the history of the rail industry. The result was my realization that the ultimate benefit to society and to me personally is the existence of these amazing outdoor recreation venues. Here in Western PA we have many hundreds of miles of rail-to-trail. My recent realization is that it would be totally impossible for our modern society to create these trails today. They were built with blood, sweat, tears and much dynamite - and not a single thought towards environmental impact studies. I estimate that only ten percent of the rail lines built around here are still used for rail. Another ten percent have become recreational trails. That percent continues to rise as more abandoned rail lines transition to recreational use. Here in Western PA we add a couple dozen miles every year.
After reading this very interesting discussion, I've come to believe that the AI arms race is mainly just transferring capital into the pockets of the tool vendors - just as was the case with the railroads. The NVidia chips will be amortized over 10 years and the models over perhaps 2 years. Neither has any lasting value. So the analogy to rail is things like dynamite and rolling stock. What in AI will maintain value? I think the data center physical plants, power plants and transmission networks will maintain their value longer. I think the exabytes of training data will maintain their value even longer.
What will become the equivalent of rail-to-trail? I doubt that any of the laborers or capitalists building rail lines had foreseen that their ultimate value to society would be that people like me could enjoy a bike ride. What are the now unforeseen long-term benefit to society of this AI investment boom?
Rail consolidated over 100 years into just a handful of firms in North America, and my understanding is that these firms are well-run and fairly profitable. I expect a much more rapid shakeout and consolidation to happen in AI. And I'm putting my money on the winners being Apple first and Google second.
Another analogy I just thought of - the question of will the AI models eventually run on big-iron or in ballpoint pens. It is similar to the dichotomy of large-scale vs miniaturized nuclear power sources in Asimov's Foundation series (a core and memorable theme of the book that I haven't seen in the TV series).
The financials here are so ugly: you have to light truckloads of money on fire forever just to jog in place.
It may be like looking at the early Google and saying they are spending loads on compute and haven't even figured how to monetize search, the investors are doomed.
Oh, I'd love to get a cheap H100! Where can I find one? You'll find it costs almost as much used as it's new.
At some point the AI becomes good enough, and if you're not sitting in a chair at the time, you're not going to be the next Google.
In practice that hasn't borne out. You can download and run open weight models now that are spitting distance to state-of-the-art, and open weight models are at best a few months behind the proprietary stuff.
And even within the realm of proprietary models no player can maintain a lead. Any advances are rapidly matched by the other players.
More likely at some point the AI becomes "good enough"... and every single player will also get a "good enough" AI shortly thereafter. There doesn't seem like there's a scenario where any player can afford to stop setting cash on fire and start making money.
Why?
I don't see why these companies can't just stop training at some point. Unless you're saying the cost of inference is unsustainable?
I can envision a future where ChatGPT stops getting new SOTA models, and all future models are built for enterprise or people willing to pay a lot of money for high ROI use cases.
We don't need better models for the vast majority of chats taking place today E.g. kids using it for help with homework - are today's models really not good enough?
Because training isn't just about making brand new models with better capabilities, it's also about updating old models to stay current with new information. Even the most sophisticated present-day model with a knowledge cutoff date of 2025 would be severely crippled by 2027 and utterly useless by 2030.
Unless there is some breakthrough that lets existing models cheaply incrementally update their weights to add new information, I don't see any way around this.
Humans do this to a minimum degree, but the things that we can recount from memory are simpler than the contents of an entire paper, as an example.
There's a reason we invented writing stuff down. And I do wonder if future models should be trying to optimise for rag with their training; train for reasoning and stringing coherent sentences together, sure, but with a focus on using that to connect hard data found in the context.
And who says models won't have massive or unbounded contexts in the future? Or that predicting a single token (or even a sub-sequence of tokens) still remains a one shot/synchronous activity?
Oh wait, the computer I'm typing this on was manufactured in 2020...
Business looks a lot like what it has throughout history. Building physical transport infrastructure, trade links, improving agricultural and manufacturing productivity and investing in military advancements. In the latter respect, countries like Turkey and Iran are decades ahead of Saudi in terms of building internal security capacity with drone tech for example.
But… I don’t think there’s an example in modern history of the this much capital moving around based on whim.
The “bet on red” mentality has produced some odd leaders with absolute authority in their domain. One of the most influential figures on the US government claims to believe that he is saving society from the antichrist. Another thinks he’s the protagonist in a sci-fi novel.
We have the madness of monarchy with modern weapons and power. Yikes.