The RAM shortage could last years
theverge.com
theverge.com
It doesn't stop there though. OpenAI is currently mired in a capital crunch. Their last round just about sucked all the dry powder out of the private markets. Folks are now starting to ask difficult questions about their burn rate and revenue. It is increasingly looking like they might not commit to the purchase order they made which kick-started this whole panic over RAM.
Soo ... how sure are we that the memory makers themselves are not going to be the ones holding the bag?
We aren't. The remaining memory manufacturers fear getting caught in a "pork cycle" yet again - that is why there's only the three large ones left anyway.
China has memory makers who are creeping up through the stages of production maturity, and once they hit then there's no going back.
If the existing makers can't meet supply such that Chinese exports get their foot in the door, they may find they never get ahead again due to volume - that domestic market is huge so they have scale, and the gaming market isn't going to care because they get anything at the moment, which is all you'll need for enterprise to say "are we really afraid of memory in this business?"
The answer to that is government regulation. Ban anything Chinese or slap it with tariffs. That is what tariffs are intended for - not for the BS the current administration has done.
Wasn't the problem here that OpenAI was negotiating with Samsung and SK Hynix at the same time without the other one knowing about it? People only realized the implications when they announced both deals at once.
There’s virtually infinite capital: if needed, more can be reallocated from the federal government (funded with debt), from public companies (funded with people’s retirement funds), from people’s pockets via wealth redistribution upwards, from offshore investment.
They will be allowed to strangle any part of the supply chain they want.
> more can be reallocated from the federal government (funded with debt)
While this is the most reliable funding, it's still not very accessible. OpenAI is a money pit, and their demands are growing quickly. The US government has started a bunch of very expensive spending. If OpenAI were to require yearly bundles of it's recent "$120B" deal, that's 6% of the US' discretionary budget. 12.5% of the non-military discretionary budget. (And the military is going to ask for a lot more money this year) Even the idea of just issuing more debt is dubious because they're going to want to do that to pay for the wars that are rapidly spiralling out of control.
None of this is saying that the US government can't or wouldn't pay for it, but it's non trivial and it's unclear how much Altman can threaten the US government "give me a trillion dollars or the economy explodes" without consequences.
Further deficit-spending isn't without it's risks for the US government either. Interests rates are already creeping up, and a careless explosion of deficit may well trigger a debt crisis.
> from public companies (funded with people’s retirement funds)
This would be at great cost. OpenAI would need to open up about it's financial performance to go public itself. With it's CFO being put on what is effectively Administrative Leave for pushing against going public, we can assume the financials are so catastrophic an IPO might bomb and take the company down with it. Nobody's going to be investing privately in a company that has no public takers.
Getting money through other companies is also running into limits. Big Tech has deep pockets but they've already started slowing down, switching to debt to finance AI investment, and similarly are increasingly pressured by their own shareholders to show results.
> from people’s pockets via wealth redistribution upwards
The practical mechanism of this is "AI companies raise their prices". That might also just crash the bubble if demand evaporates. For all the hype, the productivity benefit hasn't really shown up in economy-wide aggregates. The moment AI becomes "expensive", all the casual users will drop it. And the non-casual users are likely to follow. The idea of "AI tokens" as a job perk is cute, but exceedingly few are going to accept lower salary in order to use AI at their job.
There's simply not much money to take out of people's pockets these days, with how high cost of living has gotten.
> from offshore investment.
This is a pretty good source of money. The wealthy Arabian oil states have very deep slush funds, extensively investing in AI to get ties to US businesses and in the hope of diversifying their resource economies.
...
...
"Was". Was a good source of money.
Just look at Cuba, which could be a very rich country and one of the prime tourist destinations of the world.
Another point is I often see the money argument - like country X has more money, so they can afford to do more and better R&D, make more stuff.
This stuff comes out of factories, that need to be built, the machinery procured, engineers trained and hired.
[1]https://www.tomshardware.com/tech-industry/semiconductors/ym...
That is, memory capacity is reserved for datacenters yet to be built, but this will do weird things if said datacenter construction is postponed or cancelled altogether.
Are the Netherlands a large proportion of global datacenters?
In most other places the percentage is significantly less than that and then you can easily add more of the cheap-but-intermittent stuff because a cloudy day only requires you to make up a 10% shortfall instead of a 50% one, which existing hydro or natural gas plants can handle without new storage when there are more of them to begin with.
What's more common is that they don't have the transmission capacity itself, but that one's pretty easy in this case too, because what that means is that you have an existing transmission line which is already near capacity with generation on one end and customers on the other. So then you just build the data center on the end of the transmission line where the generation is rather than the end where the existing customers are, at which point you can add new generation anywhere you want -- and if you put it near the existing customers you've just freed up transmission capacity because you now have new customers closer to the existing generation and new generation closer to the existing customers.
> The Dutch power grid is already almost 50% renewables
I was a bit stunned when I read this. Your estimate is very close for 2025 here: https://en.wikipedia.org/wiki/Electricity_sector_in_the_Neth...I calculate about 43.5% was solar or wind. What is way crazier is the "bend in the curve" of production sources in the last 10 years. Look here at how fast solar and wind is growing! https://en.wikipedia.org/wiki/File:Netherlands_electricity_g...
In every country? Citation needed.
That's where AWS us-east-1 is, i.e. the oldest AWS region where they got started to begin with. Google and Microsoft also have a large presence there. It's not just the US government, it's everybody, and it's not new.
> How did they do it?
Here's the US nuclear plant map, guess where a bunch of them are:
https://www.eia.gov/todayinenergy/detail.php?id=65104
The area around Virginia is also a major coal producer and when this was getting started it was a source of cheap electricity, but coal is quickly being replaced with natural gas via pipelines from the Gulf coast. Their current power mix is ~30% nuclear, ~12% renewables (solar) and almost all the rest natural gas.
That's because you don't live in Maryland.
Our energy bills are through the roof and our transmission company is talking about rolling blackouts in 2027.
https://www.thebanner.com/community/climate-environment/cont...
EDIT
The opening paragraph:
> State regulators’ review of the controversial power line proposed to stretch across three rural Maryland counties will extend to at least February 2027, officials announced Thursday, a timeline that prevents developers from meeting the grid operator’s deadline to ensure reliable electricity.
I bet this is pure NIMBYism. Just this phrase alone is a dead giveaway: "controversial power line". LOL: What is controversial about a power line? Hint: They aren't, but NIMBYism exists.If demand increases you have to build more generation and power lines etc. This is not a problem but for NIMBYs, it's just the logical consequence. If the local population increases and you don't have enough grocery stores you don't say "grocery stores are stressed" and regard it as an insurmountable problem, people just open more of them.
The value of an IX isn't just in the IX itself, but also in the presence of hundreds of parties for direct peering, and excellent connectivity to the rest of the world.
It makes a lot of sense to build your DC near one - even if you have no intention of actually participating in the IX itself.
They don't need entire IX worth of connectivity. You're sending mostly text back and forth and any media is in far lower volume than even normal far less dense DC would generate, all the major traffic is inside the AI DC.
All it needs is fiber to nearest IX
Had we done more 10 years ago we would have been better of. The second best time to start is now.
Then why all the anti-coal mining diktats coming down from Brussels?
Brussels is trying to reduce "tiny" to zero, because of this: https://en.wikipedia.org/wiki/Tragedy_of_the_commons
China, like Brussels, is trying to reduce coal for similar reasons. They don't like the air pollution health hazard (fully believable), and they say they don't like global warming (somewhat believable).
Renewables deployment is happening fast. Grid upgrades are not. Batteries .. it depends.
Even nuclear darling France has set solar records: https://www.pv-magazine.com/2026/04/15/france-germany-set-da...
(We used to build it at a fraction of the cost and less than half of the time that we do with our modern fuckups and fuel can come from just about anywhere if need be. It might be a lot more expensive than the stuff kazachstan and still be a fraction of the cost.)
I think ideally we would've done both to press the cost of nuclear down and given the fact that the renewables rollout turned out to be a lot lot more expensive than proponents claimed it would be whilst still tying us up into gass to cover winter.
The problem in the EU is not renewables, it's the same problem that Democratic states in the US face. Regulations and permitting hurdles that block private renewable energy developers.
It says that in 2025, Netherlands was a net exporter of electricity (~14,000 GWh). My guess: Where they want to build data centers, the grid cannot handle it, but the overall system has more than enough power to build data centers. Do you think that sounds like a resonable guess?
Oh no!
If they add enough capacity to meet current demand quickly then if demand crashes they still have billions of dollars in loans used to build capacity for demand that no longer exists and then they go bankrupt.
The biggest problem is predicting future demand, because it often declines quickly rather than gradually.
If you suppose you have cracked the smooth-ramping problem, perhaps you should throw your hat in the ring and soak up all the pent-up demand that SK Hynix, Samsung and Micron are neglecting.
If he can do all that that fast, the RAM makers should be able to at least 1000X their fab capacity on earth in one year. One year for scaling up existing tech is an eternity compared to Elon's timeframe for moon-fabs given the relative complexity of the challenge.
The real issue is everyone wanting to upgrade to hbm, ddr5, and nvme5 at the same time.
If they could make this stuff and sell it to regular people a decade ago for very palatable prices, why do they come up with the idea that this is the technology of the gods, unaffordable by mere mortals?
R9700 has 32GB and is cheaper than most NVidia consumer GPUs, even though it's a "pro".
5090 has 1.8 TB/s?
5090s are certainly expensive compared to most other GPUs, but not expensive enough to be unobtanium for nearly any professional who could utilize one as part of their job
Even a RTX 5080 has a lower memory throughput than a Radeon VII from 2019, 7 years ago, while being much more expensive.
The memory throughput of GPUs per dollar has regressed greatly during the last 5 years, despite the fact that the widths of the GPU memory interfaces have been reduced, in order to decrease the production costs.
RTX 5080 has a 256-bit memory interface, while the much cheaper Radeon VII had an 1024-bit memory interface. RTX 5080 has almost 4-times faster memories than Radeon VII, but it has not used this to increase the memory throughput, but only to reduce the production costs, while simultaneously increasing the product price.
And it's faster for gaming, I guess? Which is what matters for the typical user.
Anyway you can buy much faster GPUs now than in 2019. They are also much more expensive, yes.
I suppose that most games are limited by computation, so they are indeed much faster on modern GPUs.
However, there are applications that are limited by memory throughput, not by computation, including AI inference and many scientific/technical computing applications.
For such applications, old GPUs with higher memory throughput are still faster.
This is why I am still using an old Radeon VII and a couple of other ancient AMD GPUs with high memory throughput.
Last year I have bought an Intel GPU, which is still slower than my old GPUs, but it at least had very good performance per dollar, competitive with that of the old GPUs, because it was very cheap, while the current AMD and especially NVIDIA GPUs have poor performance per dollar.
That's correct if you're targeting gamers, but local AI inference changes this picture substantially.
Heck, I have a phone with a 16bit memory bus for instance. The high(ish) clock rate only makes up the difference slightly.
But with general prices on all components going up, it might not be such a big factor any more.
HBM migght make sense for higher end products which can free up space for the lower end that will never use the tech.
Designing a part with a wide bus and putting the traces down on the board is what I would expect to be the easy part these days (surely).
But yield, yield comes for us all.
AMD Hawaii GPUs still had 1:2 FP64:FP32, while the consumer variant of Radeon VII dropped to 1:4. The following AMD consumer GPUs dropped the FP64 performance to levels that are not competitive with CPUs.
Nowadays the only consumer GPUs with decent FP64 performance are the Intel Battlemage GPUs, which have a 1:8 performance ratio, which provides very good performance per dollar.
because the gods want it all and are willing to pay top dollar.
I wonder whether this is some kind of a racket.
No.
The GB202 die that's in the GDDR7 based RTX 5090 and RTX 6000 Pro literally needed to be this big to support the 512bit memory bus. It's probably only getting worse with smaller node sizes. (see https://www.youtube.com/watch?v=rCwgAGG2sZQ&t=65s).
BTW: The 1TB/s is matched by RTX4090 and surpassed by the RTX5090 (1,79 TB/s).
I'd absolutely buy another hbm consumer GPU if it had at least 8gb (and if I got the vibe/hope AMD will actually support for a couple years...)
We've been projecting both FTL and AGI as future possibilities for almost 100 years now. Do LLMs get us a lot closer to AGI? I think they get us a little closer and Moore's "law" making compute faster probably is a much bigger factor, but I think we're still a very very long ways away.
I think Douglas Hofstadter satisfactorily answered this question.
> We can't even answer if we have free will or not.
Sure we can, it's just that most people don't like the answer.
He didn't prove anything. It's a theory, just like many others: https://plato.stanford.edu/entries/consciousness-higher/
> Sure we can, it's just that most people don't like the answer.
Again, not proven
Don't mistake continued debate amongst "experts" as being a signal that something is unsettled.
Do recent actions of Open AI give you the impression of a company that believes it is about to attain AGI imminently?
Hell, all it matters to investors is not being left holding the bag in the end so they don't even need to believe it
I am not sure if OpenAI has that. Their edge regarding models is small, their strategy currently seems to be "buy ALL the hardware so nobody else can". Users can quite easily switch to other models.
I hope they do, they did not have to agree to sell so much RAM to one customer. They’ve been caught colluding and price fixing more than once, I hope they take it in the shorts and new competitors arise or they go bankrupt and new management takes over the existing plants.
Don’t put all your eggs in the one basket is how the old saying goes.
They act as a de-facto monopoly and milk us. Why is this allowed?
It started with raegan, and even parties on the “left” in the west believe in it with very few exceptions.
The thing that enables this is pretty obvious. The population is divided into two camps, the first of which holds the heuristic that regulations are "communism and totalitarianism" and this camp is used to prevent e.g. antitrust rules/enforcement. The second camp holds the heuristic that companies need to be aggressively "regulated" and this camp is used to create/sustain rules making it harder to enter the market.
The problem is that ordinary people don't have the resources to dive into the details of any given proposal but the companies do. So what we need is a simple heuristic for ordinary people to distinguish them: Make the majority of "regulations" apply only to companies with more than 20% market share. No one is allowed to dump industrial waste in the river but only dominant companies have bureaucratic reporting requirements etc. Allow private lawsuits against dominant companies for certain offenses but only government-initiated prosecutions against smaller ones, the latter preventing incumbents from miring new challengers in litigation and requiring proof beyond a reasonable doubt.
This even makes logical sense, because most of the rules are attempts to mitigate an uncompetitive market, so applying them to new entrants or markets with >5 competitors is more likely to be deleterious, i.e. drive further consolidation. Whereas if the market is already consolidated then the thicket of rules constrains the incumbents from abusing their dominance in the uncompetitive market while encouraging new entrants who are below the threshold.
How is this more efficient? You'd still be applying all of the inefficient regulatory rules intended to mitigate a lack of competition to the smaller companies trying to sustain a competitive market, and those rules are much more deleterious for smaller entities than higher tax rates.
If you have $100M in fixed regulatory overhead for a larger company with $10B in profit, it's only equivalent to a 1% tax. The same $100M for a smaller company with $50M in profit is a 200% tax. There is no tax rate you can impose on the larger company to make up for it because the overhead destroys the smaller company regardless of what you do to the larger one.
As a counterpoint: Look at very high value goods, like jet engines and MRI machines. I went for an MRI the other day and wondered to myself (then asked an LLM) what the international MRI market looks like. They are vanishingly small number of manufacturers and are usually dominated by a few international players. How are you going to apply this tax to non-domiciled (international) companies? Also, companies like General Electric, Mitsubishi Heavy, and Seimens are enormous and incredibly diverse. This idea falls apart quickly.
it was clinton who delivered the democrats to wall street and vice versa
Maybe it’s time for a refresher on what neoliberal means? It’s not simply “new liberalism”. Reaganomics was the start of neoliberalism in the US, tho of course it shifted and developed its character further over time into the monster that drives 99% of our problems today
Nobody is "allowing" this. It's a natural property of being both advanced technology and a commodity at the same time.
Recently they had a second price fixing lawsuit thrown out (in the US).
Now with the state of things I'm sure another lawsuit will arrive and be thrown out because the government will do anything to keep the AI bubble rolling and a price fixing suit will be a threat to national security, somehow. Obviously thats speculative and opinion but to be clear, people are allowing it. There are and more so were things that could be done.
(Well that and collusion)
If they had actually been communicating or colluding with each other, they would have put the screws to him, making it harder for OpenAI to assert control over the vast majority of the DRAM market.
Failing that, you'd like to think a regulatory agency somewhere would step in to keep a single player from hosing everybody else, but...
Up until AI there weren't really players being able to gobble 40% of the market so nobody was looking.
I don’t buy it that two of the largest manufacturers of DRAM in the world, from the same country, didn’t know this. Even of you ignore each company’s intelligence teams, that’s also the job of the country’s internal intelligence services, to make sure they know what all companies are doing and then make it so they have the best leverage to gain as much as possible. Both companies would have known “somehow” and played hardball.
By spying?
The companies also do a lot of spying themselves, every bit of info could give them an edge.
You get market signals that the demand is there, you acquire the necessary capital, you spend 5 years to build capacity, but guess what, 5 other market players did the same thing. So now you are doomed, because the market is flooded and you have low cash flow since you need to drop prices to compete for pennies.
Now you cannot find capital, you don't invest, but guess what, neither your competitors did. So now the demand is higher than the supply. Your price per unit sold skyrocketed, but you don't have enough capacity!
Rinse and repeat.
Capitalists claim that this is optimal.
Because that does not happen exactly as you say for all players. The demand signals will be processed and long-term risk is balanced against short-term gain in a distributed fashion, so not everyone will do the same.
It's more optimal than planned economies until we have AI planned economies with realtime feedback, I guess.
Consumers get cheap goods during oversupply and most inefficient companies get elliminated during bust while consolidation leads to economies of scale.
There is an alternative where legislation dampens this behavior but the short term profits will be lower. Hence the hawks don’t like it.
Potentially. Well meaning and thought out legislation still distorts the markets, possibly making things objectively worse.
We don't even expect companies to plan long-term anymore, it's just moving wealth as fast as possible.
That isn't really a change, very few people could ever have been said to be ideological capitalists. (capitalist is not a word with a hard definition, but I'm considering it a different thing than the more modern pure libertarian zero-regulation ideology)
This is distinct from someone who is a proponent of capitalism as a system, which appears to be the way you are using capitalist. For which I don't blame you.
If anything, it shows it's possible for you to arbitrage this and in doing so help "smooth out the cycle."
At least with capitalism you have many different people with different perspectives on the risk making independent bets. That mitigates the more extreme negative outcomes.
> Capitalists claim that this is optimal.
Compared to starving under communism coz someone at top got the number wrong, yes. And it only really happens when there are massive, unpredictable market movements and governments not doing their job. Govt should look at the whole thing and just say "no", blame them.
No market system self regulates well enough, and it's government job to file down the edge cases like this. But the revolution happened in country which has two utterly incompetent parties, both in pockets of billionaires, fighting for power, and the clowns from one that won last battle use AI to smokescreen the economic growth their actions cratered
That's not what capitalists claim. Capitalists claim communism is responsible for tens, if not hundreds, of millions of death due to famine. And overall miserable ways of life, as if humans were termites deprived of any individualism and any freedom.
Capitalists claim that France importing 500 000 people from the third world each year, out of which only 10% are ever going to work (and these are official numbers) and yet offering them all a safety net is unsustainable. And that that socialism is only going to lead to one thing: running out of taxpayers' money.
Capitalists don't claim they have the best system: what they claim is that they haven't seen a less worse one.
It is unlike a socialist system in that there are signals to read in the first place. What, do socialists claim that failure is optimal when you can't even tell if you've failed?
The Fiji XT architecture after it had 512GB/S on a 4096b HBM bus in 2015.
The Vega architecture did have 400GB/s or so in 2017, which was a bit of a downgrade.
Very few applications other than GPUs need HBM.
At least as I understand it.
Edit: also, that demand pressure is going to be applied constantly; there isn’t going to be a shock, it’s just going to keep prices high longer.
OpenAI (or whoever) crashes and can't pay for the order leaving the memory makers in a tough spot.
Oh noes! Think of a poor memory makers!
The amount of money flowing both from the AI bubble and from quite literally scalping both the server and consumer market... They gambled on the opportunity and if they fail - it's their problem.
Capitalists did their gamble things. If they fail in that gamble what forbids them to sell the regular RAM they made for AI bubbleists to the regular consumers? Besides HBM it's just the regular chips which are exactly the same for the consumer/server market, why it would be any different?
The memory makers specifically did not scale up capacity to avoid being left holding the bag.
this view isn't updated correctly post-claude code and codex. there will clearly be sufficient demand.
The customer ran out of money. In terms of where you are in line of debtors when you haven't even delivered the product to a customer, it's so far back as to be assured you won't get your money.
If the memory makers got a deposit from OpenAI as part of this deal, that is likely to be the only money they will get for any undelivered memory, particularly if OpenAI runs out of capital.
The specific mix of factors could change at any time, but the supply chain is relatively inelastic, it will take some time to show up on price labels.
Are they really such a big RAM buyer?
> OpenAI’s rapid growth, fueled by the success of ChatGPT and other AI products, led to a landmark agreement in October to purchase 900,000 DRAM wafers per month from Samsung and SK Hynix—amounting to roughly 40% of global supply. This surge in demand, coupled with limited manufacturing capacity, sent prices for memory kits skyrocketing. [0]
[0]: https://peq42.com/blog/openai-canceling-many-large-purchase-...
If I booked half a hotel's rooms then suddenly said "yeah never mind. Half my friends cancelled and we're not staying", basically any hotel would be coming at me for my money because there's no way they can fill their rooms now and they're losing revenue. But OpenAI can really get the whole world to pivot towards it then say "cool but we don't need your product anymore" and RAM makers are just going to let it go.
Whoever decided that was a good idea needs to be fired and publicly shamed.
if anything, OpenAi might be in on it
Which also explains why production is falling behind demand, companies aren't going to sink billions into creating product for a market that could dry up overnight.
It's worse. HBM have lower yields so they are essentially making less GB per wafer too
The others...they did not. Memory makers won't be holding the bag because Apple/Google/Samsung are contractually obligated to purchase, after the panic OAI caused.
Given that TurboQuant results in a 6x reduction in memory usage for KV caches and up to 8x boost in speed, this optimization is already showing up in llama.cpp, enabling significantly bigger contexts without having to run a smaller model to fit it all in memory.
Some people thought it might significantly improve the RAM situation, though I remain a bit skeptical - the demand is probably still larger than the reduction turboquant brings.
That is the sad reality of the future of memory.
Given the current tech, I also doubt there will be practical uses and I hope we’ll see the opposite of what I wrote. But given the current industry, I fully trust them so somehow fill their hardware.
Market history shows us than when the cost of something goes down, we do more with the same amount, not the same thing with less. But I deeply hope to be wrong here and the memory market will relax.
Current "TurboQuant" implementations are about 3.8X-4.9X on compression (w/ the higher end taking some significant hits of GSM8K performance) and with about 80-100% baseline speed (no improvement, regression): https://github.com/vllm-project/vllm/pull/38479
For those not paying attention, it's probably worth sending this and ongoing discussion for vLLM https://github.com/vllm-project/vllm/issues/38171 and llama.cpp through your summarizer of choice - TurboQuant is fine, but not a magic bullet. Personally, I've been experimenting with DMS and I think it has a lot more promise and can be stacked with various quantization schemes.
The biggest savings in kvcache though is in improved model architecture. Gemma 4's SWA/global hybrid saves up to 10X kvcache, MLA/DSA (the latter that helps solve global attention compute) does as well, and using linear, SSM layers saves even more.
None of these reduce memory demand (Jevon's paradox, etc), though. Looking at my coding tools, I'm using about 10-15B cached tokens/mo currently (was 5-8B a couple months ago) and while I think I'm probably above average on the curve, I don't consider myself doing anything especially crazy and this year, between mainstream developers, and more and more agents, I don't think there's really any limit to the number of tokens that people will want to consume.
For example Gemma 4 32B, which you can run on an off-the-shelf laptop, is around the same or even higher intelligence level as the SOTA models from 2 years ago (e.g. gpt-4o). Probably by the time memory prices come down we will have something as smart as Opus 4.7 that can be run locally.
Bigger models of course have more embedded knowledge, but just knowing that they should make a tool call to do a web search can bypass a lot of that.
I hate to mention Jevons paradox as it has become cliche by now, but this is a textbook such scenario
> Given that TurboQuant results in a 6x reduction in memory usage for KV caches
All depends on baseline. The "6x" is by stylistic comparison to a BF16 KV cache; not a state of the art 8 or 4 bit KV cache scheme.
mind that you're quoting marketing material that's largely based on unfair baseline testing (like comparing 4 bit vs 32 bit to get "8x speed")
Supposedly AI drives down the cost of producing software,not the "price".
> How are software companies going to make enough revenue to pay for AI, when the amount of money being spent on AI is already multiples of the current total global expenditure on software?
Currently, the cost of AI is between $20/month and around $200/month per developer.
I think the huge billions you're seeing in the news are the investment cost on AI companies, who are burning through cash to invest in compute infrastructure to allow both training and serving users.
> This demand for RAM is built on a foundation of sand, there will be a glut of capacity when it all shakes out.
Who knows? What I know is that I need >64GB of RAM to run local models, and that means most people will need to upgrade from their 8Gb/16GB setup to do the same. Graphics cards follow mostly the same pattern.
Depends how big the models are, how fast you want them to run and how much context you need for your usage. If you're okay with running only smaller models (which are still very capable in general, their main limitation is world knowledge) making very simple inferences at low overall throughput, you can just repurpose the RAM, CPUs/iGPUs and storage in the average setup.
You can run huge local models slowly with the weights stored on SSDs.
Nowadays there are many computers that can have e.g. 2 PCIe 5.0 SSDs, which allow a reading throughput of 20 to 30 gigabyte per second, depending on the SSDs (or 1 PCIe 5.0 + 1 PCIe 4.0, for a throughput in the range 15-20 GB/s).
There are still a lot of improvements that can be done to inference back-ends like llama.cpp to reach the inference speed limit determined by the SSD throughput.
It seems that it is possible to reach inference speed in the range from a few seconds per token to a few tokens per second.
That may be too slow for a chat, but it should be good enough for an AI coding assistant, especially if many tasks are batched, so that they can progress simultaneously during a single read pass over the SSD data.
Batching inferences doesn't necessarily help that much since as models get sparser the individual inferences are going to share fewer experts. It does always help wrt. shared routing layers, of course.
Claude Max subscriptions have gone up, but do you think every Netflix user will pay for one?..
https://www.tomshardware.com/tech-industry/artificial-intell...
the hope is that Ai is "the next semiconductor" and "the next internet"
Not exactly.
LLMs are already quite useful today if you use them as a tool, so they are there to stay. The remaining problem is scalability, a.k.a. how to make LLMs cheap to use.
But scalability is not really a requirement when you look the bigger picture. If smaller software company/projects can't afford to use AI, the bigger ones might just. Eventually they will discover variable use cases for such tech, even if it only serves big firms i.e. defense, resource extraction, war, finance etc.
To the other end, if scalability is achieved, the use of LLM products will be cheaper too, so smaller project can also use them. But of course, if LLM usage is too cheap, then many were-to-be-consumers will just create software projects by themselves at their homes.
The advanced workflows such as programming are not the only use cases for LLMs. There are much simpler work such as answering simple questions or arrange simple tasks. Both will improve productivity. And more advanced model should improve productivity even more.
RAM is built on a foundation of sand.
I would like a source for that statement. Additionally, I want to know by who? Because it certainly isn't end users. Inflating token usage doesn't make it any more economically viable if your user base, b2b or not, hasn't increased with it. On the contrary, that is a worse scenario for providers.
The recent enterprise revenue numbers of Anthropic
1. As a consultant pretty much every company I have worked with in the last 2 years are doing some kind of in-house "AI Revolution", I'm talking making "AI Taskforce" teams, having weekly internal "AI meetings" and pushing AI everywhere and to everyone. Small companies, SMEs and huge companies. From my observation it is mainly due to C-level being obsessed by the idea that AI will replace/uplift people and revenue will grow by either replacing people or launching features 10x quicker.
2. Did you see software job-boards recently? 9/10 (real) job listings are to do with AI. Either it is fully AI company (99% thin wrapper over Anthropic/OpenAI APIs) or some other SME that needs some AI implementations done. It is truly a breath of fresh air to work for companies that have nothing to do with AI.
The biggest laugh/cry for me are those thin wrappers that go down overnight - think all the "create your website" companies that are now completely useless since Ahtropic cut the middleman and created their own version of exactly that.
I know plenty of engineers being forced to use these tools whether they want to or not. A lot of which are okay with using AI liberally, but don't particularly like generative AI and see it as pretty irresponsible (which feels more true by the week and it is clear from first hand experience). I don't know, there is a huge gradient of users, but I would argue that in previous revolutionary technologies, we didn't have to force people to use a good tool. I didn't have to be forced to use Google search or Google Maps, tech that is now ubiquitous with western society. It seems really suspect that suits have to enforce the use of something that is supposed to change the way we work and be a force multiplier.
C-level strongly believes that AI will fix all these issues. They believe that AI will fix their broken processes.
I see strong resemblance with "Agile Development" ~15 years ago. Extremely hyped, noone asked if their org even is a fit for it or need it, and most importantly - the only way to fix agile is to do more agile. Same with AI right now.
Ton of software out there where optimisation of both memory and cpu has been pushed to the side because development hours is more costly than a bit of extra resource usage.
Pressure to optimize can more often imply just setting aside work to make the program be nearer to being limited by algorithmic bounds rather than doing what was quickest to implement and not caring about any of it. Having the same amount of time, replacing bloated abstractions with something more lightweight overall usually nets more memory gains than trying to tune something heavy to use less RAM at the expense of more CPU.
From what I understand, increasing cache locality is orthogonal to how much RAM an app is using. It just lets the CPU get cache hits more often, so it only relates to throughout.
That might technically offload work to the CPU, but that's work the CPU is actually good at. We want to offload that.
In the case of Electron apps, they use a lot of RAM and that's not to spare the CPU
Cache misses mean CPU stalls, which mean wasted CPU (i.e. the CPU accomplises less than it could have in some amount of time).
> In the case of Electron apps, they use a lot of RAM and that's not to spare the CPU
The question isn't why apps use a lot of RAM, but what the effects of reducing it are. Redcuing memory consumption by a little can be cheap, but if you want to do it by a lot, development and maintenance costs rise and/or CPU costs rise, and both are more expensive than RAM, even at inflated prices.
To get a sense for why you use more CPU when you want to reduce your RAM consumption by a lot, using much less RAM while allowing the program to use the same data means that you're reusing the same memory more frequently, and that takes computational work.
But I agree that on consumer devices you tend to see software that uses a significant portion of RAM and a tiny portion of CPU and that's not a good balance, just as the opposite isn't. The reason is that CPU and RAM are related, and your machine is "spent" when one of them runs out. If a program consumes a lot of CPU, few other programs can run on the machine no matter how much free RAM it has, and if a program consumes a lot of RAM, few other programs can run no matter how much free CPU you have. So programs need to aim for some reasonable balance of the RAM and CPU they're using. Some are inefficient by using too little RAM (compared to the CPU they're using), and some are inefficient by using too little CPU (compared to the RAM they're using).
Yeah, I was saying CPU cache hits would result in better performance. The creator of Zig has argued that the easiest way to improve cache locality is by having smaller working sets of memory to begin with. No, it's not a given this will always work in every case. You can reduce working memory and not have better cache locality. But in a general sense, I understand why he argues for it.
> So programs need to aim for some reasonable balance of the RAM and CPU they're using
I agree with this, but
> but if you want to do it by a lot, development and maintenance costs rise and/or CPU costs rise, and both are more expensive than RAM, even at inflated prices
I would like you to clarify further, because saying CPU costs are more expensive than RAM costs is a bit misleading. A CPU might literally cost more than RAM, but a CPU is remarkably faster, and for work done, much cheaper and more efficient, especially with cache hits.
You had originally said
> It could be effective in some specific situations, but I would definitely not say that those situations are more common than the other ones
This is what I'm confused on. Why do you think most cases wouldn't benefit from this? Almost every app I've used is way on one end of the spectrum with regards to memory consumption vs CPU cycles. Don't you think there are actually a lot of cases where we could reduce memory usage AND increase cache locality, fitting more data into cache lines, avoiding GC pressure, avoiding paging and allocations, and the software would 100% be faster?
Andrew is not wrong, but he's talking about optimisations with relatively little impact compared to others and is addressing people who already write software that's otherwise optimised. More concretely, keeping data packed tighter and reducing RAM footprint are not the same. The former does help CPU utilisation but doesn't make as big of an impact on the latter as things that are detrimental to the CPU (such as switching from moving collectors to malloc/free).
> Why do you think most cases wouldn't benefit from this?
The context to which "this" is referring to was "Reducing your RAM consumption is not the best approach to reducing your RAM throughput is my point." For data-packing, Andy Kelley style, to reduce the RAM bandwidth, the access patterns must be very regular, such as processing some large data structure in bulk (where prefetching helps). This is something you could see in batch applications (such as compilers), but not in most programs, which are interactive. If your data access patterns are random, packing it more tightly will not significantly reduce your RAM bandwidth.
I'm getting lost. What are we talking about if not that? Because if you're talking about unoptimized software, you can absolutely reduce RAM consumption without putting extra load on the CPU. Using a language that doesn't box every single value is going to reduce RAM consumption AND be easier on the CPU. Which is what most people are talking about on this post.
> The context to which "this" is referring to was "Reducing your RAM consumption is not the best approach to reducing your RAM throughput is my point."
I'm more interested in the original claim, which was
> Using a lot less RAM often implies using more CPU
There are a lot of apps using a lot of RAM, and it's not to save CPU. So where is "often" coming from here? I think there are WAY more apps that could stand to be debloated and would use less CPU.
It feels like you're coming at this from a JVM perspective. Yeah, tweaking my JVM to use less RAM would result in more CPU usage. But I don't think there's a single app out there as optimized as the JVM is. They use more RAM for other reasons.
> If your data access patterns are random, packing it more tightly will not significantly reduce your RAM bandwidth
Packing helps random access too. A smaller working set means more of your random accesses land in cache. Prefetching is one benefit of packing, but cache and TLB pressure reduction is the bigger one, and it applies regardless of access pattern
What popular language does that? I admit that rewriting the software in a different language could lead to better efficiencies on all fronts, but such massive work is hardly "an optimisation", and there are substantial costs involved.
But more importantly, I don't think it's right. Removing boxing can certainly have an impact on RAM footprint without an adverse effect on CPU, but I don't think it's a huge one. RAM footprint is dominated by what data is kept in memory and the language's memory management strategy (malloc/free vs non-moving tracing collectors vs moving collectors), and changing either one of these can very much have an adverse effect on CPU.
> There are a lot of apps using a lot of RAM, and it's not to save CPU. So where is "often" coming from here?
That the developers may not be conscious of the RAM/CPU tradeoff doesn't mean it's not there. Keeping less data in memory (and computing more of it on demand) can increase CPU utilisation as can switching from a language with a moving collector to one that relies on malloc/free.
> Packing helps random access too. A smaller working set means more of your random accesses land in cache.
Unless your entire live set fits in the cache, what matters much more is the temporal locality, not the size of the live set. If your cache size is 50MB, a program with a 1GB live set could have just as many or just as few cache misses as a program with a 100MB live set. In other words, you could reduce your live set by a factor of 10 and not see any improvement in your cache hit rate, and you can improve your cache hit rate without reducing your live set one iota.
For example, consider a server that caches some session data and evicts it after a while. Reducing the allowed session idle time can drastically reduce your live set, but it will barely have an effect on cache locality.
Tighter data layouts absolutely improve cache behaviour, but they don't have a huge effect on the footprint. Coversely, what data is stored in RAM and your memory management strategy have a large effect on footprint but they don't help your cache behaviour much. In other words, Andy Kelley's emphasis on layout is very important for program speed, but it's largely orthogonal to RAM footprint.
> What popular language does that?
Other than C, Rust, Go, Swift? C# can use value types, Java cannot. So famously that Project Valhalla has been highly anticipated for a long time. Obviously the JVM team thinks this is a gap and want to address it. That is enough in itself to make someone consider a different language.
> I admit that rewriting the software in a different language could lead to better efficiencies on all fronts, but such massive work is hardly "an optimisation", and there are substantial costs involved
That's a pivot to a totally different discussion, which is dev experience. We can say using a different language is not an optimization, I don't care to argue about that. But the fact is some languages have access to optimizations others do not. My dad has 8gb of RAM. I'm not going to install a JavaFX text editor on his computer and explain to him that "it's really quite good value for what the JVM has to do."
> Removing boxing can certainly have an impact on RAM footprint without an adverse effect on CPU, but I don't think it's a huge one
Removing boxing can improve layout, footprint, and CPU utilization simultaneously. That would lie outside the framework "You can't improve one without harming the other."
And it can be a huge effect. Saying it's always a big or small difference is like saying a stack of feathers can never be heavy. It depends on the use case. For a long-running server dominated by caches and session state, sure, although you're not hurting your performance to do it. For data heavy code? The difference between a HashMap<Long, Long> and an equivalent contiguous structure in C# is huge.
>> There are a lot of apps using a lot of RAM, and it's not to save CPU. So where is "often" coming from here? > That the developers may not be conscious of the RAM/CPU tradeoff doesn't mean it's not there
I'm saying Electron uses a lot of RAM and it has nothing to do with offloading work from the CPU, and everything to do with taking the most brute force approach to cross app deployment that we possibly can. I'm not saying anything about the intentions of these developers.
> Unless your entire live set fits in the cache, what matters much more is the temporal locality, not the size of the live set. If your cache size is 50MB, a program with a 1GB live set could have just as many or just as few cache misses as a program with a 100MB live set. In other words, you could reduce your live set by a factor of 10 and not see any improvement in your cache hit rate, and you can improve your cache hit rate without reducing your live set one iota
That's all true. You are fitting more data into each cache line, but your access pattern can be random enough that it doesn't make a difference. It would technically reduce your ram footprint, but as you say, not by much. I only brought this up as an example of something that could reduce RAM footprint without harming CPU utilization, not because it's a worthwhile optimization.
But one way to shrink the live set and improve cache behavior at the same time is to stop boxing everything.
> Other than C, Rust, Go, Swift? C# can use value types, Java cannot. So famously that Project Valhalla has been highly anticipated for a long time. Obviously the JVM team thinks this is a gap and want to address it. That is enough in itself to make someone consider a different language.
As someone working on the JVM, I can tell you we're very much interested in Valhalla and largely for cache-friendliness reasons, but Java certainly doesn't box every value today, and you are severely overstating the case. If you think you can save on both RAM and CPU by preferring a low-level language (or Go, which is slower almost across the board), you're just wrong. But I want to focus on the more important general point you made first.
> My feeling, and the feeling of most people, is that dev experience has been so heavily prioritized that we now have abstractions upon abstractions upon abstractions, and software that does the same thing 20 years ago is somehow leaner than the software we have today. The narrow claim "within a fixed design, reducing RAM often costs CPU," is true.
The problem here is that in some situations there's truth to what you're saying, but in others, it is just seriously wrong. I think the misconception comes precisely because "most poeple" these days don't have the long experience with low level programming that people in my generation of developers do, and you're not aware that many of these abstractions are performance optimisations that come from deep familiarity with the performance issues of low-level programming (I started out programming in C and X86 Assembly, and in the first long job of my career I was working on hard- and soft-realtime radar and air traffic control systems in C++).
Low-level languages aren't meant to be fast (and aren't particularly fast). They're meant to give you direct control over the use of hardware. When it comes to small software, this control does frequently translate to very good performance, but as programs get larger, it makes low-level languages slow. It is true that Java was intended to help developer productivity, but it's also meant to solve some of the intrinsic performance issues in low-level languages, which it does rather well. After all, our team has been made up of some of the world's biggest experts in optimising compilers and memory management, and removing some of C++'s overheads is very much a central goal.
So where do things go wrong for low-level languages? The core problem is that these languages split constructs into fast and slow variants, e.g. static vs dynamic dispatch and stack vs heap allocation. The programmer needs to choose between them. What happens is:
1. As programs grow larger and more complex, the direction is almost completely monotonical in the direction of the more expensive, and more general, variants.
2. There is a big difference between "a fast program could hypothetically be written" and "your program will be fast". Getting good perfomance out of low-level languages requires not only experience, but a lot of effort. For example, you can write a small benchmark and see that malloc/free are pretty fast these days, but that's often true only for the benchmark, where objects tend to be of the same size, and their allocation and deallocation patterns are regular. Memory allocators degrade over time, and they're quite bad when patterns are irregular, which is what happens in real programs, especially large ones. There's also the question of meticulous care around correctness. When Rust first came out I was very excited to see a few important correctness issues solved without loss of control, but was then severely disappointed. Almost anything that is interesting from a performance perspective for us low-level programmers requires unsafe. Even a good hashmap requires unsafe. The performance cost of safety in Rust is higher than it is in Java, and non-experts end up writing slower programs (when they're not small at least).
Such performance issues have plagued low-level programming forever, and Java is reducing these overheads. The idea that high abstractions can improve performance was possibly first stated in Andrew Appel's paper, "Garbage Collection Can Be Faster than Stack Allocation" in the eighties, in which he wrote: "It is easy to believe that one must pay a price in efficiency for this ease in programming... But this is simply not true."
Instead of a static/dynamic dispatch split, Java offers only the general construct (dynamic), and the compiler can "see through" dynamic dispatch and inline it better than any low-level compiler ever could. You can say that surely there has to be some tradeoff, and there is, but not to peak performance. The tradeoff is that 1. you lose control and can't guarantee that the optimisation will be made, so you get good average performance but maybe not the best worst-case performance (which is why it's not hard to beat Java in small programs if you know what you're doing), 2. the compiler needs to collect profiles as the program runs, which results in a "warmup" period.
(If, like me, you like Zig, you might have seen Kelley talk about the "vtable barrier" in low level languages; this doesn't exist in Java. You may also be interested in this talk, "How the JVM Optimizes Generic Code - A Deep Dive", by John Rose: https://youtu.be/J4O5h3xpIY8,
As for memory, not only do moving collectors do not degrade (or fragment) over time, they can use the RAM chip as a hardware accelerator. Unfortunately, when a program uses the GPU for acceleration it's considered clever, but when it uses the RAM chip for accelaration it's considered bloated, even though every CPU core these days comes with at least 1GB of RAM that you might as well use if you're using up the core, as that's effectively free.
The people who consider that bloated are mostly those who haven't struggled with low-level programming long enough or on software that's large enough (they're people who say, I wrote this lean and fast gizmo by myself in 5 months; 99% of value delivered by software is in software written by large teams and maintained over many years). When I was working on a sensor-fusion and air-traffic control software in the nineties, it wasn't "lean"; we just had no choice. We constantly had to sacrifice performance for correctness. Of course, once machines got better, we switched to Java for better performance. God could have written a faster program of that size in C++, but not a large team made up of people with different levels of experience. People who think C++ (or Rust) is particularly efficient are people who haven't written anything big and long-maintained with it.
In conclusion:
1. Sometimes layers of abstractions add performance overheads, and sometimes they remove it. It is not generally true that more abstraction/generality have a performance cost, especially when comparing different languages, although it is almost always true within one language (e.g. dyanmic dispatch is never faster than static dispatch, and is often slower, in C++, but dynamic dispatch in Java can be faster than even static dispatch in C++, and the tradeoffs are elsewhere). If you didn't believe that, you'd be writing all your code in Assembly (which is what I did to get the fastest programs in the early nineties, but it's just not generally faster today thanks to good optimisation algorithms in compilers).
2. Low-level languages give you control, not speed. This control typically translates to better performance in small programs and to worse performance in large ones. This performance problem is intrinsic to low-level programming.
> Removing boxing can improve layout, footprint, and CPU utilization simultaneously. That would lie outside the framework "You can't improve one without harming the other."
First, the footprint won't reduce by much. E.g., in Java, boxing could cost you 10% of your footprint, but the RAM-assisted acceleration could be 80% of the footprint.
Second, yes good layouts help CPU utilisation, but today you can't get that without giving up on other things that harm performance. Dynamic dispatch and memory management in C++ and Rust are just too slow, and while Zig can be blazing fast, it's not easy to write large software in it without compromising performance any more than in any other low-level language. I hope that with Valhalla, Java will be the first language to let you enjoy everything at once, but it's not really an option today.
> I'm saying Electron uses a lot of RAM and it has nothing to do with offloading work from the CPU, and everything to do with taking the most brute force approach to cross app deployment that we possibly can.
That developers choose it because it's "brute force approach to cross app deployment" doesn't necessarily mean that it doesn't also offload work from the CPU, but yes, Electron apps are probably very inefficient from some perspectives. But I think this is also overstated by people who are overly sensitive. When we say something is inefficient, it means that we spend on it more than we have to, but what we really mean is that we could spend that resource that we save on something else instead. On my M1 laptop, I comfortably run three electron apps and two browsers simultaneously without much harming the speed at I can, say, compile HotSpot, probably because SSDs are fast enough for virtual memory in interactive GUIs. I can't think of anything else I could use my laptop's resources for if the apps were leaner on RAM. Reducing the consumption of a resource that can't be meaningfully used for other work isn't real efficiency, and if it comes at the expense of anything useful, it's downright inefficient.
I understand the JVM is not only very efficient, but the JIT gives it unique opportunities to optimize where a compiled language couldn't. You may not get those optimizations consistently, but you don't necessarily need to go into that level of minutia.
You also pointed out that these JIT characteristics can be easily gamed against Java in microbenchmarks, so it's not difficult to make Java look slower than it is in a complex application.
That being said, I am not understanding this narrative that low level projects, as they grow, always devolve into an inefficient dynamic soup. The Linux kernel is millions of lines and uses function pointers sparingly and deliberately. SQLite is huge, mature, and almost entirely static. High-frequency trading systems, embedded software, browser rendering engines, database storage layers. There are entire industries of large, long-lived, performance-critical codebases that do not "devolve" into dynamic dispatch.
If you're saying it's just hard to do that, and Java makes it easy to get close enough with its already dynamic model, then fine. But if you're saying this is an inherent problem as low level programs grow, I would like to understand why.
But you also said
> If you think you can save on both RAM and CPU by preferring a low-level language (or Go, which is slower almost across the board), you're just wrong
Really? Ignoring gamed benchmarks, I don't think it's controversial to say Rust consistently beats Java at the same tasks in RAM and CPU. Maybe that's not important to you because they are too small and you're talking about what happens to programs as they grow in complexity. So I'd like to hear more about why you wrote I'm wrong.
I mean Java and Go are pretty much neck and neck here, with Go using way less RAM - https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
> Second, yes good layouts help CPU utilisation, but today you can't get that without giving up on other things that harm performance
Like what? I'm not understanding. You seem to be implying that without boxing we'd be stuck with a lot of dynamic dispatch and fragmented memory, and I'm not seein the connection.
I brought up unboxing because pointer chasing is expensive and trying to make a collection in Java that you can efficiently loop through can be a frustrating thing.
? Does it not box
> Electron apps are probably very inefficient from some perspectives. But I think this is also overstated by people who are overly sensitive
I also have an m1 laptop and can run things fine. But I'm probably not going to budge on that, because I am consistently exposed to people with low RAM systems, and they are forced to use stuff like Teams in their day to day. Yes, I understand it's cross platform and saves on dev time. Nobody likes using WinForms. But I think Electron has been a net negative on the ecosystem of apps for people with ok computers.
Low-level languages are designed for direct and complete control over hardware, and that is also the job of an OS kernel. Their level of abstraction is a perfect match. But the things at which low-level languages are slow - heap allocations and dynamic dispatch - are exactly the things that applications (not kernels) naturally gravitate towards needing over time.
Of course, it's possible to keep redesigning the architecture as the software evolves to avoid low-level languages' slow operations, but that costs a lot. This isn't some new discovery. The motivations for Java's bet on a JIT and moving collectors were a result of seeing what happened with C++: it was very easy to write nice-looking and fast programs. It was very hard and very costly to keep them that way over time.
> SQLite is huge, mature, and almost entirely static
SQLite is not only not huge but is quite small. ~150KLOC.
> I don't think it's controversial to say Rust consistently beats Java at the same tasks in RAM and CPU.
I don't know if it's controversial, but it's certainly very wrong.
Let's look at one of the most famous terrible benchmarks: The Computer Language Benchmarks Game (it's terrible not only because it compares different algorithms, but also because it has no benchmarks that are long-running, none with interesting memory management, and no concurrent benchmarks - the very things most programs today do): https://benchmarksgame-team.pages.debian.net/benchmarksgame/... In all but one, the C++ and Java results are mixed, i.e. some Java entries are faster than some C++ entries and vice-versa, and this is despite the benchmarks penalising JITs and being minuscule, which is where low-level languages shine. This goes to my point about the important difference of "some program can be very fast" vs. "your program will be fast". Low level languages and Java are on different sides of the tradeoff here: low-level languages focus on control, which often means "someone could write fast code", while Java focuses on compiler and runtime optimisations of high abstractions with the goal of making your code fast.
If we look at another famous benchmark, techempower, we see the same thing: Java, Rust, and C++ results are intermixed, despite the benchmarks being small and thus favouring low-level languages: https://www.techempower.com/benchmarks/#section=data-r23
Of course, there aren't cross-language application benchmarks, i.e. benchmarks that measure the performance developers really care about. All I can say is that a developer of one of the world's largest tech companies told us that his new team lead wanted to migrate some service from Java to Rust for the performance. What happened was that they experienced a large drop in performance, but to save face, they spent 6-12 months carefully optimising the Rust code, and in the end managed to match, though not exceed, Java's performance.
C++ and Rust are simply not particularly fast for applications, and Java is. It's possible to spend a lot of effort optimising them, but it's effort that needs to be spent continuously as the program evolves. That's exactly what led compilation and memory management experts to design the JVM the way they did in the first place: It's hard to make low-level code efficient for large applications.
> I mean Java and Go are pretty much neck and neck here, with Go using way less RAM
Go uses way less RAM because it uses an inefficient non-moving collector, which is why you see Go shops complaining constantly about the poor performance of Go's GC and why they try to avoid it (as Java developers used to do in the past). The speed is similar only because the benchmarks are not very interesting, but while, broadly speaking, C++, Java, and Rust are roughly at the same time "performance level" (ignoring all the tradeoffs I mentioned before), Go is strictly in a lower class. While you have to get pretty large to see Java beating C++ and Rust, it's fairly easy to see Java leaving Go in the dust even on fairly small programs. The programs just need to be a little more interesting than those in the Benchmarks Game.
But I don't think Go is even playing the same game. Its goal wasn't to be a super-optimised language that takes advantage of progress in compilation and memory management technologies. It was meant to be good enough for some things while keeping a small and simple implementation. It's faster than Python and JS, and that's the goal. It's not really trying to compete with C++/Java on performance.
> Like what? I'm not understanding. You seem to be implying that without boxing we'd be stuck with a lot of dynamic dispatch and fragmented memory, and I'm not seein the connection.
I'm saying that the languages that give you good layout today happen to be languages that are bad at other things (like memory management, dynamic dispatch, and concurrent data structures). So if you win in one area you lose in another (but depending on the program, some of these areas may matter more than others).
> Does it not box?
In Java, valus in an int/long/double/etc. array or fields in a class like `class A { int a, b; boolean c; String d; }` are just as boxed as they are in C++, which is to say they're not. Instances of the class will not be flattened into arrays or fields, which is exactly why we have Valhalla, but the problem is not that severe in big program (which is why we haven't dropped everything to just do Valhalla). Also, remember that boxing has a cost in low-level languages beyond cache-locality - due to heap allocations - that don't exist (at least not as significantly) in Java. Boxing in Java is much cheaper than it is in C++/Rust, except fot the cache locality cost, but while in some programs that can be a problem, in many that's not the main one.
> I also have an m1 laptop and can run things fine. But I'm probably not going to budge on that, because I am consistently exposed to people with low RAM systems, and they are forced to use stuff like Teams in their day to day
Of course if you deploy a program that uses a lot of resource X to machines where X is more restricted than the other resources the program uses, you should optimise the consumption of X.
> But I think Electron has been a net negative on the ecosystem of apps for people with ok computers.
That depends on what else these people want to use their computers for while running an Electron app. By far the largest group of people I've seen complain are people here on HN who like counting MBs rather than look at the overall utilisation picture.
It also compares un-optimised single-thread #8 programs transliterated line-by-line from the same original.
However long (programs run) they never seem to become "long-running".
There's always some programmer who replaces "interesting memory management" with array and int.(Many complaints about Go binary-trees programs seemed to be: they should implement a custom arena.)
What does "no concurrent benchmarks" mean when:
import java.util.concurrent.CyclicBarrier;
> Of course, there aren't cross-language application benchmarksMaybe something like
Most application servers are expected to run without issue for at least a day. Our acceptance tests run high workloads for 1, 7, and 30 days. The longest running Benchmarks Game benchmark doesn't break one minute. You can maybe argue whether long running is 3 hours or 3 days, but under one minute isn't long running by anyone's definition.
> What does "no concurrent benchmarks" mean when: import java.util.concurrent.CyclicBarrier;
I believe it's used to coordinate parallelism. Parallelism (where tasks cooperate) and concurrency (where they compete) result in completely different machine workloads.
> Maybe something like https://link.springer.com/article/10.1186/s12859-019-2903-5
It's obviously more interesting than the benchmarks game as it exercises things in a more realistic way, but as much as I like seeing Java winning as it did in this benchmark [1] (even an ancient version of Java, before the new GC generations and new compiler optimisations) it's still very small, and as a batch program, not very representative of most software people write.
The problem with benchmarks is that they tell you how fast a specific program is (the benchmark itself) but it's very hard to generalise from that result to what you're interested in, unless the benchmark is very similar to your program (microbenchmarks never are; larger benchmarks could be, but the space is large so you need to be lucky).
[1]: It's interesting that they made a common mistake when interpreting the results. The program seems to try to get the CPU to 100%. In this situation it's not hard to see that a program that runs even 1% faster and uses 10x more memory is more memory efficient than a program that's 1% slower and uses 10x less memory. That's because while a program runs at 100% CPU, no RAM can be used for any purpose by any other program. So either way you capture 100% of RAM, but in one case you capture it for less time. This idea is at the core of using RAM chips as hardware accelerators (using up CPU effectively uses up RAM because using RAM requires CPU cycles).
JavaOne long ago, there would be mixed messages: both "So a benchmark that ends in less than 10 sec probably does not measure anything interesting." and in blog post benchmarks "100000000 hashes in 5.745 secs … 100000000 primes in 1.548 secs"
(Goldilocks would know.)
> … different machine workloads…
I'm happy to accept that you didn't mean no parallel programs.
> … very hard to generalise …
Indeed.
I didn't say that short-running benchmarks don't measure anything interesting, only that they don't say much about long running programs, where the same mechanisms can exhibit very different behaviour.
I suppose when you write "because it compares different algorithms" you didn't say that there were no comparisons based on the same algorithm.
We've certainly not attempted to prove that these measurements, of a few tiny programs, are somehow representative of the performance of any real-world applications — not known — and in-any-case Benchmarks are a crock.
If I write a malloc benchmark I may think, oh, this measures the cost of malloc/free. In reality, it only measures the cost for a program whose concurrency, allocation/deallocation patterns, and duration match exactly what I wrote, and bear little resemblance to the numbers I'd get if any of those were different.
So I'm not saying that the Benchmark Game is lying. It is telling the truth about how long those programs ran. It's just that what we can generalise from those benchmarks is even less than what we can from more "interesting" ones, but given that even that is close to nothing anyway, maybe it doesn't matter.
There seem to be people who find those brute facts surprising in themselves.
But a good benchmark suite is one that covers a variety of different problems and/or programs similar to a significant portion of production software. The Benchmark Game is neither, plus it's confusing because it often compare things that measure the sophistication of the algorithm while making it seem it measures something about a language (you don't need to be deceitful to confuse). So no, I don't think it's a good benchmark suite at all.
And makes no claim to be.
Here's something that could reasonably make those claims:
https://dl.acm.org/doi/10.1145/3669940.3707217
Oh! It's only Java.
I know. I don't understand why you think I have a problem with the site's honesty. It's a poor benchmark suite, and it admits it is. We're in agreement.
> Here's something that could reasonably make those claims
I'm not familiar with this paper, but you seem to think I was complaining about false claims, which I wasn't. Benchmarks are problematic these days because results no longer generalise as they did a couple of decades ago, but some benchmarks are of higher quality than others (again, I'm not talking about what they say they are but about what they actually are) by at least covering a wider and possibly more relevant set of use cases, and by offering comparisons that are less confusing.
It presents "DaCapo Chopin, a major release of the DaCapo benchmark suite for Java". It's a benchmark suite. It says so.
> I'm not talking about what they say they are but about what they actually are
“When I use a word,” Humpty Dumpty said in rather a scornful tone, “it means just what I choose it to mean—neither more nor less.”
“The question is,” said Alice, “whether you can make words mean so many different things.”
“The question is,” said Humpty Dumpty, “which is to be master—that's all.”
[1]: In particular, Java was designed to overcome some of the biggest performance issues of low-level languages that has plagued a large number of applications: memory management when objects are of varying sizes and lifetimes, concurrency (especially lock-free data structures), and dynamic dispatch, which grows in use as applications grow in size and complexity. Not a single one of these is covered in the Benchmark Game, which focuses on small, very regular, batch workloads, the very things that low-level languages have always been good at, and none of the areas where the performance of low-level languages has traditionally (and to this day) suffered and which led to different compiler and memory management designs.
It's not that the benchmarks game is not a good benchmark suite, it isn't a benchmark suite.
It's not that the benchmarks game is not a good comparison of language speeds, it's that comparison of "language speeds" is so under-specified as-to-be wishful thinking.
> Java was designed to…
"… build software for the next generation of consumer electronics – think smart toasters, interactive TVs, and other futuristic gadgets." Things change.
>… the very things that low-level languages have always been good at…
Which is why there are people who find those kind-of Java programs being in-any-way comparable, somewhat surprising.
OK, but I was responding to someone who did consider it to be a benchmark suite. As long as we agree it's not a good benchmark suite whatever it considers itself to be, we're in agreement.
> It's not that the benchmarks game is not a good comparison of language speeds, it's that comparison of "language speeds" is so under-specified as-to-be wishful thinking.
With that I completely agree. But if you group results by language, that's exactly what you're inviting, and if your suite of benchmarks or whatever you want to call it covered a wider range of problems, that point could be more easily seen. Let's say that the combination of grouping results by language and covering only a very narrow (and niche) set of problems that also happens to be the sweet spot of some languages that have other significant performance failings in other use cases doesn't exactly help people get the right impression.
Close enough.
> … help people get the right impression.
The target audience wonder "Which programming language is fastest?"
A table or chart sorted by elapsed time is the answer they expect.
The target audience have various (perhaps un-examined) ideas about the question.
The sources and measurements can be a way to examine and discuss some of those ideas.
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
BTW, what we do is compare our suite of micro-benchmarks to our (much smaller) suite of macro-benchmarks. This way we get at least some sense of how relevant the microbenchmarks are (i.e. we're looking at the correlation of the deltas). Some microbenchmarks are more correlated with the macrobenchmarks than others. If an optimisation helps some microbenchmarks that we think are not representative of many programs and doesn't help with any macrobenchmark - we take it out.
Just to give an example, we may want to measure some optimisation that helps some allocation pattern. Sometimes it turns out that if that pattern is diluted by other allocation patterns the program does for other tasks, the advantage is completely erased. Some optimisations in free-list allocators are particularly susceptible to this: if your program allocates only in this specific way, it will be super fast. If, in addition, there are some sporadic allocations that follow a different pattern, then after an hour you'll see performance start to drop.
Hopefully, some of the target audience might try to confirm that programs are what they think of as "comparable".
> Having more domain coverage is easier and more valuable…
So where are the examples of that being done? (It's been decades.)
Whenever people want to get valuable information. As I said, we in OpenJDK have a couple hundred benchmarks, some macro, many micro, which are meant to give a decent coverage of the things that affect performance.
If a website wants to group results by languages, it should think about performance from the perspective of how languages work (which include compilers, linkers, and runtimes).
For example, what compiler/linker optimisations are done can depend a lot on whether the program is in a single compilation unit or multiple (and in the case of C and C++ - it does).
On the runtime front, think about memory management. These mechanisms often have different behaviour depending on whether the objects are of similar size or not, whether they're allocated and freed by multiple threads or a single one, and whether the heap is "young" and unfragmented or old and fragmented.
Another area in runtimes is data structures. Are they single-threaded or concurrent, and if concurrent, how do they behave under low and high contention?
Some mechanisms, in all of these levels, have great performance under some conditions and not so great performance in others, and sometimes where they perform great is actually a condition that is encountered less often in real programs.
If you're asking what multi-lingual benchmark suites offer good coverage - I don't know. But that we don't have good information doesn't mean that it's good to offer bad information. Imagine that in American presidential elections there were no national polls and no polls in most states. Would having a poll only in Alabama or only in California offer good insight into who's likely to win? Probably not, because such a poll offers a very partial view of the situation. Is it better than nothing? Maybe, but not by much, because the outcome in Alabama and California is easy to predict without any polls, so it's only helpful in the most extreme cases.
My point is that bad information is bad information, and if people don't understand how different languages behave under different conditions (e.g. that the optimisations the compiler does can differ depending on whether the program is in a single file or not) then they can get the wrong impression. Imagine that someone has no idea about the regional polarisation in the US, and you tell them, well, there are 50 states, but since we don't have polls for all of them, here's the poll for Alabama. Is that information helpful at all?
In any event, any increase in the coverage makes the information a little better, and because the audience may not know whether multiple benchmarks exercise the same or different behaviour in the language, it's the role of the website to pick problems that trigger the different codepaths in the languages' infrastructure. Otherwise, there's the wrong impression of variety, like saying we don't poll only in Alabama but also in Mississippi. Or it's like testing the structure of a bridge by driving a car across it, and then doing it with ten different car models. Testing a bridge does require variety, but the different car models are not what triggers different conditions for the bridge.
In which case, given: The target audience wonder "Which programming language is fastest?" :there doesn't seem to be support for your claim that: "Having more domain coverage is easier and more valuable…".
> Is it better than nothing? Maybe, but not by much…
The benchmarks game: provisional and modest.
Since programming languages specifically optimise differently for the different conditions I listed above, the "support" for my "claim" is that it's obviously true. No one who implements languages or runtimes will dispute it.
But I don't understand the logical implication. The existence or nonexistence of good information says nothing about the value of the information we do have. If all you know is how much cash I have in my wallet, the fact that no one has ever published how much money I have in my bank account doesn't make the information you have more relevant as an estimate of my wealth. That information is irrelevant regardless of whether or not you have access to the relevant information. That information being available is not what's needed to "support" my "claim" that what you know is irrelevant. All you need to know is how people keep their money.
> The benchmarks game: provisional and modest.
I would say it's more like a website comparing US presidential candidates through polls only in Alabama. A more appropriate description than "provisional and modest" would be that it doesn't actually give us valuable information about the candidates' chances.
If people know how US elections work, such information could be put in context, but I don't know how many programmers understand how languages and runtimes optimise performance. Merely saying it's partial/provisional/modest is insufficient to give people the appropriate context.
Sorry for yet another analogy, but it's just like me telling my boss that my test suite passes but it's incomplete without telling him that I've tested only 20% of the program's functionality.
The website should at least explain just how skewed the results are in favour of conditions that arise in small, short-lived programs without (competitive) concurrency and that larger and/or longer-lived and/or concurrent programs can exhibit very different behaviour. These conditions aren't a small matter. They're among the primary motivations for the huge investment over the last few decades in moving GCs and JIT compilers.
Always.
> … website should at least explain…
Whatever the explanation, it would never be sufficient for you.
As you said -- "If you're asking what multi-lingual benchmark suites offer good coverage - I don't know."
> Whatever the explanation, it would never be sufficient for you.
Since I work on compilers and runtimes I know just how small the coverage offered on that website is. For me, no benchmark suite may ever be enough. Still, offering some of the information that is needed to put the results into context would be better than offering none.
Again, I'm not saying that the website is intentionally misleading and I have nothing against the people behind it. They may have a very good reason for why it's of such low quality. I'm just pointing out that it is.
Hmm. To put a pin in this, you're saying the following is harder to do as an application grows in complexity: - Avoiding lots of little allocations (using arenas, value types) - Avoiding dynamic dispatch
Two things that Java doesn't need to avoid, because it's optimized for it. What I'm unclear on is the perspective that it's inevitable. I don't know what scale of apps you're talking about.
In the case of Rust, it has Arenas, stack allocated structs, and generics via monomorphization. Not only can you avoid both of these things, it doesn't even seem that difficult. If you're saying the borrowchecker just becomes too cumbersome to do for sufficiently large applications, that's fine.
> What happened was that they experienced a large drop in performance, but to save face, they spent 6-12 months carefully optimising the Rust code, and in the end managed to match, though not exceed, Java's performance.
There's really not enough detail here to draw from it. But they got to the same performance as Java with less RAM, at the cost of dev experience. Does that not support what I said? Maybe that 6-12 months for the same performance was way too much, but for a smaller app, that could actually be a worthy tradeoff, no? Like...a desktop app?
> Let's look at one of the most famous terrible benchmarks: The Computer Language Benchmarks Game (it's terrible not only because it compares different algorithms, but also because it has no benchmarks that are long-running, none with interesting memory management, and no concurrent benchmarks - the very things most programs today do)
If Benchmark game is not long enough, dynamic enough, or allocate-y enough, then it's not worth talking about. But what is worth talking about?
This is difficult because you won't accept benchmarks that are too small because they are unfair to the JIT, but you also won't accept applications that are too small. Apparently Sqlite, 150k LOC, is not big enough to be relevant to this discussion. So all we have is anecdotal experience that Java is more performant for large, long lived processes with many contributors. I've certainly read a lot of reports from people rewriting to Rust and getting much leaner applications, but maybe that they weren't working on apps complex enough to force them into dynamic dispatch or many allocations. Or maybe it was pure cope. I don't know because their anecdotes and your anecdotes are not very detailed.
But see how far we have drifted from what the original claim was, which is the claim you cannot reduce RAM consumption without harming CPU utilization. You said Rust isn't particularly fast, but Java is. That's a really strong claim considering we have painted a much narrower scope of when that's true. Most of us aren't working on 3M LOC faang apps. We are working on smaller things, or desktop apps. In those cases, the ratio of RAM consumption to CPU efficiency is much better in Rust than it is in Java. Isn't using Java just a straight up RAM loss for those cases?
> That depends on what else these people want to use their computers for while running an Electron app. By far the largest group of people I've seen complain are people here on HN who like counting MBs rather than look at the overall utilisation picture
That doesn't seem terribly fair. Grandma doesn't have the vocabulary to complain about RAM, true. But her computer is slow, she asked her grandson for help, and her grandson told her to use Spotify in the browser, not download the app. And now she has to be mindful of what she has open, even though 8gb of RAM is actually a lot, we've just lost sight of it.
> Of course if you deploy a program that uses a lot of resource X to machines where X is more restricted than the other resources the program uses, you should optimise the consumption of X
The problem is CPU and RAM usage are fundamentally different. I don't get to know what will be run with my program, so I don't get to know how restricted RAM is. If a computer is CPU limited, at the very least, it won't be terribly busy with programs that aren't being used. But for most programs, RAM is allocated and then just sitting there, whether the program is being used or not. So deploying an Electron app kinda feels like a middle finger to your users, because even though it doesn't need to, it limits the amount of programs they can have open. Not to save on CPU, but because Chromium needs that RAM to work in the first place. It's purely a dev experience decision because people don't know how to ship desktop apps. It's pure waste for the user.
Not really.
First, let's look at stack-allocated structs and think about how much data can live in them. The typical stack size is 2MB, but because the only live data in stacks are in caller functions, we can say that on average, the amount of live data that a stack holds is 1MB. Now, look at how much RAM an application uses in MBs and divide it by the number of threads. Usually, the ratio is much higher than RAM, which means that data in stacks is not a significant portion of the program's data. (Async changes this calculus a bit, but async is extremely limited in Rust as it doesn't allow recursion, FFI, or dynamic dispatch; proper user-mode threads, like the ones in Java and Go actually make stack allocation more useful.)
Now let's look at arenas. Arenas are extremely efficient because they offer a similar RAM/CPU tradeoff knob as moving GCs, but they're not as general (if your allocation pattern supports them well, they're great, but you can't generally use them). But in Rust, things are much worse, because arenas are quite limited; too many standard-library data structures, including strings, vectors, and maps can't easily be plugged into an arena. The only language that gives you arenas' full power (which, again, is not completely general) is Zig. This is one of the reasons hardcore low level programmers find Rust so underwhelming (the other being that too many things that are important in low-level programming, including basic data structures but also benign concurrency, require unsafe).
> But they got to the same performance as Java with less RAM, at the cost of dev experience.
You say "dev experience" as if it's some quality-of-life thing. They traded off a cheap resource, RAM, for an eternal maintenance and evolution cost that would only grow higher as the program grows. And remember that RAM isn't entirely fungible. It's hard to get less than 1GB per core (either in bare metal or in cloud VMs/containers) so using less RAM often saves you $0.
> but for a smaller app, that could actually be a worthy tradeoff, no? Like...a desktop app?
I said that low-level languages can offer good performance in small programs, but many desktop apps aren't small. Claude Code's CLI is over 500KLOC.
> So all we have is anecdotal experience that Java is more performant for large, long lived processes with many contributors
If anything, benchmarks are much more anecdotal. Not only are there fewer benchmarks than applications, but they don't even resemble real programs. But yeah, ever since operations lost their intrinsic costs some 20 years ago - with CPU cache hierarchies, branch prediction, and ILP, more powerful optimising compilers, and more elaborate GCs/memory allocators, the ability to generalise from one program to another is close to nil. So yeah, experience is all we have to go with. Going with the numbers we have (some benchmarks that don't extrapolate) rather than the numbers we need but don't have doesn't help.
I can tell you that the loss of intrinsic operation costs has made our lives as compiler/runtime developers much harder, because we can no longer tell people that this operation is generally fast or generally slow. But that doesn't change the fact that this is our reality.
> what the original claim was, which is the claim you cannot reduce RAM consumption without harming CPU utilization
No. I wrote, and I quote: "Using a lot less RAM often implies using more CPU."
> That's a really strong claim considering we have painted a much narrower scope of when that's true.
Did we? Most of software (measured by the distribution of paid programmers) is in large applications.
> Most of us aren't working on 3M LOC faang apps. We are working on smaller things, or desktop apps
I don't think that's true at all. Forget FAANG. Most software isn't written by software companies at all, but is in-house software (well, Netflix isn't a software company, so I guess it's one FAANG letter). The bulk of software is in things like telecom management and billing, banking and finance, manufacturing control, logistics and shipping, healthcare and hospitality, retail and payment processing, travel, government, defence. 3MLOC is quite typical. People who work on smaller software are overrepresented in Silicon Valley and, I'm guessing, among HN readers, but they're the outliers.
> Grandma doesn't have the vocabulary to complain about RAM, true. But her computer is slow, she asked her grandson for help, and her grandson told her to use Spotify in the browser, not download the app. And now she has to be mindful of what she has open, even though 8gb of RAM is actually a lot, we've just lost sight of it.
Does she, though? I doubt she's running anything intensive in the background, so she's really only using one program at a time, and SSDs are fast enough these days to page in virtual memory when she switches programs, unless the one program she's currently using eats up the 8GB. I agree that if her OS - the one thing she needs to run in the background - is taking up a lot of RAM that could be a problem, but the OS is special. Her computer is slow not because shes using a program that eats up a lot of RAM, but because she's inadvertently running a lot of stuff in the background that shouldn't be running at all (browser plugins? some programs that add themselves as login items?) A Surface Laptop comes with 16GB of RAM. No single program uses even half of that.
> The problem is CPU and RAM usage are fundamentally different. I don't get to know what will be run with my program, so I don't get to know how restricted RAM is.
You'd think that, but that's not the case. I admit that I only recently started thinking deeply about this, thanks to some conversations with a colleague who's one of the world's leading experts on memory management, and it was so eye-opening that I gave a talk about this at the recent Java One (because my colleague wasn't available). There are two sides to this:
1. On the demand side, the key is that the use of RAM necessitates the use of CPU (and vice versa): writing and reading to/from RAM requires CPU, but also we write to RAM only when we expect the program to read it in the future. This means that any CPU we use, takes away the ability of another program to use some RAM (because using RAM requires CPU). To give the basic intuition for this, I mentioned the extrme example of a program that uses 100% of CPU. Such a program effectively captures 100% of RAM no matter how much of it it actually uses because no other program can use any RAM, as no other program has the CPU available to access it. You don't need to know anything about what other programs do. Another way to think about this is that the machine is spent whenever the first of RAM and CPU is exhausted.
2. On the supply side, RAM and CPU - whether in metal or in virtualised hardware - are effectively sold as a package (it's hard to get less than 1GB per core, except on embedded devices these days). Furthermore, both moving GCs and (to a far lesser extent) memory allocators can trade RAM and CPU (sophisticated memory allocators aren't quick to return RAM to the OS and maintain internal buffers).
So even though it is true that different programs may have different CPU/RAM usage patterns, you have to think about the ratio rather than CPU and RAM in isolation, and try to achieve some approximate balance. To put it simply, if a program uses a lot of CPU it doesn't make sense for it to use little RAM, because by using a lot of CPU it is effectively depriving other programs of their ability to use RAM (as that requires CPU). There are some exceptions, such as large caches, but the tradeoffs there are very different and too complicated to go into here (I did cover that in my talk).
> It's purely a dev experience decision because people don't know how to ship desktop apps. It's pure waste for the user.
No. I mean, some of it is probably waste, but:
1. What you call "dev experience" also affects the user because it directly impacts the cost of software. Users want cheaper software.
2. More relevant to this particular discussion is what else the user could do. Having "more programs open" isn't a problem thanks to SSDs and virtual memory. So we're talking about programs that are actively using the CPU for something, and they, too, need a balance of the RAM/CPU ratio.
I'm not trying to be dogmatic in the other direction and assert that Electron is necessarily the best tradeoff. But I'm saying that efficiency is ultimately about money that is spent on a combination of RAM, CPU, and software, and when you look at the full picture you see that it's more complicated than it seems. It's not that the software industry has decided to waste users' money. If it did, there would be a competitive edge to programs that use less RAM, but we don't see that competitive edge. What we do see is a few people on HN saying how they simply can't live with VS Code's 50ms keystroke latency and how amazing is some other editor with only 20ms latency that's likely to go out of business soon [1]. The people who made these decisions aren't some early-career developers who just like hot code reloading or some such.
[1] Yes, I do think Rust is more hardware-efficient than JS, but here I'm looking at an even bigger picture. And yes, if you rewrite from JS to C++/Java/Rust/Go you can win on hardware, but as I said at the very beginning, any such rewrite is not really "an optimisation".
I wish I could find you a few reports from people on here basically renouncing Java because they could not optimize it any further after 20 years programming in it, and moving to Rust. I'd be curious what you'd think.
> You say "dev experience" as if it's some quality-of-life thing. They traded off a cheap resource, RAM, for an eternal maintenance and evolution cost that would only grow higher as the program grows
No, I say dev experience because I brought it up earlier, and staked my claim on it. Because it's an umbrella term that covers nice-to-haves and how ergonomic the language and ecosystem are. The antithesis would be lots of repetitive plumbing that slows down feature release. It's one part of the triangle. They now are shipping way behind but wound up with a product that is just as fast and uses less RAM. That kind of decision can matter to other projects. Smaller projects likely wouldn't have such a slow turnaround.
Again, we're not advocating for everyone to do rewrites. Threads like this are people begging app developers to stop using stuff not appropriate for desktop apps.
> I said that low-level languages can offer good performance in small programs, but many desktop apps aren't small. Claude Code's CLI is over
Err, Claude Code is very special, yes. Couldn't tell you why that tool needs to use that much, but most desktop apps don't. They are built on vendor code and keep the actual app code small, and that makes them excellent candidates for what we're talking about.
> 3MLOC is quite typical. People who work on smaller software are overrepresented in Silicon Valley and, I'm guessing, among HN readers, but they're the outliers
I don't agree, unless you have some stats I don't know about. I mean I'm really jealous that you even know someone that worked on an app of that size. Most people are coding for one of the millions of mid sized businesses dotted all over the country. They are on 20 year old code bases that make great revenue. Everyone there is nice and meets with you weekly. The devs answer to the clients directly. They don't really need to grow their business endlessly, but there's lots of maintenance to do. I've worked with Healthcare companies, they were certainly not 3M lines of code. What you're describing is an extremely narrow class of software that most people will never touch. But I think it sounds cool.
> If anything, benchmarks are much more anecdotal
Not really. I concede they only measure what they do, and it might not be much, but at least they measure something, and it's public and reproducible. Anecdotes are vague stories that are impossible to evaluate. I had an anecdote of someone saying they can't use Java anymore because, even with 20 years of experience, they cannot optimize Java any further for what they need. They rewrote it in Rust. It works much better now, and it's not even close. What am I to make of that?
> Her computer is slow not because shes using a program that eats up a lot of RAM, but because she's inadvertently running a lot of stuff in the background that shouldn't be running at all
There are absolutely apps that run 8gb of RAM. And pageswaps are not good, even with an SSD. They're a real problem.
I just find it lame to tell people this when we could ship leaner apps and it wouldn't even be that hard. It's 2026 and people should be able to have whatever open windows they want. Even 8gb of Ram is a lot, we've just forgotten about it. Shit, my web browser uses 4GB ram idle.
> So even though it is true that different programs may have different CPU/RAM usage patterns, you have to think about the ratio rather than CPU and RAM in isolation, and try to achieve some approximate balance. To put it simply, if a program uses a lot of CPU it doesn't make sense for it to use little RAM, because by using a lot of CPU it is effectively depriving other programs of their ability to use RAM (as that requires CPU). There are some exceptions, such as large caches, but the tradeoffs there are very different and too complicated to go into here (I did cover that in my talk).
That's a cool realization. But I think it's slippery. Only in extreme scenarios will your CPU actually block RAM. The 100% usage scenario makes sense. But most of the time, your CPU is going to be underutilized and capable of letting every app use RAM freely. Obviously the more direct problem would be someone's RAM was sucked up by different apps.
> It's not that the software industry has decided to waste users' money. If it did, there would be a competitive edge to programs that use less RAM, but we don't see that competitive edge. What we do see is a few people on HN saying how they simply can't live with VS Code's 50ms keystroke latency and how amazing is some other editor with only 20ms latency that's likely to go out of business soon
How? If you're forced to use work software, you have no competition to go to. Same for your music app, your social network, your team's chat tool. The "competitive edge" argument requires users to actually have a choice, and for most desktop software they don't they use what their employer, school, or social network has standardized on. Where users do have free choice, they gravitate toward leaner options constantly. Sublime kept paying customers against free Electron alternatives. Mobile platforms enforce resource discipline and have no Electron equivalent. The competitive edge for leanness exists; it just can't express itself when ecosystem effects lock users in.
So I still feel strongly there are cases where you could make sufficiently small programs in Rust that wouldn't devolve into spaghetti. That would give you great performance and ram usage. I still want to find the example of my anecdote, but I'm tired.
> You say "dev experience" as if it's some quality-of-life thing. They traded off a cheap resource, RAM, for an eternal maintenance and evolution cost that would only grow higher as the program grows No, I say dev experience because I brought it up earlier, and staked my claim on it. Because it's an umbrella term that covers nice-to-haves and how ergonomic the language and ecosystem are. The antithesis would be lots of repetitive plumbing that slows down feature release. It's one part of the triangle. They now are shipping way behind but wound up with a product that is just as fast and uses less RAM.
... and will cost much more to evolve, costs will never drop and may well rise. These costs are higher than the RAM they saved, which was free anyway, because it couldn't be used for anything else.
> That kind of decision can matter to other projects. Smaller projects likely wouldn't have such a slow turnaround. Again, we're not advocating for everyone to do rewrites. Threads like this are people begging app developers to stop using stuff not appropriate for desktop apps.
The people begging are sometimes right and sometimes really wrong. They are sensitive to certain things but don't consider the full picture.
> I said that low-level languages can offer good performance in small programs, but many desktop apps aren't small. Claude Code's CLI is over
> Couldn't tell you why that tool needs to use that much, but most desktop apps don't. They are built on vendor code and keep the actual app code small, and that makes them excellent candidates for what we're talking about.
VS Code is something like 5MLOC. Slack is probably around a million. Every desktop app (that isn't bundled with the OS) that I or people I know use are roughly that size or bigger.
> What you're describing is an extremely narrow class of software that most people will never touch. But I think it sounds cool.
First of all, this is the class of software that most people rely on the most by far. It certainly contributes more economic value than other software. You use that software every time you tap your bank card; every time you place or receive a call or send or receive a text message; every time a package arrives at your doorstep; every time you watch any video on any platform; every time you receive medical treatment or stay at a hotel. Clearly, it's the class of software that contributes the majority of value that software delivers. I don't know exactly how many people work on such software. I think around 50% of developers at least, but even if I'm wrong, we're talking at least a few million developers.
> but at least they measure something, and it's public and reproducible.
There is zero value to a measurement that is not relevant to you, and negative value to making you think it's relevant ("well, it's not the number I need but it's some number I have so I'll go by that") when it's not.
> Anecdotes are vague stories that are impossible to evaluate.
You can evaluate them at least as well as you can benchmarks - by asking questions that will allow you to know if the information is relevant to your use case - only they at least have a chance of being relevant, while benchmarks really rarely are. I will just say that if you don't know how exactly a memory allocator is implemented (e.g. whether and how it degrades with time), it is absolutely impossible for you to evaluate a benchmark that purports to measure memory management. There is nothing you can learn from it because you don't really know what it is that's been measured.
> I had an anecdote of someone saying they can't use Java anymore because, even with 20 years of experience, they cannot optimize Java any further for what they need. They rewrote it in Rust. It works much better now, and it's not even close. What am I to make of that?
You are to make of it that, in the past 20 years at least, we have no ability to extrapolate from one program to another. Given that languages like C++, Rust, Java, Zig, and C# are all at the topmost performance category, it makes a lot of sense that in some situations X will be faster than Y and in others Y will be faster than X.
> There are absolutely apps that run 8gb of RAM. And pageswaps are not good, even with an SSD. They're a real problem.
I've not seen that in quite a few years now. Right now I have Chrome open on four streaming platforms. Just over 500MB. I have over 100 tabs open in Safari. About 1.2GB. I'm sure there are some apps that use 8GB of RAM, but these aren't apps that grandma uses or wants to use.
> That's a cool realization. But I think it's slippery. Only in extreme scenarios will your CPU actually block RAM. The 100% usage scenario makes sense.
No. This scales. Again, every CPU cycle you consume takes a cycle away from some other program and reduces its ability to use RAM (as that requires the cycle).
> But most of the time, your CPU is going to be underutilized and capable of letting every app use RAM freely.
No. Using RAM means using CPU. I you're idle, then you're not using RAM. Might as well be paged to SSD. You get zero points for keeping RAM significantly more plentiful than free RAM. You save $0.
> How? If you're forced to use work software, you have no competition to go to. Same for your music app, your social network, your team's chat tool. The "competitive edge" argument requires users to actually have a choice, and for most desktop software they don't they use what their employer, school, or social network has standardized on. Where users do have free choice, they gravitate toward leaner options constantly. Sublime kept paying customers against free Electron alternatives. Mobile platforms enforce resource discipline and have no Electron equivalent. The competitive edge for leanness exists; it just can't express itself when ecosystem effects lock users in.
The competitive edge does not require users to have a choice. It requires a real edge. If some software really makes you more productive, then your employer would be foolish not to buy it. If some software allows the school to significantly save on hardware - the same. The reason these products don't catch on is because they're written by hackers with certain sensitivities who do not understand the economics of their users' hardware and software.
> So I still feel strongly there are cases where you could make sufficiently small programs in Rust that wouldn't devolve into spaghetti. That would give you great performance and ram usage.
Sure. Like I said, low-level programs are fast and efficient for small programs. But even then, you need to look at the full picture to know how much, if any, money you're saving.
> The question isn't why apps use a lot of RAM, but what the effects of reducing it are. Redcuing memory consumption by a little can be cheap, but if you want to do it by a lot, development and maintenance costs rise and/or CPU costs rise, and both are more expensive than RAM, even at inflated prices
That sums up what I've been getting at much better.
> and will cost much more to evolve, costs will never drop and may well rise. These costs are higher than the RAM they saved, which was free anyway, because it couldn't be used for anything else
I guess I'm suggesting If it evolves. My experience with smalller apps is spread across dozens of businesses that are 20+ years old. 1M LOC is not a required thing. In fact, for many businesses, you probably couldn't reach that many LOC without inventing busy work for devs. Sometimes my job is ripping features OUT that a different dev company put in, and nobody knows why anymore.
> The people begging are sometimes right and sometimes really wrong. They are sensitive to certain things but don't consider the full picture
That's a true statement in isolation. I think it's reasonable to not want a web browser embedded in multiple apps on your computer. Slack and Spotify use more RAM than Steam. For what each app does, that seems absurd to me. Again, that's not a bad tradeoff from a development velocity perspective.
> First of all, this is the class of software that most people rely on the most by far. It certainly contributes more economic value than other software.
But the fact that this type of software has more devs is different from saying the average project has the same considerations. I wouldn't tell someone to use Kubernetes because FAANG uses it, and that means a lot of devs use it. If you estimate 50% of developers to be working on this kind of software, I estimate 5% of them has any choice over what tech stack they are using in the first place. So when you are making tech stack recommendations and saying "C++ is not fast for apps", you are talking to the other 50%.
> There is nothing you can learn from it because you don't really know what it is that's been measured
That's true. I downloaded the benchmarks and ran them myself and played around with them. But I lean on others for technical evaluation. My understanding of low level programming ends at toy projects and what I've read about cache/cpu. I imagine if you develop the JVM it's frustrating to continuously talk to people about isolated benchmarks.
Edited out a lot of rhetorical arguing after rereading the conversation
Just to be clear, the cost of maintaining a program in a low-level language is always higher. That's easily the #1 reason the use of low-level languages has been declining steadily for a few decades now with no hint of a change in direction. What happens in large programs is that low-level languages become slow. So yes, if that program doesn't grow, it will probably not become slow, but they've already paid more on development and continue to pay more on maintenance than any savings they could have made on memory, which are probably zero or nearly that.
The point of my explanation about the RAM/CPU relationship is that a well-balanced ratio is free. If your CPU usage amounts to some X% of RAM "captured" any memory savings below it translates to $0 in savings. It's sort of like ink and paper. They're used in combination so reducing the consumption of one without the other doesn't really save you anything.
> I think it's reasonable to not want a web browser embedded in multiple apps on your computer.
I don't know why that would be reasonable unless you can show me it's a waste of money. Maybe it is, but I'm not sure.
> Slack and Spotify use more RAM than Steam. For what each app does, that seems absurd to me.
But software is written to deliver value to users. Most software has no intrinsic value. Sometimes in an economy you get things that may seem absurd - I can't think of a good example, but say that you can only buy rope in units of 1m - but make sense once you consider the entire system. Could Slack use much less RAM than Steam? Of course! Should it, though? I don't know.
> Again, that's not a bad tradeoff from a development velocity perspective.
And again, what you call "development velocity" is not some vanity metric, but something that can translate to actual money savings for the user more than reducing RAM consumption.
> Do you disagree that making the equivalent app in Avalonia, JavaFX, QT would likely use less RAM and CPU than Electron? Is there not room to trip RAM usage in the Desktop world without harming the CPU?
There probably is, but as I said in the beginning, switching a language is a large investment, not exactly an optimisation (and it might not be worth it).
> But the fact that this type of software has more devs is different from saying the average project has the same considerations.
Well, that depends what you mean by "the average project". If we're counting by number of programs/repos, the median project size may well be a 100 line script. We have to weigh it by something. Number of devs and lines of code are probably highly correlated, so either one would do.
> I estimate 5% of them has any choice over what tech stack they are using in the first place
I don't understand the point you're trying to make. I don't really care what someone working on some small website does because getting that tech stack wrong is of little consequence anyway. For software that "matters", the choice of tech also matters, and you're right that the junior developers (and probably many senior developers) working on those projects don't choose the tech, but somebody does, and these are the choices that matter. For example, you care about Slack's tech choice. That was also some high-level decision. If they got it wrong, it wasn't their junior programmers who made the mistake.
> I imagine if you develop the JVM it's frustrating to continuously talk to people about isolated benchmarks.
Yes, but everyone who deals with software performance has been frustrated by this for a long time. Benchmarks used to be at least somewhat more informative until the late '90s. I don't know how to educate developers more about this, but I hope someone manages to do it.
> Well that is a very different statement from what you said earlier, which is "C++ and Rust are simply not particularly fast for applications, and Java is." You have been painting a picture that it is essentially impossible to top Java with Rust except in the narrowest of situations.
It is generally hard to beat Java in large programs. It is always theoretically possible because you can view every Java program as a C++ program (which is what the HotSpot JVM is) running on some data, but it's hard, and I would say close to impossible for similar costs.
I agree with this entirely.
> I don't really care what someone working on some small website does because getting that tech stack wrong is of little consequence anyway.
For server applications and not tools, probably not. Should Curl have a slower startup time? Probably not.
> but they've already paid more on development and continue to pay more on maintenance than any savings they could have made on memory, which are probably zero or nearly that
One key detail of desktop apps is every performance compromise is multiplied across all your users. It doesn't make much business sense to spend hundreds of dev hours to spare your server 4gb of RAM. But it has a much bigger impact across a growing number of users.
In the server case, you are offloading Dev velocity to your own RAM cost. In the Desktop app case, you are offloading Dev velocity to everyone else's ability to run programs on their own computer.
> I don't know why that would be reasonable unless you can show me it's a waste of money. Maybe it is, but I'm not sure
/rant
Well it's not reasonable as a singular business decision. If we look at everything from the lens of how you can make the most money as a company, Electron is probably the route to go right now. You can reach more users even if you upset more as you grow. For software that is targeting the things I want, I care about the health of the company because it determines how fast improvements get to me. So, in isolation, it doesn't upset me if an app uses Electron. I may not have ever had a chance to use the app if they chose something else.
The problem is when everyone does that. If the expectation becomes "everyone has a lot of RAM so just use Electron", then where does that end? Now we need more RAM, even as it gets more expensive, to run apps that are not particularly novel. It's not surprising people are growing frustrated. Dev Velocity is not a vanity metric, but it's not inherently moralistic either. Sure, a company could "scale" faster replacing all of their support staff with an AI chatbot, but that doesn't mean I have to like it.
I have much more sympathy for the small dev team than a large corporation when it comes to using Electron, specifically because their software is smaller. If I vet and use their product, it likely has less features overall, but does more of what I want. I'm more forgiving of their compromises and I can't be as picky because I chose this product.
In the case of apps like Slack, I don't get a choice to use it. I use it for Work, and they develop a lot of stuff we simply don't use. I literally just need it to send text. And so I am a bit less sympathetic to their decision to offload dev velocity costs onto my computer when I don't particularly want to use their app in the first place, and don't believe the majority of those dev hours will be used to help me.
In the case of VSCode, from what I understand, they have to spend a lot of dev time anyway to make Electron work for them. I don't think the value add is as cut and dry when your product needs to run fast.
Of course, but I don't think anyone would consider curl to be of little consequence. There are many small programs that are very important, but in general more value is in larger programs, and I don't think it's hard to see that. A large program costs tens of millions of dollars per year. Companies don't pay that unless the software more than pays for itself. When it comes to small programs, because they're small, competition is also easier. Five different people may identify the same small problem to solve with five different programs. One may end up being consequential, and the rest won't be.
> One key detail of desktop apps is every performance compromise is multiplied across all your users. It doesn't make much business sense to spend hundreds of dev hours to spare your server 4gb of RAM. But it has a much bigger impact across a growing number of users.
Yes, and I don't want to appear as if I claim that, say, Electron isn't a problem. It's just that I'm not sure it's a problem, and I'm trying to say that things are more complicated. If there is a problem, of course it affects many people, but I'm not sure there actually is one. My point about RAM isn't that it's no big deal if you waste it, but that using a lot of it might not be waste at all (in other words, that the RAM you're using is effectively free, as it cannot be used for anything else). So if an Electron program uses 6GB of RAM, and 4 of them are a waste - even if 2 of them are a waste - that's a problem. But even if you can write such a program that only uses 1GB, that doesn't mean that the other 5 are a waste at all. Using less RAM isn't necessarily more efficient if the RAM you saved can't be put to good use.
I'm also making a separate claim that even if some of that RAM is a waste, it could be offset by a lower cost of development, but these are two different claim.
In short, what I'm saying is that it's complicated.
> The problem is when everyone does that. If the expectation becomes "everyone has a lot of RAM so just use Electron", then where does that end? Now we need more RAM, even as it gets more expensive, to run apps that are not particularly novel.
It's not so simple! First, we need more RAM because we have more compute. To some degree it's like ink and paper. You can't enjoy more ink unless you also have more paper. Second, because some RAM can be converted to CPU (through moving collectors or arenas) the overall cost of running some computation can be lower if you buy more RAM. Third, once you already have that RAM, how much of a problem is it if some silly program uses a lot of it? The Electron apps I've seen have little problem being paged out to SSD, and they page in fast (paging in even 5GB takes about 2s, and you usually don't need to page in so much at once).
> It's not surprising people are growing frustrated. Dev Velocity is not a vanity metric, but it's not inherently moralistic either. Sure, a company could "scale" faster replacing all of their support staff with an AI chatbot, but that doesn't mean I have to like it.
Who is growing frustrated? Hackers on HN? If there's actual demand, and if the economics really support the claim that it could and should be done, then alternative products will have a competitive advantage. I'm always dubious when people make claims that seem to me to run counter to how the market behaves. That doesn't necessarily mean they're wrong, but it is a significant point against the claim. If a lot of people think they're paying to much for what they're getting and it's possible to pay less, such a product offering would be a huge success.
> I use it for Work, and they develop a lot of stuff we simply don't use.
Serious question: What would you be using the RAM Slack consumes for while at work?
If you could use that RAM for something more productive or if it meant your work machine could be significantly cheaper, then that's a very good argument. But if it's just about not liking to see a large number of something your boss has already paid for when a smaller number could do, even though they don't really make smaller hardware, then that could explain why there isn't a real pressure to do things differently.
> I literally just need it to send text.
Let's say the job could be done in 500KB and that Slack uses 5GB. But you already paid for 8 or 16 GB of RAM. Unless you could use that 5GB for something better while you're sending the text, why do you care that the number goes up? It doesn't cost you anything.
> In the case of VSCode, from what I understand, they have to spend a lot of dev time anyway to make Electron work for them. I don't think the value add is as cut and dry when your product needs to run fast.
I have no idea why VSCode chose Electron and whether it's a good or bad decision (I don't know how much it played a role, but I think that the ability to write plugins in JS/TS helps them, as there are so many JS/TS developers), but it doesn't bother me because the performance is good enough and it doesn't seem to hinder my use of my machine for anything else I run. If at some point it starts bothering me, I'll look for leaner alternatives.
The fallacy in that reasoning is that a program that's using a lot of CPU (especially if it's a huge MLOC-sized app) is most likely using up its CPU on memory throughput, not pure number-crunching compute! So at least for the enterprise app case (not pure number crunching), you'd actually need a tunable tradeoff between memory throughput and total RAM footprint, and adopting copying/moving GC's just doesn't give you that. Collections cycles are a huge burden on memory throughput: thus, indirectly, on the very thing you're calling "CPU". The theoretical prospect of winning by forgoing collections cycles outright (pure bump arena allocation) is explicitly excluded here since we're talking about long-running programs that will at some point need to garbage collect.
Heap allocation may have marginally higher "CPU" use in the pure compute sense, but that's exactly the kind of CPU use that does trade off successfully with a lower RAM footprint.
Similarly, non-moving concurrent garbage collectors like Go's also successfully navigate this tradeoff compared to moving/copying collectors, because their collection work, while compute- and to some extent memory-traffic intensive (though less so than if copying/moving memory was involved!) can be largely (though not completely - some minor compute overhead on the hot path is still present) shunted off to a lower-priority background thread.
On the other side of the tradeoff, arenas and caches increase memory footprint in a way that's low-impact on memory throughput (unlike pervasive use of a copying/moving GC) because only live data is accessed as needed, and deallocating the arena is a single operation. The tradeoff is actually highly favorable to low-level languages, which commonly use arenas to manage challenges with heap allocation such as fragmentation.
No, there's no such fallacy here because that assumption is not needed for the conclusion. The point is that CPU is needed to use RAM, and so if you use CPU for whatever reason - even to loop for an hour over some integers - you are consuming a resource that is needed to use RAM (by another program). So the use of CPU "captures" RAM whether it uses RAM or not, so it might as well use it.
The extreme example I gave was that a program that uses 100% CPU (again, even if it uses zero RAM) effectively capures 100% of RAM because no other program can use any RAM while that program is running. This extreme example is just to build some intuition, but it scales to lower CPU utilisations.
> Collections cycles are a huge burden on memory throughput
I don't even know where to start. The whole point of moving collectors is that the can make the cost of memory management arbitrarily low, reduing the overhead compared to free list approaches. A collection cycle does a constant amount of work (per program & workload), but the frequency of the collections can be made arbitrarily low. This is memory management 101.
The problem of moving collectors has traditionally been the impact on latency, not on throughput (they were always better on throughput than free lists) - until the advent of pauseless moving collectors.
> Heap allocation may have marginally higher "CPU" use in the pure compute sense, but that's exactly the kind of CPU use that does trade off successfully with a lower RAM footprint.
Except it doesn't given the actual economics of RAM and CPU. You'll need to wait for my talk to be posted to YouTube (I can't reproduce it all here), but in the meantime you can watch this one, by my colleague, which was a keynote at the most recent ISMM (International Symposium on Memory Management): https://youtu.be/mLNFVNXbw7I
The problem is that Erik is one of the world's leading experts on memory management, and he's talking to other experts, so his talk assumes quite a bit of familiarity with the subject. Also, his comparison focuses on tracing GCs, leaving the one to malloc/free implicit.
> Similarly, non-moving concurrent garbage collectors like Go's also successfully navigate this tradeoff compared to moving/copying collectors
Again, except they do not. Go users experience severe problems with the GC that Java users no longer do precisely because of the inefficiencies of their simple GC (the JDK used to have such a GC, but we removed it five years ago when newer, more sophisticated algorithms yielded better results.
> On the other side of the tradeoff, arenas and caches increase memory footprint in a way that's low-impact on memory throughput (unlike pervasive use of a copying/moving GC)
This is simply not true, and shows unfamiliarity with how modern moving collectors actually work (remember that the first open-source high throughput, paseuless moving collector was first released two and a half years ago). Moving collectors offer pretty much the same tradeoff as arenas. The key points are:
1. A generational design makes copying a relatively rare operation to begin with (only a relatively small number of objects are copied).
2. The frequency of collections can be made arbitrarily low.
If the usage happens to be arena-like, i.e. no objects survive, nothing is copied (unrelated long-lived objects are already in the old gen, and because the old gen is untouched, there's no need to compact anything there).
The reasons for not using moving collectors have nothing to do with throughput:
1. Latency used to suffer. Low-latency moving collectors were a very advanced technology. The first open-source one is younger than ChatGPT.
2. Moving collectors impact the design of FFI (with C, etc.), as C (etc.) does not support moving pointers for reasons having nothing to do with performance. It's very hard for languages that want a very direct and simple FFI (as FFI is very common in code) have a hard time implementing efficient moving collectors.
3. Good moving collectors require large expert teams (let alone pauseless moving collectors). The languages that have them (to varying degrees of sophistication and performance) are well funded ones. In particular, they are the JDK team, the .NET team, and the V8 team. Good allocators (for malloc/free) are also big and sophisticated beasts these days, but they're much easier to reuse in different languages (i.e. not only by C/C++/Zig/Rust, but also by Python). Effectively, large pieces of those languages' runtime is "offshored" to unrelated specialist teams.
> The tradeoff is actually highly favorable to low-level languages, which commonly use arenas to manage challenges with heap allocation such as fragmentation.
This is also not true (and I say this because I'm primarily a low-level programmer, and have been doing low-level programming for over 25 years). First, C++ and Rust in particular make arenas hard to use to their full power (which is one of the several reasons low-level programmers prefer Zig). Second, if you're not familiar with the severe costs of memory management in low level languages, I can only conclude you haven't been doing it for very long.
Java was designed, among other things, to reduce the severe and hard-to-fix performance problems that many C++ programs had experienced (and do to this day). The hard problems concern both the limitations of AOT compilation and of free-list-based memory management. I'm not saying all C++ programs suffer from these issues, but a huge class of them do, which is one of the reasons large programs have migrated to Java.
If it was really true that "we can make the cost arbitrarily low" by simply delaying collection, that would be equivalent in the limit to not using GC at all and resorting to pure bump arena allocation. However as I mentioned, you can't really achieve this for long running programs. The chickens will always come home to roost: eventually the memory limit will be reached and you'll have to pay that collection cost in full. All that speeding up the mutation phase does is make you reach the limit even faster, compared to using heap allocation. Heap allocation may pay a small cost over time in CPU overhead (searching metadata for free memory ranges or taking locks, not doing huge memory copies - very limited impact on the memory subsystem!), but most such programs simply aren't CPU compute limited. That cost is almost free to them. The tradeoff in adding a memory copying/moving workload due to GC, even one that's invoked "rarely", is simply not worthwhile. Properly understood, they are well into the "pay more CPU compute cost in order to save on memory footprint and memory traffic" side of the tradeoff.
I'm not dismissing the theoretical possibility that such cases may exist, but there's an argument that (1) they're surprisingly rare, and (2) even when they do apply, it's inherently unlikely that adopting tracing and moving GC would be a relevant knob, compared to increased use of plain old arenas or domain-specific ways of using more memory (such as caching).
> Except it doesn't given the actual economics of RAM and CPU
These economics have been heavily impacted by the RAMpocalypse. Tradeoffs that may have been valid to some extent back when that YouTube talk was given are no more applicable today.
This sounds like you don't understand how GCs work. I can't cover all the basics, but the point is that they do constant work at a frequency that is arbitrarily low.
> The tradeoff in adding a memory copying/moving workload due to GC, even one that's invoked "rarely", is simply not worthwhile.
Again, I don't think you know the basics, so this conversation is a bit pointless. No one, and I mean no one, debates that moving collectors offer the best throughput of all general-purpose memory management solutions. There are tradeoffs, but not the ones you mention. You are debating the fundamentals of Memory Management 101.
> These economics have been heavily impacted by the RAMpocalypse.
No, they haven't. You're saying that not because that's something you've looked into, but just because it's something you imagine might be true. It could affect the economics, but for that RAM prices would have to be many, many times higher than they are now.
Of course memory safety has a quality all its own.
Whatever little CPU they waste is often worth more than the RAM they save.
> For cases where they are we've got stuff like arena allocators.
... that work by using more RAM to save on CPU.
Far less for moving collectors. That's why they're used: to reduce the overhead of malloc/free based memory management. The whole point of moving collectors is that they can make the CPU cost of memory management arbitrarily low, even lower than stack allocation. In practice it's more complicated, but the principle stands.
The reason some programs "avoid the heap like the plague" is because their memory management is CPU-inefficient (as in the case of malloc/free allocators).
> Meanwhile I'm not sure where you got this idea about the value of CPU cycles relative to RAM
There is a fundamental relationship between CPU and RAM. As we learn in basic complexity theory, the power of what can be computed depends on how much memory an algorithm can use. On the flip side, using memory and managing memory requires CPU.
To get the most basic intuition, let's look at an extreme example. Consider a machine with 1 GB of free RAM and two programs that compute the same thing and consume 100% CPU for their duration. One uses 80MB of RAM and runs for 100s; the other uses 800MB of RAM and runs for 99s (perhaps thanks to a moving collector). Which is more efficient? It may seem that we need to compare the value of 1% CPU reduction vs a 10x increase in RAM consumption, but that's not necessary. The second program is more efficient. Why? Because when a program consumes 100% of the CPU, no other program can make use of any RAM, and so both programs effectively capture all 1GB, only the second program captures it for one second less.
This scales even to cases when the CPU consumption is less than 100% CPU, as the important thing to realise is that the two resources are coupled. The thing that needs to be optimised isn't CPU and RAM separately, but the RAM/CPU ratio. A program can be less efficient by using too little RAM if using more RAM can reduce its CPU consumption to get the right ratio (e.g. by using a moving collector) and vice versa.
In the young generation, few objects survive and so few are moved (the very few that survive longer are moved into the old gen); in the old generation, most objects survive, but the allocation rate is so low that moving them is rare (although the memory management technique in the old gen doesn't matter as much precisely because the allocation rate is so low, so whether you want a moving algorithm or not in the old gen is less about speed and more about other concerns).
On top of that, the general principle of moving collectors (and why in theory they're cheaper than stack allocation) is that the cost of the overall work of moving memory is roughly constant for a specific workload, but its frequency can be made as low as you want by using more RAM.
The reason moving collectors are used in the first place is to reduce the high overhead of malloc/free allocators.
Anyway, the general point I was making above is that a machine is exhausted not when both CPU and RAM are exhausted, but when one of them is. Efficient hardware utilisation is when the program strikes some good balance between them. There's not much point to reducing RAM footprint when CPU utilisation is high or reducing CPU consumption when RAM consumption is high. Using much of one and little of the other is wasteful when you can reduce the higher one by increasing the other. Moving collectors give you a convenient knob to do that: if a program consumes a lot of CPU and little RAM, you can increase the heap and turn some RAM into CPU and vice versa.
Anyway I'm not at all inclined to blindly believe your claim that malloc/free is particularly expensive relative to various GC algorithms. At present I believe the opposite (that malloc/free is quite cheap) but I'm open to the possibility that I'm misinformed about that. You're going to need to link to reputable benchmarks if you expect me to accept the efficiency claim, but even then that wouldn't convince me that any extra CPU cycles were actually an issue for the reasons articulated in the preceding paragraph.
This doesn't matter because if you're running a single program on a machine, it might as well use all the CPU and all the RAM. As long as you're under 100% on both, you're good. But we want to utilise the hardware well because we typically want to run multiple programs (or VMs) on a single machine, and the machine is exhausted when the first of CPU or RAM is exhausted. So the question is how should your CPU and RAM usage be balanced to offer optimal utilisation given that the machine is spent when the first of CPU and RAM is spent. E.g. you can only run two programs, each using 50% of CPU; if they each use only 5% of RAM, you've saved nothing as no third program can run. So if you spend either one of these resources in an unbalanced way, you're not using your hardware optimally. Using 2% more CPU to save 200MB of RAM could be suboptimal.
I'm not saying that for every program that uses X% CPU should also use exactly X% of RAM or it must be wasting one or the other, but that's the general perspective of how to think about efficiency. Using a lot of one and little of the other is, broadly speaking, not very efficient.
> Anyway I'm not at all inclined to blindly believe your claim that malloc/free is particularly expensive relative to various GC algorithms. At present I believe the opposite (that malloc/free is quite cheap) but I'm open to the possibility that I'm misinformed about that.
You are.
> You're going to need to link to reputable benchmarks if you expect me to accept the efficiency claim, but even then that wouldn't convince me that any extra CPU cycles were actually an issue for the reasons articulated in the preceding paragraph.
I don't believe there are any reputable benchmarks of full applications (which is where memory-management matters) that are apples-to-apples. I'm speaking from over two decades of experience with C++ and Java.
The important property of moving collectors is that they give you a knob that allows you to turn RAM into CPU and vice-versa (to some extent), and that's what you want to achieve the efficient balance.
And hopefully kill Electron.
I have never seen the point of spinning up a 300+Mb app just to display something that ought to need only 500Kb to paint onto the screen.
I don’t see how design workflows matter in the conversation about cross-platform vs native and RAM efficiency since designers can always write their mockups in HTML/CSS/JS in isolation whenever they like and with any tool of their choice. You could even use purely GUI-based approaches like Figma or Sketch or any photo/vector editor, just tapping buttons and not writing a single line of web frontend code.
Yikes. I spent 15 years developing native on both mobile and desktop. If you think that native has the same design flexibility as HTML/CSS, you're objectively wrong.
By design, each operation system limits you to their particular design language, and styling of components is hidden by the API making forward-compatible customisation impossible. There's no escaping that. And if you acknowledge that fact, you can't then claim native has the same design flexibility as HTML/CSS. If you don't acknowledge that fact, you're unhinged from reality.
There's pros and cons to the two approaches, of course. But that's not what's being debated here.
They do. But not in the way that you think.
I recently switched from Spotify (well known Electron-based app) to Apple Music (well known native app). The move was mostly an ethical one, but I must say, the UI functionality and app features are basically poverty in comparison. One tiny example, navigating from playlist entry to artist requires multiple interactions. This is just one of many frustrations I've had with the app. But hey, it has beautiful liquid glass effects!
In short: iteration time matters. Times from design to implementation, to internal review, to real user feedback, and back to design from each phase should be as fast as possible. You don't get the same velocity as you do in native. Add to that you have to design and implement in quadruplicate, iOS design for iOS, Android for Android, MacOS for Mac, Windows design for windows. All that is why people use Electon.
It's bad enough having to run one boated browser, now we have to run multiples?
This is not the right path.
Anyways, I'm both cases you don't really have to write it twice.
Native to the OS: write only the UI twice, but implement the Core in Rust.
Native to the machine: Write it only once, e.g. in iced, and compile it for every Plattform.
Now that everyone who cant be bothered, vibe codes, and electron apps are the overevangelized norm… People will probably not even worry about writing js and electron will be here to stay. The only way out is to evangelize something else.
Like how half the websites have giant in your face cookie banners and half have minimalist banners. The experience will still suck for the end user because the dev doesnt care and neither do the business leaders.
About the only thing they share is curly braces.
If a js dev really wanted to it wouldn’t be a huge uphill climb to code a c app because the syntax and concepts are similar enough.
This comment makes no sense.
There ought to be a short one-liner that anyone can run to get easily installable "binaries" for their PyQt app for all major platforms. But there isn't, you have to dig up some blog post with 3 config files and a 10 argument incantation and follow it (and every blog post has a different one) when you just wanted to spend 10 minutes writing some code to solve your problem (which is how every good program gets started). So we're stuck with Electron.
and if not?
If the alternative is memory-safe and easy to build, then maybe people will switch. But until it is it's irresponsible to even try to get them to do so.
It likely would use less, and doesn't use a browser for rendering.
> And I'm pretty sure Avalonia is even worse
Definitely not
> The people who hate Electron hate JavaFX just as much if not more
In my opinion, I only see this from people that seem to form all of their opinions on tech forums and think Java=Bad. These are the people that think .NET is still windows only and post FUD because they don't know how to just ask for help.
We're not doing Electron because some popular software also using it. We're doing Electron because the ability to create truly cross-platform interfaces with the web stack is more important to us than 300 MB of user memory.
May I never have to use or work on your project's software.
It's closer to 1GB but trust me, everyone is well aware of your priorities.
Native apps are so poorly optimized that they don't offer any advantage over Electron apps.
At a cost of simplicity and beauty. And two lost decades of mediocre performance. Sigh
Then again, after many, many years of claims that the following year would be the year of the Linux Desktop, there seems to be more and more of a push into that direction. Or at least into a significant increase in market share. We can thank a current head of state for that.
[0] https://techwireasia.com/2026/04/chinese-memory-chips-ymtc-c...
>CXMT still trails Samsung, SK Hynix, and Micron by approximately three years in advanced DRAM node development, and yield rates on new production lines remain the variable that determines whether capacity targets translate into reliable supply. Liu notes that lines launched in the second half of 2026 are unlikely to change the global supply-demand balance until 2027.
The Verge article talks about demand exceeding supply in 2028. Your article suggests it'll take until 2029 before Chinese production catches up to current technology.
It'll help drive prices down in five yearss, but the Chinese memory production won't be ready and efficient enough to prevent the shortages from continuing to grow.
Assuming China takes TSMC in one piece (unlikely without internal sabotage in the best case scenario), it would still probably take years before it produces another high end GPU or CPU.
We would probably be stuck with the existing inventory of equipment for a long time…
The risk with China taking over Taiwan is that they mostly expedite their own production research by a couple of years.
Anyone trying to spin up a competitor to TSMC would have to first overcome a significant financial hurdle: the capital investment to build all the industrial equipment needed for fabrication.
Then they'd have to convince institutions to choose them over TSMC when they're unproven, and likely objectively worse than TSMC, given that they would not have its decades of experience and process optimization.
This would be mitigated somewhat if our institutions had common-sense rules in place requiring multiple vendors for every part of their supply chain—note, not just "multiple bids, leading to picking a single vendor" but "multiple vendors actively supplying them at all times". But our system prioritizes efficiency over resiliency.
A wealthy nation-state with a sufficiently motivated voter base could certainly build up a meaningful competitor to TSMC over the course of, say, a decade or two (or three...). But it would require sustained investment at all levels—and not just investment in the simple financial sense; it requires people investing their time in education and research. Dedicating their lives to making the best chips in the world. And the only reason that would work is that it defies our system, and chooses to invest in plants that won't be finished for years, and then pay for chips that they know are inferior in quality, because they're our chips, and paying for them when they're lower quality is the only way to get them to be the best chips in the world.
They have the other system.
> A wealthy nation-state with a sufficiently motivated voter base could certainly build up a meaningful competitor to TSMC over the course of, say, a decade or two (or three...).
They just need to blockade Taiwan.
Have you seen how many states and countries look enviously at Silicon Valley’s tech companies, China’s manufacturing dominance, or London’s financial sector and try to replicate them?
Turns out it’s way harder than you’d expect.
Hell, Intel can’t match TSMC despite decades of expertise, much greater fame, and regulators happy to change the law and hand out tens of billions in subsidies.
But software optimisation helps all hardware and that doesnt drive sales.
Linux however, they dont have to worry about that. Maybe it is finally the era of Haiku OS as the ghost of BeOS rises!
Basically, the optimizing that can happen is that I ditch heavy tools in favour of lighter ones, and hopefully enough other people do the same to help lighter tools with finances/dev resources.
If I look at the Activity Manager in macOS, of apps that are less trashy but currently taking up a lot of memory, they mostly aren't apps that I'm willing/able to move away from to save on resource use: Firefox, Safari, 1Password. (For browsers, you can blame poorly optimized websites for a lot of it, but I just don't see anyone rushing to create lightweight clones of websites in order to save users' RAM.)
Then, mostly by chance, I saw that my local Microcenter had some pre-builts for sale, and I ended up picking one up for <$5k that had "best in slot" components across the board, including a 5090 and even a high-end power supply.
The last time I built a gaming PC was upwards of a decade ago, and at that time the prevailing wisdom was to never buy a pre-built unless you had a massive amount of disposable income and couldn't spare even just one weekend to dedicate to a hobby project that could benefit you for years. Now, it was absolutely a no-brainer.
That's still the case, and always will be — with a pre-built you're at the very least paying for someone to assemble it for you, so it's always going to be more expensive as a baseline.
Beyond that, the chance they've chosen good components and haven't tried to screw you over on less flashy ones like the motherboard and power supply is low.
That's not to say it's literally impossible to ever find a good deal. You very well might have. Doesn't change anything though.
Except isn't it possible that pre-built companies actually get better deals on hardware bought in bulk, and therefore could offset the labor costs with cheaper materials?
Hardware pricing and availabilty pre-COVID was pretty predictable and stable, which meant the consumer could extract a meaningful cost advantage if they were willing to do the relatively modest amount of work of sourcing components individually and personally assembling the build. Right now, though, some places like Microcenter appear to have a cost advantage that fundamentally relies on market and pricing instability and can only be achieved through deeper integration with the supply chain and bulk purchasing in advance -- something a retailer like Microcenter can do, but I personally cannot.
I'm struggling to put this in context. For comparison, what was your budget for refreshing the pc you had? Were the planned upgrades going to exceed $5k at current prices? Or is the situation that a pre-build machine with far better components was now only marginally more?
Or is it that pre-built gaming PCs have stopped being a joke? I had the experience building a bicycle: I was certain I was taking the frugal path sourcing each component individually and putting it together myself. At the end I was horrified to realize I spent far more than a new bike with superior components. It was pointed out that bicycle makers are buying by the pallet and will beat diy every time — so long as they're building something I want to buy.
The lawsuits in the past prove that statement to not be basically but actually.
Demand increased, everyone built new fabs, then prices dropped and they couldn't pay off their investments. Many went out of business. It happened in the 80s, it happened in the 90s, it happened in the 2000s.
Now there's only three manufacturers left, and they know very well that demand for their product tends to be cyclical.
I've been in the industry for 30 years and I've worked at companies with fabs were demand was high and customers would only get 30% of what they ordered. Then just 2 years later our fab was only running at 50% capacity and losing money. It takes about $20 billion and 3-4 years to make a modern new fab. If you think that AI is a bubble then do you want to be left with a shiny new factory and no products to sell because demand has collapsed?
From now on, RAM will always be super costly for consumers, because they can't make massive deals like Apple/OpenAI/etc. We are the bagholders.
Have they really ever been cheap? Also Tesla 3 is cheaper now, Yaris is still cheap as well.
Now it's high again, but give it a couple years and it'll once again crash.
even if gaming is and will remain very popular for years, it and the desire to upgrade gaming rigs is still a discretionary activity with more price elasticity of demand than corporate uses for RAM in the dawn of the AI age. gamers live on the margin of this market, where low prices will stimulate upgrades and high prices will lead to holding out. The complaints about price are real, but that segment of the market is some combination of less large and less important.
Everybody’s getting pinched, not just the gamers.
letting the market set prices ensures that the chips go to the critical markets and uses. less critical uses will not allocate funds for purchases.
Can you please elaborate what you mean by "critical market"?
Edit: formatting
At the moment, nothing is certain. Could this last? Sure. Could it not last? Yup.
the current relative spike in the prices misses the medium-term trend of the vast decrease in memory price post-covid that led to the recent surge. the cartel got another opportunity to make bank and they will use that lever to the max.
funnily enough i've been personally stuck with 16 gigs since 2015, across three memory generations! but i am used to the past when you would spend 80-100 on an 8gb stick (jdec timings, nothing fancy but from a major brand) without accounting for inflation.
Another thing I've been thinking about is what happens when the next generation of NVidia chips comes out? I suspect NVidia is going to delay this to milk the current demand but at some point you'll be able to buy something that's better than the H100 or B200 or whatever the current state-of-the-art for half the price. And what's that going to do to the trillions in AI DC investment?
I'm interested when the next bump in DRAM chip density is coming. That's going to change things although it seems like much of production has moved from consumer DRAM chips to HBM chips. So maybe that won't help at all.
I do think that companies will start seeing little ot no return from billions spent on AI and that's going to be aproblem. I also think that the hudnreds of billions of capital expenditure of OpenAI is going to come crashing down as there just isn't any even theoretical future revenue that can pay for all that.
They'll just spend whatever they were planning to spend and get more performance.
I don't want to pay more because of AI companies driving the price up. That is milking.
Think I will scrap my PC and sell its parts.
I wonder if there are any niche companies building decent rigs with DDR3 and 5/6th generation Intel CPUs out there, it is cheap and might be a business opportunity?
All computers in my household are 8+ years old.
so it's 5x as expensive as Opus then.
There's a future where RAM makers tool up for this massively increased demand, then the AI companies go broke as the bubble bursts, so RAM is cheap as. So laptop manufacturers get on that and start making laptops with 1TB+ memory so we can run decent LLMs on the local machine. Everyone happy :)
We have RAM shortage now, we will have very cheap RAM tomorrow. It’s not like production is bottlenecked by raw materials. Chip companies just need to assess if the demand by AI companies will last so it’s better to scale up, or perhaps they should wait it out instead of oversupplying and cutting into their profits.
There are two RAM suppliers...
I cannot stand how you and people like you try to justify everything by supply and demand. Also you act like it's some natural law of nature. It's not a law of nature- if you took an economics class you would realize it's try to maximize PROFIT. It's not for the good of the people.
All of these things are a CHOICE that people are making to now completely screw the average person for, again, the needs of big corporations and the top 0.01%.