But in performance work, the relative speed of RAM relative to computation has dropped such that it's a common wisdom to treat today's cache as RAM of old (and today's RAM as disk of old, etc).
In software performance work it's been all about hitting the cache for a long time. LLMs aren't too amenable to caching though.
I can understand why they just decide to bake the cache algorithms into hardware, validate it and be done with it. Id love if a hardware engineer or more well-read fellow could chime in.
https://www.intel.com/content/www/us/en/developer/articles/t...
And because the abstraction is simple and easy enough to understand that when you do need close control, it's easy to achieve by just writing to the abstraction. Careful control of data layout and nontemporal instructions are almost always all you need.
Also, additional information on instructions costs instruction bandwidth and I-cache.
That is very context-dependent. In high-performance code having explicit control over caches can be very beneficial. CUDA and similar give you that ability and it is used extensively.
Now, for general "I wrote some code and want the hardware to run it fast with little effort from my side", I agree that transparent caches are the way.
I guess this is one place where it seems possible to allow for compiler annotations without disabling the default heuristics so you could maybe get the best of both.
CAR is often used in early boot before the DRAM is initialized. It works because the x86 disable cache bit actually only decouples the cache from the memory, but the CPU will still use the cache if you primed it with valid cache lines before setting the cache disable bit.
So the technique is to mark a particular range of memory as write-back cacheable, prime the cache with valid cache lines for the entire region, and then set the bit to decouple the cache from memory. Now every access to this memory region is a cache hit that doesn't write back to DRAM.
The one downside is that when CAR is on, any cache you don't allocate as memory is wasted. You could allocate only half the cache as RAM to a particular memory region, but the disable bit is global, so the other half would just sit idle.
;-)
my issue is that my company won't issue laptops with more than 16 gbs of ram
guess i'm not virtualizing anything...
Back to native apps without bloated toolkits!
Mail.app is sitting here using 137Mb of RAM. Outlook 1270Mb.
My main machine has 16Gb of RAM and I don't think I've ever seen it go over 4Gb and that was when I had a 200gb mmap'ed sparse array.
Will help reduce E-Waste, and to the end user there won't be a different. A machine from 5 years ago feels just as fast as a brand new machine.
Big Corporations offen trash IT equipment thats only 3 - 4 years old. And there is no recycling etc. Very sad.
Where are these luxurious big corporations that give their employees nice new equipment? :(
Except you can't install Windows 11 on it, and the org has to trash it anyway to keep up with security requirement (I know people on that line of work, they're all angry about it)
Also RAMsan will have a renaissance then? :-D
I don't think that ever happened. Using relatively sparse amount of memory turns into better cache management which in turn usually improves performance drastically.
And in embedded stuff being good with memory management can make the difference between 'works' and 'fail'.
This actually generalizes in a rather clean way: compared to the 1980s, you now want to cheaply compress data in memory and use succinct representations as much as practicable, since the extra compute involved in translating a more succinct representation into real data is practically free compared to even one extra cacheline fetch from RAM (which is now hundreds of cycles latency, and in parallel code often has surprisingly low throughput).
Your login isn’t slow because the developer couldn’t do leetcode
Not for fun but for convenience (laziness occasionally?). Someone needed to "pay" for the app being available on all platforms. Either the programmer by coding and optimizing multiple times, or the user by using a bloated unoptimized piece of software. The choice was made to have the user pay. It's been so long I doubt recent generations of coders could even do it differently.
Maybe a bit of engineering and planning could help here. Shipping always half finished products is it usually not a recipe to success.
They actively prefer keeping confortable margins than competing between each other. They have already been condemned for active collusion in the past.
New actors from China could shake things up a bit but the geopolitical situation makes that complicated. The market can stay broken for a long time.
Rapid increase in capacity leads to oversupply which leads to negative margins. They've been there before, and they don't want to go there again.
RAM manufacturers do routinely setup new fabs and decommision old fabs. Maybe they're trying to hurry up new fab construction in times like these, and they would likely defer shutting down old fabs or restart them where possible. But they're less likely to build new fabs that weren't already part of their long term plans.
But well, I think there is no right answer and there always be a trade off case by case depending on the context.
As 'just' a user in the 1990s and MS-DOS, fiddling with QEMM was a bit of a craft to get what you wanted to run in the memory you had.
* https://en.wikipedia.org/wiki/QEMM
(Also, DESQview was awesome.)
I do embedded Linux and ram usage is a major concern, same for other embedded applications.
I’m partying like it’s the 90s, on a 32-bit processor and a couple hundred MB of ram.
it's just a cartel cycle of gaining profits while soon eliminating all investments into competitors when flood of cheap ram "suddenly" appears
But it doesn't really need a nefarious plot for the price spikes. There is a serious lack of VRAM deployed out there. Filling that gap will take quite some time. Add to that the nefarious plot and the situation will most likely get even worse....
Something something, 2000 dot-com bubble, something
Although their stated reason for hoarding is that they "really need it", I think it was a strategic move to make their competitors' lives more difficult with little regard for the collateral consequences to non-competitors, such as regular people or companies needing new computers.
Key DRAM Factory Construction Projects:
Micron Technology (USA): Building a $100 billion, 4-fab complex in Clay, New York (first production expected around 2030) and a new $15 billion, 2-fab project in Boise, Idaho.
Micron (Global): Investing in expanding capacity in Singapore and Taiwan.
Nanya Technology (Taiwan): Previously initiated a $10.69 billion DRAM facility in New Taipei, Taiwan.
SK Hynix
The current HBM market leader is fast-tracking multiple "megafabs" and packaging centers. Cheongju, South Korea (P&T7): A new $13 billion advanced packaging and testing plant dedicated to stacking and testing HBM chips. Construction is set to begin in April 2026, with completion by late 2027.
Cheongju, South Korea (M15X): This fab is being fast-tracked for HBM4 mass production, with the first cleanroom now expected to open in February 2027.
Yongin, South Korea: SK Hynix is investing roughly $22 billion in the first fab of a massive new semiconductor cluster. Operations are planned to start in February 2027.
West Lafayette, Indiana, USA: A $3.87 billion advanced packaging site that will integrate HBM directly onto GPUs. Construction fencing was installed in February 2026, with production targeted for late 2028.
Samsung Electronics
Samsung is accelerating its "Shell First" strategy to secure production space ahead of competitors.
Pyeongtaek, South Korea (P4 & P5): Samsung has advanced the construction of the P5 cleanroom by several months, with a new operational target of late 2027. The P4 line is expected to come online even earlier, likely during 2026.
Taylor, Texas, USA: This $17 billion "megafab" is designed for advanced logic and HBM packaging. While hit by delays, it is now targeting a late 2026 opening.
Micron Technology
Micron is diversifying its HBM production across the U.S. and Asia to grow its market share.
Boise, Idaho, USA (ID1 & ID2): The ID1 fab reached a key milestone in June 2025 and is expected to start wafer output in the second half of 2027. ID2 is planned to follow shortly after.
Onondaga County, New York, USA: Micron officially broke ground in January 2026 on a $100 billion "megafab" complex, though significant supply is not expected until near 2030.
Hiroshima, Japan: A planned $9.6 billion HBM-focused fab is expected to come online between 2027 and 2028.
Singapore & Taiwan: Micron began construction on a $24 billion wafer facility in Singapore in January 2026 and acquired a fab in Taiwan for $1.8 billion to rapidly expand DRAM capacity by late 2027.
For lower end GPUs, like what goes into Apple machines.
New LPDDR Production Facilities
Samsung (Pyeongtaek P4 & P5): Samsung is converting several NAND flash lines to DRAM and accelerating the P4 and P5 fabs in South Korea. While these fabs support HBM, they are also designed for mass-producing 6th-generation 1c DRAM, which will form the basis of the next-gen LPDDR6 modules expected to debut in 2026.
SK Hynix (Icheon & M15X): SK Hynix is planning an 8-fold increase in 1c DRAM production by the end of 2026. This capacity will be split between HBM and "general-purpose" DRAM, which includes the LPDDR variants used in mobile and laptop chips.
Micron (Boise, Idaho - ID1): Micron's new ID1 fab in Boise is currently under construction, with structural steel completion reached in late 2025. It is scheduled to begin wafer output in the second half of 2027, focusing on leading-edge DRAM that includes LPDDR for the U.S. market.
The "Memory Wall" for Apple
The primary challenge is that HBM production requires significantly more wafer area than standard LPDDR. Consequently, even as these new factories open, the shortage of commodity DRAM (LPDDR5X/LPDDR6) is expected to persist through 2028 because manufacturers find HBM far more profitable.