AMD Demonstrates Stacked 3D V-Cache Technology: 192 MB at 2 TB/SEC
anandtech.com
anandtech.com
DRAM is incompatible with the logic family used in CPU making, though attempts were numerous at trying to work around that. Every few years, there are somebody coming with claims of passable CMOS+DRAM tech, but none got adoption so far.
A very curious case, and a very clever hackery with a custom SOI process. It's still very, very far from coming to mainstream chips.
"Both it and all levels of cache in the main processor from level 1 use eDRAM, instead of the traditionally used SRAM."
So if you just first make CMOS, and leave empty space nearby protected by something to do DRAM later, it will be very hard to fit into the tiny remaining thermal budget past which CMOS devices will turn into schmoo.
[a]: what’s the upper limit? As in, when does it not work anymore? For example, would it work at 100+ nm?
First, you need a wafer on which you can make a capacitor, which already means an SOI process, and SOI wafers.
The lowest nodes with SOI available to mortals are 40nm, and 14nm on GlobalFoundries, but god knows how one gets GloFo to collaborate on that.
Then, you need CMOS device which has enough thermal budget to survive both own, and DRAM creation.
Third, I think on-die DRAM will only make sense when it wins over SRAM. DRAM cells cannot physically go smaller than the minimal size of a trench capacitor. I believe at 5-7nm nodes, 6T SRAM will already be smaller than a reasonably fast eDRAM per area.
We know that IBM's Z15 is a 14nm FinFet chip, and it has DRAM on board, with them probably somehow doing DRAM first.
As you said, I'm not sure about the density part. Unless it's completely done in BEOL, I don't see designers trade precious chip real estate for memory (unless pin or power-limited, of course).
What do you think on FRAM vs. next generation MRAM?
Well, hafnium can be used to make ferroelectric crystals, that are necessary for ferroelectric memories (FRAM). The most used ferroelectric material is PZT, which has lead, and is a nightmare for CMOS compatibility due to contamination issues, and temperatures.
It's a bit light on the details, but maybe I should point you towards the video that explains our project: https://www.youtube.com/watch?v=M8tL-nN7G-A
The most exciting part might not be the performance (though it looks good), but the way ferroelectrics can be used for new circuits (variable treshold transistors thanks to FeFETs). Hafnium has been used in gate oxides for a few years now so it's quite compatible with CMOS.
No word on latency though.
Depending on what you mean by "memory-on-chip" we've already seen this with AMD's Fiji (2015) & Vega (2017) line of GPUs which used on-package memory in the form of HBM/HBM2 memory ( https://www.anandtech.com/show/9390/the-amd-radeon-r9-fury-x... and https://www.anandtech.com/show/11717/the-amd-radeon-rx-vega-... ).
It was incredibly fast, but also very expensive and limiting in SKU configurations. The resulting 16GB SKU for the Vega 56/64, for example, made basically nobody happy. It was too much for gamers, who then didn't want to pay money for something that didn't help, and it wasn't enough for the professional crowd, who were getting used to 24GB offerings from Nvidia.
> Actually isn't that already the case for mobile CPU's?
Nope. They "just" replace the slot with solder more or less. It's still externally packaged DRAM modules.
I see AMD's effort as a way to sidestep the whole problem of putting memory and CPU on the same chip, by making separate chips and stacking them. It makes sense. You can get CPU packages with CPU + DRAM + Flash, but these are separate chips which are wired together inside the package.
E.g.
https://en.wikipedia.org/wiki/Package_on_a_package
The whole point of these systems is to avoid putting memory and CPU (and flash) on the same chip. AMD's version is smaller & more integrated than package-on-package but it still achieves the same goal: multiple chips.
The thing here is balancing core speed vs “effective” memory latency (factoring in cache hits, misses, etc).
Cache management is a hard problem, and the equilibrium point is load dependent (i.e. depends on what type of program you use).
AMD has been smart enough to understand that sometimes it’s better to just brute force your way in (with higher cache sizes) than being super clever about how you handle it.
They add massive caches as a side effect of having loads of cores enabled by their chipset designs.
I generally err on the side of “I’m not as smart as I think I am” in these kinds of discussions. It’s not that I can never do it. I’m sure if I studied it a bit I could. The reason is that 99% of the lines of code I’ve ever written don’t warrant that kind of attention. I’m sure there are uses like when you’re writing some core fundamental algorithm in the fields of compression, crypto algorithm, hash, video processing, etc etc etc. I don’t work in those spaces though so the benefit is much more marginal.
It’s possible 192mb is enough to start needing some explicit memory management. It makes the coding model much more complex though. And in a greedy software system where your code isn’t the only one running, such complexity doesn’t necessarily net overall wins. It’s the reason we have drivers and OSes even though we started with each piece of content bundling explicit HW support (at much better perf generally)
1. Heat doesn't go away with layers and gets harder to cool. 2. These are still layers. You want each layer to be as flat as possible. As the layers deform you lose the ability to cleanly add more layers.
Current lithography is very much based on the notion that everything is flat.
You also don't want to run your CPU at 80°C, else the RAM will produce much more errors. (ECC is a must anyway, I suppose.)
Now what we have over the years seen interest in is adding processing cores to the RAM itself and maybe having a small dedicated processor for some tasks attached to the RAM may well prove viable.
Imagine if we didn't have any CPU or RAM sockets/slots and just a row of slots you addedd a module that had CPU/RAM in one and can just add more upto the slot limit. But then, that is kinda how GPU's have gone already in many aspects and look at how much RAM they hold and how large the cooling solution for them is. That gives you an idea of the cooling needed for large amounts of processing and RAM when closely packaged.
I think we call these 'GPUs' today.
> just a row of slots you addedd a module that had CPU/RAM
As you mentioned, GPUs fit this description, but some PCIe devices are full-blown embedded systems.
I think we're at the point where the only major improvements will take place on the CPU die itself.
Whereas, what the parent poster is describing would be true MIMD: a bunch of tiny cores each with its own on-board RAM, its own instruction pointer, and then a bus (or a bunch of busses) fast enough to feed them all data from (probably NUMA) main memory.
GPUs don't provide any advantage for running e.g. a high-concurrency Erlang application server. But a true-MIMD system would.
Bryan Cantrill talked extensively about how long it took for Sun Microsystems to open source Solaris, and a lot came from code that was often outsources because it was "boring" and basically non-core tech (the example was i18n/l10n)
AMD is something like 10% of the total Linux kernel size now which is ridiculous. But their commitment in recent years is admirable as hell
NVIDIA need to pull their fucking head out and just start working on Nouveau
The driver code itself also has lots of code duplication. It's just strange.
I've had enough issues with ideologues precenting any meaning full progress
- it is a driver, not a core module
- the constants are implementation details of the driver
- active maintenance of code is a necessary condition for inclusion in the Kernel. The other Kernel developers are not supposed to maintain and refactor code dumps.
There was no hardware at all on which you could rely on it, on top of it being much worse than CUDA.
Me memory isn't perfect, but IIRC the situation was roughly the following: We were quite short on resources (both devtime and money), which meant that we had to choose our scope wisely. Optimally we would have implemented both CUDA and OpenCL 2.0, but we had to settle for OpenCL 1.2 (which offered reduced performance, but was "good enough" for inference). IIRC OpenCL 2.0 was very very similar in what capabilities it assumed and offered to the CUDA version at the time, and cards like the GTX Titan X had "compute capabilities" that supported features like shared virtual memory between CPU and GPU in CUDA at the time. In fact the advances around memory management (and async copying) that were present in CUDA and not in OpenCL 1.x were the main source for the performance differences between the two.
From everything that I can tell at that point in time, if NVIDIA would have wanted to support OpenCL 2.0 they could have done so based on technical requirements. What the reason for not doing so is, is just pure speculation (lack of internal resources due to focusing on devtools?), but to me it always looked like they were using the edge they got via their proprietary libraries like cuDNN to get a foot into the field of ML and then purposefully neglected OpenCL to prevent any competitors from catching up. Classic Embrace, Extend, Extinguish.
Then again, it sounds like OpenCL 2.0 required some flexibility the nvidia drivers or hardware wasn't able to provide.
It's pretty hard to speculate to which degree nvidia was intentionally sandbagging here, and to which degree it really was stuck.
However, it's a member of khronos, and it's hard to believe they as such a major manufacturer could not have either said beforehand that the spec was a problem or simply complied with it; as both AMD and intel did.
Also, given CUDA's success it's all rather convenient for nVidia - at the very least it looks like they didn't mind leveraging their market position for continued market dominance, even if it's unclear whether that was an intentionally anti-competitive aim from the get go, or simply a fortunate happenstance they didn't try to avoid.
Then again; with antitrust enforcement mostly remaining in vaporware mode it's a little hard to blame them.
Intel's stack is already pretty open on Linux.
Now they got a solution to add hundreds of megs of cache for cheap.
More importantly SRAM can be binned/tested/KGDed separately!
And it can be fabbed on its own customised process!
And also, they can throw in an eDRAM there instead on a moment notice.
And as they already have silicon interposer here, adding HBM2/3 will be also a triffle.
They can seriously reduce the L3 cache area, which is like 50% of the die currently, or get rid of it alltogether.
Imagine, 2 times more dies per wafer at near no cost except for extra packaging.
If a particle kills a repairable part of the die (which can be downgraded, or cut-off) you still get from a quarter, to a half of die area dead, and a second grade chip.
But if you can make 2x dies from wafer even with the same defect rate (it will usually be less.) Even if the die will completely die from a local defect, you still get much more perfect dies in total, which is more important for your bottom line.
Standalone SRAM can be very repairable, and high yield with a custom process. Adding few spare columns, or SRAM banks should be covering for much more defects than if binning was done per entire CPU die.
Plus there's an upper limit to the computing power that people need from their phones as well, so the sleek ads for phones simply don't work on many people nowadays. They are content with what they have.
See Intel's Foveros https://www.anandtech.com/tag/foveros or Ponte Vecchio that they've been talking about for 2 years now ( https://www.anandtech.com/show/16453/intel-teases-ponte-vecc... )
A quick googling brings up this article, which I think may be the same thing: https://www.bbc.com/news/science-environment-24571219
Both litography and the desperate search for infinite growth is going to hurt in the long run!
I'm stacking up on 14nm Atom that I can passively cool at 42 degrees celcius full blast on 8 cores and 50nm (100.000 writes per bit) X25-E SSDs!
[1] https://www.newegg.com/supermicro-mbd-a1sri-2758f-o-intel-at...
I think these boards are the first and last to be able to run 100 years! Earlier models consume too much energy and subsequent will be too fragile/complex!
Don't be so sure. If it was going to be the case, we'd be changing CPUs like early Seagate disks at our datacenter. There's no measurable longevity loss in the CPUs that we observed in our data center.
We use every system for ~7 years at full load.
Since you haven't tried the 7nm CPUs yet I guess we'll have to see?
7 years sounds really short to me, I'm planning on running my machines for 100 years!
That CPU's reported temperatures are 81 for high, 91 for critical (in degrees Celsius).
We get new systems almost every year, so we have a rolling set consisting many systems. So, we have a cross section of systems to observe.
Temperature kills hardware, you need to bring those temperatures down, and the only way to do that is to lower the wattage!
Atom is the perfect design, no crap and low power!
From my experience, that doesn't happen like that in enterprise hardware. Either wrong voltage (inside the system) or defective design causes premature death. If the server BMC says it's fine, it's fine.
> Temperature kills hardware, you need to bring those temperatures down, and the only way to do that is to lower the wattage!
High Performance Computing doesn't work like that, unfortunately :)
> Atom is the perfect design, no crap and low power!
I'm sure it has its own uses and can accomplish a lot, but in HPC, it won't cut it. I use small SBCs at home to do and try a lot of fun and useful stuff, but it has limits.
The best way to keep your CPU intact for a long time is keep it powered off. But then you have probably no use for this CPU…
> it's too hot and it's going to break very soon!
“very soon” on your scale from now to 100 years, probably. But most people prefer having a CPU working at full capacity for a few years than having a useless brick of silicon sitting around for a hundred year.
Nonetheless, there is a continuous aging of the metal traces (electromigration) and especially of the insulating layers, e.g. the MOS transistor gates, due to the difusion of atoms.
This continuous aging is accelerated by steady-state high temperatures and it eventually results in either open circuits or short circuits somewhere, destroying the device.
Good MOS integrated circuits are designed for a lifetime at their maximum specified temperature of at least 10 years or even 20 years or more for the better of them.
Nevertheless, this is puny in comparison with the lifetime of many semiconductor components produced 40-50 years ago, before the continuous shrinking of the active device sizes, which could have lifetimes of hundreds of years, when free from fabrication defects.
I even think like you that my 45nm D510MO probably will outlive these! But it's so underpowered (like 1/10 of the perf. at 15W) compared to these that I'm willing to take the risk!
The risk of new boards being so much better that I'll have to throw these away ever is zero at this point, memory being the bottleneck!
That's only like ~45 °C or so.
Are you talking about junction temperature, package temp, or heatsink temp? The only CPU I own with a package temp that's cold enough to touch is inside my phone.
Meanwhile some of my data center machines are a decade old and run at 75C all day. I've never had a CPU fail before the machine became obsolete
The biggest offender in my experience is OrangePi Zero, but I need a thermal probe to verify it.
That's 50C (bit less)... which is quite low operating temperature and very far from anything that damages silicon. Temperature fluctuations are far worse due to thermal expansion/contraction.
There are plenty of things that damage chips, especially those at 7nm and beyond but this too-hot-to-touch "rule" is rubbish.
So, per die sensors are probably reading somewhat lower numbers, but the core's cooking at 70 degrees C internally. Nevertheless, bigger die surface inevitably allows better heat conduction and reduces internal stresses considerably, when compared to a desktop part.
Nevertheless the numbers you mention are neither unrealistic, nor impossible in stock cooling and/or sustained load scenarios.
I really want to buy a proper serious workstation and I can afford it in a few months but I keep wondering: is now the best time for it?
There's never a good time to buy a new computer. It's always going to be obsolete before you open the box. Though in this case, one might wait for the Zen 3 Threadrippers.
--
As for the Zen 3 TRs, well, to be fair, I am not looking to blow $50k on a workstation. :) I am more interested how -- and if -- they will drive down the prices of the TR 3900 SKUs...
You won't spend $50,000 on a workstation just by using current-generation parts. I think when you see a workstation that costs that much, it's because it has multiple GPUs in them. Pro GPUs are always artificially overpriced, and given the GPU shortage, they're now even more overpriced.
I did a quick pcpartpicker expedition and found that using last-gen parts saves you about $1000 on a $6000 32 core Threadripper workstation. I compared last-gen SSDs, consumer GPUs, processor, and motherboard, and picked relatively high-end parts. You will save more money by dropping to 24 (or heaven forbid, 16) cores, not getting an extreme motherboard, getting 64G of RAM instead of 128, etc.
This could all be invalid in a few months. It is hard to separate "market is always crazy" from "this whole COVID thing is going on". Building a workstation during the pandemic was a pain -- I bought a used GPU, and didn't get ECC memory because nobody would sell me any. If you wait a year, that is likely to improve, and there will be newer hardware. But, if you need to do some computing between now and then, you don't have much choice but to buy what's available now, and it certainly makes for a very good computer.
I realized I was looking at several custom-made TR Pro workstations and they of course have a hefty markup on top -- custom cases, custom cooling, pre-added several PCIe NVMe riser cards, plus style price tag I suppose, etc. Should have looked in PC Part Picker indeed.
And yeah, last-gen tech very rarely gets discounted even by normal people who just post ads in local Craigslist-like websites (like OLX). Stuff that's 2 or more generations ago is discounted, but not last gen. Puzzling indeed, especially having in mind that this last gen tech is very soon going to be "two gens ago". Oh well.
As for ECC RAM, I hear you. I was unable to find such anywhere officially but lucked out that one guy in the local OLX was having loads of it but just didn't post the ads due to being very busy (we communicated because of other ads) and then I just bought 64GB ECC DDR3 RAM from him for a home NAS that I am gradually expanding.
I'm definitely very interested in having a TR Pro workstation; started getting sick of Macs and their artificial slowness. I got the iMac Pro and granted, I have it connected to my TV where it plays Twitch streams all day but hell, a lot of stuff on the terminal (that's not Python) just works slower than it does on a meager i3 Linux machine that I have lying around.
So I do want a TR Pro machine but I am very curious about TR 5000. A release is expected in August which isn't that far away. On the other hand, the TR Pro 5000 might take several more months on top. Hmmm. Decisions, decisions. :)
There is something to be said for the prebuilt workstations from reputable vendors. It shows up in a box, and starts working. I have strongly considered that angle; being a system integrator is tough. If something doesn't work, it's a week cycle time where you find a new part, organize the RMA of the faulty part, etc. and there is a potentially unbounded amount of time spent tweaking. Meanwhile, HP or Lenovo just ships you a computer; they tested it and it works. You pay several thousand dollars for the privilege, but it might be worth it. (And if you want a TR Pro, you have no choice. AMD doesn't sell them to consumers.)
If so, why? I can't imagine a workload that is so unchangeable in its nature, especially considering the world we live in and how often storage media changes that I'd want to plan ahead for 100 years.
My drives are from 2011, peak SSD at 50nm (100.000 writes per bit).
My CPU/motherboards are from 2017, I think 2Gflops/watt is about where peak CPU is at.
Time will tell... but there is no point waiting for improvements now, time to go all in!
Heh, that's an interesting thought that i've also sometimes had.
For doing something like that, the hardware would need to be way more resilient and have failover for various components - for example RAID for the HDDs/SSDs and some mechanism for clustering and failover for apps and other stuff you'd want to run on them. But even with those in place, i'm not sure that the hardware that's available to us wouldn't just die long before the 100 year mark.
Anyone have any idea what are the oldest computers presently in continuous use? Best i could find was this, but it was just turned on for once, instead of being continuously working: https://www.smithsonianmag.com/smart-news/watch-the-worlds-o... Apart from that, all i can think of are mainframes and such.
I really doubt that any piece of currently modern hardware could last 100 years without becoming some hard to understand mess that's incredibly out of touch with the OSes and paradigms of the future. What would you even run on it? Debian? FreeBSD? Haiku? Would there be anyone to debug Python 2 or Python 3 in 100 years? What about Java, .NET, PHP, Golang, Rust, C++ or even C? Considering that many of those ecosystems are more and more migrating to integration with the Internet, especially for the dependencies, what are the chances that any of the software will survive that long?
Voyager 1 and 2 are good candidates. In the end, they'll probably shut off because their generators won't provide enough power to operate their computers after 50 years or so.
Instead I made my own async. distributed database: http://root.rupy.se
That way I can just fix the drives as they fail.
As for OS/languages I'm 100% sure linux + Java is the final server platform for eternity.
Nothing has improved for 5+ years since NIO got epoll!
The io_uring and potentially user network/disk stuff to work around the kernel waste (in my case ~30% of the CPU) might happen but also might not...
Nevertheless, their expected lifetime is not much greater than that.
I normally replace CPUs because they become obsolete much earlier than they should show age problems.
Nevertheless, I had a pair of Opterons that were used continuously for about 10 years.
At the end, they developed a very high leakage current, leading to an idle power consumption several times higher than in the beginning.
Increased leakage current is one of the most frequent signs of aging in semiconductor integrated circuits.
Querying and recording these as time series provides very nice insights over time.
Copper 400 W / m K Aluminium 230 W / m K Silicon 150 W / m K Iron/Steel 50 W / m K
[1] Though Intel thinned their dies recently to improve thermal performance (CPUs are flip-chip, so the metal layers and active circuitry are facing the interposer). Some were concerned about stability/cracking of the thinner dies. Perhaps AMD is doing the same here, thinning the compute die, then stacking the memory die on top to end up with a stack that's exactly the same thickness as before. Since they're bonded together, the structural integrity should be similar. That additionally has the advantage that you can keep using the exact same IHS as before.
> - The processor with V-Cache is the same z-height as current Zen 3 products - both the core chiplet and the V-Cache are thinned to have an equal z-height as the IOD die for seamless integration
> - As the V-Cache is built over the L3 cache on the main CCX, it doesn't sit over any of the hotspots created by the cores and so thermal considerations are less of an issue. The support silicon above the cores is designed to be thermally efficient.
There's additional silicon stiffeners which should help with thermal transfer, granted at lower thermal efficiency than a single element.
-----
Cooling? Well, I'm using a 5900X right now, air cooled, with a high airflow case loaded with slow spinning 140mm case fans.
In games at 1440p120 (maxed settings, with RTX) I'm hotspotting at -- wait for it -- about 68C. With most of the CPU at 60C. In CPU intensive applications, more like 74C hotspot, with most of the chip at 65-67C. That's a setup that's still inaudible to me at 1m distance.
I feel this is going to be a 5900XT and 5950XT. Fills a price gap between the higher end X desktop CPUs and Threadripper for the HEDT market. Great for reasonably priced dev desktops (without falling down the workstation rabbit hole), as compilers love cache.
... though next year with 64-core EPYCs at 5nm with 768MB of L3? Oh. Dear. What's in the Xeon pipeline that can attempt to compete with A CPU that dominates on PPW and will be neither cache-starved nor core-starved? I guess it'll fuel a lot of Optane 5800-series sales, as driving down IO latency to sub-10μs will matter more.
This could also mean a GPU chiplet on package with compute. Each chiplet gets at least one cache layer. The next few years could be pretty crazy.
The power cost will really determine if we see this in Laptops or not though.
You will sometimes, but traffic that still overflows to the other chiplet is probably almost all traffic that would have gone to RAM before.
And as a reference point, cross-chiplet L3 is about 50 nanoseconds slower than local L3. That will dwarf most things.
2. You will still ( for now ) need L3 Cache on the Chiplet where you have layers of SRAM on top.
A solid brick of compute would have to run pretty slow, square chips have the surface area for heat removal.
That's not necessarily a problem if the cost of fabbing logic drops enough. It actually could make a lot of sense: throw more and more transistors and solving problems higher up the stack but use them with a lower duty cycle to avoid thermal issues. We may end up living in a world build of logic bricks, gently throwing off heat and solved problems while hiding as structural or aesthetic components.
If we’re talking science function anyway, could it be a superconductor at room temperature? Or would it be impossible to build logic out of that?
https://en.wikipedia.org/wiki/Reversible_computing#Reversibi...
I wonder how the continuous growth of caches in the recent past, as well as this approach will affect this performance requirement.
these big caches will help indirect code be a few % less bad but it will still be better by a lot to keep things in l1.
now maybe one day with some unforeseen developments, linked lists will be cool again.
I think latency has a lot to do with the speed of light - since both of these CPUs run at similar frequencies, you can cover more SRAM cells worth of distance in the one with a denser process geometry, hence you can have access to a larger cache with the same latency.
The less capacitance (and 3d finfets / other innovations reduce capacitance), the fewer electrons it will take to "turn on" or "turn off" a transistor.
That's why modern process designers are trying more and more exotic shapes to reduce capacitance and reduce the number of electrons that need to be moved anywhere for the "on-off" toggle to happen.
Leakage current is the amount of electrons that "leak" even in the offstate (which all CMOS designs are statically in the off-state. Even at 5V / on, there's a complementary transistor that's in the off-state somewhere that theoretically prevents current from flowing).
As such, the two sources of current (and therefore power usage) are capacitance (the number of electrons you need to move to turn a transistor from on to off, or vice versa). Or leakage (the number of electrons that flow despite the transistor being nominally off).
---------
The lower the leakage, the less power is used. The lower the capacitance, the less power is used AND the quicker the transistor turns on / off (because fewer electrons have to move, so the switch goes from .25 nanoseconds to .1 nanoseconds or whatever).
And regarding "Java will never be as fast as well-written C++ code." => since Java is getting inline/value/data classes [1] is will also benefit from data locality.
The CPU word size, number of registers, memory sizes and cache all matter to this since they determine the "base unit size" that the hotspot code could address without dropping down a layer.
Pointer-chasing code can be slow, but if the entire structure fits in cache, it won't be that slow. It gets slow mostly because the CPU can't predict what's needed next and keep the pipeline fed. And if you were to have a pointer structure that also existed in memory mostly-linearly, it would most likely have performance characteristics similar to a flat array. At scale, the data format really is the bottleneck to performance - shaving off a few bytes with better packing can make the difference, and in this respect, yes, Java is going to impose some limits via lack of control.
That said, as a society we still process petabytes of JSON daily. We aren't constantly doing high-performance things - inputting and editing data and making it legible to humans are tasks that have the natural result of putting some "slack" and redundancy into data. Games tend to get results fast by "cheating" and turning the full-fat editable data into a lean, single-purpose rendering structure.
It's a common feature on embedded processors, where portability is less of a concern.
The introduction talk to this is by Mike Acton at cppcon, but Jon Blow has been harping on about it as well
E.g. I want to run Quake1 right from the L3 cache.
E.g. with filesystem I'm sure when I read or write it, if I use regular IO operations, no mmap-ing (of course there's a lot of quirks with drive caches not respecting fsync, etc, but let's skip that for sake of this comment).
Would a 512MB of Cache, or say a 4GB Cache dramatically change the performance of these Dynamic Languages? It still wouldn't be C like performance, but instead of the current 10-20x for Python or Ruby, it would be closer to 3-5x?
Or would it not be much difference at all? I could imagine in the future there will be a System Level Cache built on top the IOD on EPYC where it has 4-8GB capacity.
That was HMC (a competitor to HBM), a DRAM technology that IIRC was slightly higher latency than the other DRAM on their board. So it was kind of hard to use. HMC had more bandwidth but worse latency, so it was only sometimes faster.
This V-Cache is SRAM, and therefore likely to be lower-latency than any DRAM technology, and therefore actually useful as a cache.
I don't have experience with Xeon Phi. I just remember people talking about it back then.
Can you do many writes to the same byte?