We Test PCIe 4.0 Storage: The AnandTech 2021 SSD Benchmark Suite
anandtech.com
anandtech.com
1. Extremely Power Efficient. They have come a long way with idle power and power management. Even on the Desktop where Energy efficiency isn't much of a concern. There used to be a time where SSD are 1W+ even when idling.
2. The amount of idle time, so these IO traces shown most of the time SSD isn't doing anything at all.
3. Read Latency are now extremely good. We are talking about sub 100us and sub 200us in 99th percentile. To the point Optane doesn't provide any meaningful differences in consumer usage. We will find out soon when Billy mentioned he will be testing Optane PX5800, I cant wait to see that. I should also note before someone jump in about Optane's Random Read Write advantage in QD1, I still dont believe it matters in consumer usage beyond the current NAND SSD performances.
I think with coming PCI-E 5.0 Drive we have a roadmap that has pretty much "solved" the performance category if we haven't already done so. What I want to see is slower, but large capacity SSD that is more affordable. But apart from some unforeseeable market condition, I dont see anything on the roadmap or forecast that could bring us 4TB SSD for sub $200 within the next 4 years. We seems to be entering the end of the S Curve where improvement ( cost reduction ) will be slower.
I train deep learning models, so I read random images from disk to feed data loader pipelines, and I also write back stats, logs, and model checkpoints. I frequently use 8 GPUs, each running a separate experiment, with its own data pipeline. This means every GPU is reading 1.2M jpg images in a random manner, every 30 minutes. I haven't actually tested if the disk reads are a bottleneck, but I'm guessing random reads of 20M images per hour would stress any drive, so I'd rather not risk it and get the best money can buy.
[1] https://www.anandtech.com/show/16087/the-samsung-980-pro-pci...
The next page, working with 128KB chunks, is far more relevant to your use case. Especially this chart: https://images.anandtech.com/doci/16087/sr-s-905p-1500.png
Here's a similar benchmark, to prove that it's nothing unique to sequential reads. 256KB random chunks at queue depth 4 go well over 5 gigabytes per second. https://www.legitreviews.com/wp-content/uploads/2020/09/ATTO...
So for the use case of loading image files, with the ability to grab a bunch in parallel, flash seems like a clear winner.
It will also write large checkpoint data twice as fast.
And the extra microseconds when you're writing to a log shouldn't matter.
You do make a good point though: my reads are not 4k, they closer to 256k on average, but still, they are random reads. I’m not sure how much faster those get with size.
And I’m pretty sure doing writes in parallel with those random reads makes things worse, but ok, let’s ignore that.
Well, something to keep in mind if you want to optimize it later.
> still, they are random reads. I’m not sure how much faster those get with size.
The second image I linked should demonstrate that pretty well.
> And I’m pretty sure doing writes in parallel with those random reads makes things worse
It looks like both models handle a mixed workload acceptably. Writing 128KB chunks is somewhat slower than reading, and mixing them gives you a speed somewhere in between. http://images.anandtech.com/doci/16087/sm-s-905p-1500.png
The chart is showing how fast it is to load random chunks of data with different sizes. When you load image files, you're effectively loading random chunks of data. So if you want to figure out how fast this drive could load images for you, look at the line that best represents the size of your images.
> Would it be the same as comparing how much faster it is to transfer a 256MB file than transferring a 4MB file on that plot?
No, because the relationship is non-linear. If you want to compare 4KB and 256KB you have to look at the numbers for 4KB and 256KB. 256KB random reads are 15x faster than 4KB random reads.
The point is that when you initially linked the data for 4KB random reads, you were effectively worrying about how fast the drive would be at loading 4KB files. That number doesn't matter. You want to know how fast it is at loading 128KB or 256KB files. So look at a chart that's measuring random reads of the correct size.
If the question you want answered is the one implied by "still, they are random reads. I’m not sure how much faster those get with size.", then the answer is: That is literally the point of the ATTO chart. It shows how much faster random reads get with size. The relationship is a curve, so it's best to just look at the chart for your specific size. (And also keep in mind that this particular benchmark always had 4 IOs pending at once. And with real files, there will be a minor speed impact because they're not aligned as nicely.)
The German press seems to be more cynical about marketing, c't Magazine is great, but no longer available in English.
The Register is also a good less technical site if you want to cut through BullSh_t in IT.
It is no surprise though, as my 980 Pro comes from a long series of Samsung drive purchases including 970 EVO 2TB (in same system) and others going back, having chosen Samsung after they did seriously good on some of the earlier SSD torture tests.
(I once had 4 IBM Deathstars in raid 10...)
As it happens this is precisely Anandtech problem. They have a long legacy of printing whatever toiled paper manufacturers/vendors ship them https://news.ycombinator.com/item?id=22241800
>Intel tactic was manufacturing facts and positive press stories, something we now call fake news. ...
Intel SSE also got a big push with fake 3D acceleration claims https://www.vogons.org/viewtopic.php?f=46&t=65247&start=20#p...
Intel version: "At the time, I was working for Intel and was involved in the launch of the Pentium 3, aka Katmai.
We _engaged_ a number of games manufacturers to provide demos showcasing not only Screaming Sindy's Extensions, but the arcane and mysterious Katmai New Instructions.
One such outfit was Rage Software, now sadly deceased. Rage provided demos of Incoming and an early prototype of a game called Dispatched, which as far as I know never actually saw the light of day. Dispatched featured a strangely-arousing cat riding a jet powered motorcycle. The first version I saw was running on a 400MHz Katmai and was still in wireframe. It was bloody impressive."
Reality, according to hardware.fr: "Let's start with Dispatched first. This is actually a Rage Software game that should come out late 99, which Intel showed the demo at Comdex Fall to highlight the benefits of the SSE. Big interest, it is possible to enable or disable the use of SSE instructions at any time.
Nothing to say in terms of speed, it goes squarely faster once the SSE activated, + 50% to + 100% depending on the scenes! But looking closely at the demo, we notice - as you can see on the screenshots - that the _SSE version is less detailed_ than the non-SSE version (see the ground). Intel would you try to roll the journalists in the flour?"
SSE version is less detailed? How convenient! Rage Software Dispatched never came out. The only outfit, other than Intel, in possession of this software was Anandtech. They used this exclusive press access to pimp out Pentium 3 benchmarks manufacturing fiction like this https://images.anandtech.com/old/cpu/intel-pentium3/Image94....
>so it was pretty much an Intel commissioned demo piece to showcase P3 during Comdex Fall, and was cheating with details. Two other SSE patched games mentioned on hardware.fr actually ran slower with SSE
Anandtech used Intel commissioned piece of fake software to lie to its readers about SSE for couple of years during the time Intel was getting kicked by AMD and paying bribes to vendors.
There are differences between Reporting and Benchmarking / Reviewing.
Anandtech is perfectly happy to print Intel PR without trying to verify or make common sense assertions.
https://www.asus.com/ca-en/Motherboard-Accessories/HYPER-M-2...
It costs about $70 USD.
In a single socket ryzen zen3 workstation you'd typically be able to use one of these. One PCI-E 4.0 x16 slot for your video card, one slot for this.
Realistically you need to give this card your main slot and your GPU can get an x4 connection and lose a few fps.
*: I know there are other reasons for non-K processors (binning might actually mean your non-K CPU can't overclock), but it's largely just a way to move more product and sell at a lower price point to those who can't afford unlocked or otherwise wouldn't have overclocked.
(Binning because of imperfect production yields is different though and I don't have a problem with that. It's hard to tell how much of the K/non-K split is legitimate binning, and how much is just a bool in the microcode).
https://www.amd.com/en/technologies/smart-access-memory
Or that time x470 was going to support PCIE 4, but then it was made x570 exclusive
I'd like to see a complete fin heatsink that benefits from case flow.
Not sure if this thing is smart enough to know when to kick the fans on and when to keep the temps "high". That is something worth noting though.
https://www.extremetech.com/computing/142096-self-healing-se...
https://hardware.slashdot.org/story/12/12/02/2222235/self-he...
I found this here that says lower temperatures are better for storage but higher temperatures are better for writes. I'm not sure how accurate that all is though
https://en.wikipedia.org/wiki/Field_electron_emission#Fowler...
Go to "Writing and erasing". It links to: https://en.wikipedia.org/wiki/Tunnel_injection?wprov=sfti1
For PCI-E 3.0 i can recommend a card with a Broadcom PEX 8724 controller. Works well with 4 drives and is not much more expensive. The fan is horrible: https://www.aliexpress.com/item/1005001782779648.html
Looking forward to seeing what you pick in terms of application benchmarks. I think application benchmarks help to keep us honest in terms of real-world impact of storage upgrades. LTT seeing if people could tell the difference between SATA and NVMe for games was great. I'm interested in database workloads, and the reason I was asking about the SN850 is that in Postgres benchmarks I've seen it has smoked the 980 Pro.
Could someone fill me in with PCIe. What slots are available and are these mainly used now for graphics and SSD? SSD in slots is new to me which I haven't gotten around to yet.
Typically, there are 16 lanes used per GPU and 4 per M.2 SSD.
With older CPUs, you had to either reduce the lanes assigned to the GPU to connect an SSD to the CPU, or go through the chipset.
With newer CPUs (newest intel or AMD Zen), the CPU provides 20 Lanes (16+4) for general use and in the case of AMD, 4 additional lanes for the chipset.
If you want more lanes connected directly to the CPU, you have to go to HEDT or Server Chips with up to 128 lanes available.
See this article for a typical block diagram of a AM4 system.
https://www.guru3d.com/news-story/amd-ryzen-3000-new-block-d...
Here you see, there are 1x16 / 2x8 lanes for GPU, and 4x PCIe/SATA combo ports from the CPU which can be used for connectivity.
More lanes can be provided by the PCH (Chipset).
Another example is in this artice by anandtech, which even lists the slots explicitly in the block diagram. https://www.anandtech.com/show/14657/the-asus-pro-ws-x570ace...
That is amazing. That is around DDR4-1866 speeds, and not far from DDR4-2666 (~21 GB/s). At those speeds I would happily work with dataframes sitting on the disk rather than in memory [1, 2]. Did you benchmark RAID 0 with less than four disks?
[1] R: https://github.com/xiaodaigh/disk.frame
[2] Python: https://docs.dask.org/en/latest/dataframe.html
3990X, 265GB DDR4 3200Mhz, Asus Rog Strix E-Gaming
I'm blown away, so what are you all using the even greater throughput for?
The utility is quite slim as is, because of the networking limitations and other I/O. Only certain internal processing and parsing can take advantage of this.
See the first half of https://www.youtube.com/watch?v=fqi09JnJHOo It's partly a comparison against hard drives, but it gets into why you'd want gigabytes per second.
They are the only PCIe 3.0 contender, yet beat Samsung an almost everything, but raw latency, and throughput.
They are indeed limited just by the PCIe 3.0 ceiling.
tldr: if I'm buying a new machine doing a lot of GPU work, do I get a PCIe 4.0 storage system or 3.0?
Both Sony and Microsoft are talking big about the storage subsystems of their new consoles, and how it's going to enable entirely new things, such as low-latency direct requests to the flash from the GPU. for gaming in the future. This is likely to spill into the PC ecosystem during this console generation, both for games and gpu compute.
Personally, I would not purchase a motherboard that didn't support PCIe 4.0 anymore, but I would not worry about getting the fastest possible drive to plug into it. The idea behind this being that I am almost certainly going to expand/replace the drive before it matters, but I'm quite likely to still be using the motherboard at that time.
Hynix P31 seems to be good choice overall if you are looking for an NVMe drive (speed, energy efficiency, 5 yr warranty)