Pure Storage Teases 300 TB Ultra-Large NVMe SSD with Tentative 2026 Launch
wccftech.com
wccftech.com
> A Pure Storage DFM contains NAND chips combined with the company's proprietary FlashArray and FlashBlade OS called Purity
Personally I have years of experience with their SANs and they're a dream to work with. But I also never had to write the checks for them.
You basically tell them how much usable space you need (high water mark), and they will make sure you have that amount available. It also comes with an SLA for latency and they'll even upgrade your controllers for free if it goes out of bounds.
I'm a huge pure storage fan.
And so that's how you get scenarios where you pay through the nose, but you've explicitly called out what you need and you get it, even if the company ends up doing weird things like shipping the top-end hardware (with speed disabled) so that they can make sure they meet the requirements.
4PB of cloud storage is probably a bit more than $800,000k and then you still have data "in the cloud" and not local to whatever you're doing with it.
Our X20 which is on the small end, with only 10 drives can easily do hundreds of thousands of IOPs, supporting over 1000 VMs and a high performance ERP solution.
The whole thing costs us about 50k a year. On an AWS/GCP that would cost a lot, lot more.
that's like number for consumer grade ssd which costs $100?
Doing so is a lot more complex and intensive than what a single consumer grade SSD can handle.
My point was just that if you wanted to get dedicated IOPs on AWS to match what you get with modern SANs, it’ll cost you far more.
I tend to think your number for PS is likely off.
Also, I am not sure how it will stack up against some cloud instance with bunch of ssd under software raid.
That is absolutely not representative of a real workload of mixed reads and writes, different block sizes, potentially different queue depths, all coming in on hundreds or maybe thousands of different volumes.
A single consumer Samsung SSD would hilariously crumble under a real workload, it will NOT deliver hundreds of thousands of IOPs in that environment.
Your benchmark screenshot is the equivalent of showing your pickup truck can do burnouts in the parking lot, and extrapolating that to think it could keep up with a Ferrari on the Nurburgring.
if you see bottleneck there, it could be that actual SSD speed is maybe irrelevant in your case since upstream software is not optimized.
Also 1GB test file is often fits in the SSD's RAM cache. Get iometer, 50R/50W%, blocks from 512 to 16k, at least half the size of the storage. Then you would see the real performance.
NB if you have random read way below the random write means you are measuring anything but the storage performance.
Sorry, lazy to learn how to use iometer, but you probably have SSD too and can report your results.
Okay, not a problem: https://imgur.com/a/teoPGrz
Real world performance[+Mix], Read&Write[+Mix]
One is Fujitsu DX200 S4 SSD SAN, other is HFM256GDJTNG-8310A.
Can you guess which is where?
>> but you probably have SSD too and can report your results.
Uh-uh!
I intended to show you the difference between a single NVMe drive and a SSD SAN (a bit old, but still very performant to handle ~900 VMs).
Sure, I can just ramp up queue depth and see some magical numbers, but for me the real performance is in everyday tasks and running CrystalMark isn't an everyday occurrence.
Did you guess which one is SAN?
I am bit confused how random 4k access is relevant to your task, it should be more like 1MB seq access likely.
My everyday task is tuning heavy data processing pipeline, and I am trying hard to achieve those q16t16.
> I intended to show you the difference between a single NVMe drive and a SSD SAN
and why you have such intention? It is obvious there is a difference.
Sorry? 900 VMs equals 100% full random access. There is no sequential access there, just as I said in my first comment.
> It is obvious there is a difference
Because of your comment[0].
This comment[1] pretty much summarized what I said in a more eloquent way.
it depends on workload, if they do most of the work in RAM, and most of fs traffic is snapshoting and restoring from snapshots, then you will get 99% io seq traffic. If they do some non-trivial fs operations, then you will get q16t16 io traffic. It is very unlikely you will get q1t1 random.
> Because of your comment[0]. > This comment[1] pretty much summarized what I said in a more eloquent way.
In my view you are jumping from topic (single ssd vs nas) to another topic (your speculations about benchmark not representing real world scenarios) and then back.
Running Linux LVM software RAID over 3 x Samsung NVMe SSDs. That's not a read-write measurement, but it's a satisfying number for a not particularly high end server.
(I use it for a side project's database engine experiments. That level of IOPS supports a very high random query rate.)
Now there’s no doubt, so they probably have more pricing power. And EMC most likely has an offering that competes with the performance (or at least as close as they can get). Back then it was flash vs spinning disks.
We do most of our stuff in the cloud, BTW, and PS is faster for cheaper, by a wide margin. But we’re not going to bring that compute back (I’d like to).
One thing I remember was them being very open and direct, and they wanted to know a lot of heuristics about our data so they could give an accurate estimate about speed and dedup. Which was really close to the reality. Oh, and we’re using one of their devices for an Itanium OpenVMS cluster, it works like a charm. Try getting a startup to support that kind of setup these days, lol
It’s a pretty good system especially for CFOs who prefer opex model
I really appreciate the effort they put into APIs for their appliance.
Not sure how they compare to other enterprise offerings these days, or how hard it is to roll your own RAID 6 open source thing with synchronous replication.
https://nimbusdata.com/products/exadrive/pricing/ # for the 50-100TB disks, low volumes but available now if you have no budget limits.
This means that the total time needed to read and/or write all the data on a given drive continues to increase. This is very important for things like backup or RAID rebuilds. Reading 50TB from a drive with 300 MB/s speed will take almost 50 hours.
It is common in my work to see read replicas of 2-3. Sometimes those replicas are split up into smaller "shards" across multiple disks even further.
Restoring might be pulling from 10 different disks at the same time.
Tell me what I can buy _today_, like with an add-to-cart button, or go away.
Unless you actually hit the maximum your building can sustain (heat, volume). Building datacenters is incredibly expensive, so reusing existing infrastructure and packing it with more is actually important.
Naturally as the storage amount increases, the price per GB decreases over time.
Consumers have been saying exactly what you have been saying since the dawn of storage. When storage capacities get higher, what you had before only becomes cheaper.
As always, Newegg link or it doesn't exist :)
$800,000 starting price (estimate, but .20 per GB is probably for the max size)
I guess, maybe if you are in video editing that would pay off? But what else is there that an individual may plausibly be doing that needs so much room?
Their computers share some resources for efficiency, like power supplies.
400-500 layers sounds exciting, that's the biggest news I get out of that.
Flash doesn't scale down with node shrinks, I guess the downside is that layer stacking will be a linear process, but hopefully that process has a longer runway than transistor node shrinks do.
performative IT like this is just saying "I'm willing to waste exceptional amounts of money because it implies I'm important".
conspicuous consumption via Rolls Royce arguably gives you actual, tangible differences, albeit not in the automotive equivalent of bandwidth and latency.
Meanwhile my 40GB X25-M runs flawlessly after 10+ years of 24/7 operation.
These newer drives probably have very advanced error correction systems that will hide defects for long enough that you get comfortable. Redundancy saves you from catastrophe, but not from the depression that follows the realization that everything is crap.
If you make the probability of those numbers you would panic too.
The most recent broke first. Brace for impact.