How much is one terabyte of data?
github.com
github.com
Aren't normal SSDs now at 500-600MB/s read?
And then there are nvme SSDs which can read up to 7GB/s
On the other points of the article, even if you had a huge disk array plugged into the machine, how many cores can you also plug into that computer? I suppose there will always be a (healthy, productive) race here between the vertical scaling of GPUs + NVMe SSDs and the horizontal scaling of CPUs and blob storage.
EDIT: formatting.
[1] First Google result is Tom's hardware: https://www.tomshardware.com/features/ssd-benchmarks-hierarc...
[2] https://cloud.google.com/compute/docs/disks/local-ssd#nvme_l...
[3] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/provisio...
[4] The ephemerality has two downsides. First, you have to get the data onto that local SSD from some other, probably slower, storage system (I haven't benchmarked GCS lately, but that's probably your best bet for quickly downloading a bunch of data?). Second, you need to use non-spot instances which are 3-6x the price.
Not for your average homelab budget but...
I find it interesting that the solution to a lot of the problems is to just reprocess the data and don't try to optimize anything. 10 TB is not a lot of data with NVMe.
Storage
| product | price (USD/GiB-month) | price (USD/IOPS-month) | claimed max read bandwidth (MiB/s) | minimum price to achieve bandwidth |
|----------------------|-----------------------|-----------------------------------|------------------------------------|------------------------------------|
| GCP NVMe Local SSD | 0.1046 [1] | 0.00 | 12,480 [2] | 13,000 USD/month [3] |
| AWS io2 SSD | 0.1250 [4] | 0.065 [4,5] | 4,000 [6] | 1,042 USD/month [4,6,14] |
| AWS io1 SSD | 0.1250 [4] | 0.065 [4,5] | 500 [7] | 135 USD/month [4,7,15] |
| Google Cloud Storage | 0.0200 [8] | 0.0004 per "1k Class B Op" [8,16] | 23,842 [9] | [10] |
| AWS S3 | 0.0230 [11] | 0.0004 per 1k GET [8,16] | 11,921? [12] | [13] |
You could build a similar table for compute but it gets complicated. FLOP seems like a reasonable
unit of compute, but there are things other than FLOPs (e.g. decoding your column-oriented
compression scheme).I've tried to do this comparison a few times but I usually find it hard to get clear aggregate FLOP numbers for GPUs. GPUs also require caretaker CPUs and I don't have experience using them so I'm not certain how to spec a VM that can practically saturate the compute of the GPUs. My gut instinct is that the big compute consumers must be able to arbitrage this to some extent by shifting some workloads to chase the cheapest FLOP.
EDIT(2x): Table formatting. We could really use some Markdown styling on HN.
EDIT3: Clarify incomparability of IOPS and GETs.
[1] https://cloud.google.com/compute/disks-image-pricing#localss...
[2] https://cloud.google.com/compute/docs/disks/local-ssd#nvme_l...
[3] For 12TiB. https://cloud.google.com/compute/docs/disks/local-ssd#nvme_l...
[4] https://aws.amazon.com/ebs/pricing/
[5] We only need need 16,000 IOPS for peak performance of io2, so I ignore the drop in price at higher volumes.
[6] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/provisio...
[7] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/provisio...
[8] https://cloud.google.com/storage/pricing#price-tables
[9] https://cloud.google.com/storage/quotas#bandwidth
[10] Honestly not sure. You can do 5,000 parallel "reads" per second to a single object. I'm not sure what kind of instance you need to receive 23,842 MiB/s or if a single object can actually deliver that much bandwidth. https://cloud.google.com/storage/quotas#objects
[11] It gets slightly cheaper as volume goes up. https://aws.amazon.com/s3/pricing/
[12] I could not find a clear answer, but it seems at least 100 Gbps (11,921 MiB/s) https://repost.aws/knowledge-center/s3-maximum-transfer-spee...
[13] As with Google [10], I'm not really sure.
[14] For 16GiB and 16k IOPS.
[15] For 40GiB and 2000 IOPS.
[16] Not really comparable to provisioned IOPS because you pay for the IOPS once per month whereas you pay for every individual GET request.
This very detail is the reason USB flash drives have useless performance numbers on them, which oftentimes lead to 1/100th of the advertised performance.
Everything that's bigger than the sector size (4k) is usually far lower than those numbers.
All the different levels of caches make the raw performance numbers (also how full is the disk) hard to find.
Source: I wrote S3's on-disk placement algorithm
If not, someone on Seth Markle's team over at Amazon wrote a pretty fascinating article on the S3 architecture. Never actually heard the "airplane at 75 MPH over a grass lawn comparison" before.
This [2] aggregate workload graph/video is pretty sweet too. Very matrixy while still actually being fairly clear about what's going on.
Imagery implies S3 (at writing) had a 3.7 PB bucket size with 2.3M req/sec. Worked out to 143 hard drives (storage constrained result) and 19,000 hard drives for the I/O constrained result. Kind of an interesting counterpoint to the article. Here's what really large workload / access rate systems with spikey, intermittent workloads are structured like.
[1] https://www.allthingsdistributed.com/2023/07/building-and-op...
[2] https://www.allthingsdistributed.com/videos/fast-agg-graph-c...
[1]: https://www.kitguru.net/components/hard-drives/simon-crisp/w...
Not to put stock in the assertion, I have no idea what the article intends
Claims 500 MB/s. But of course your chances of getting that from a hard disk are far less than from a SSD.
Probably they are pulling info from an older article or just using a round number to make it more clear.
Edit note: As a sibling comment says, Seagate had for a while a "2x" Exos model that had dual actuators and claimed ~500Mbytes/s of max transfer speed. But seem discontinued or unavailable and didn't have much success (out of the datacenter at least)
1 billion is 1000^3 (1 m cube contains a billion 1 mm cubes)
1 trillion is 10,000^3 (I can't visualize 0.1 mm, so increase size: the big cube is now 10 m on side, with 1 mm parts)
So the million and billion you could have on your desk made of parts that you can see and handle; the trillion is just a bit too big for that, you'd need a big room with a high ceiling.
Edit: and, of course, you can also see a million in 2D. A sheet of grid paper 1 m across, with 1mm grid squares. Or, more practically, try counting the pixels on a 720p monitor: 921600 pixels on that screen. Use checkerboard pattern.
Sure you can probably buy much more expensive examples, but it's not essential.
(Side note : don't try to use a gel ink pen across varnish, or it stops working until you clean it off with solvent)
> The Zettabyte Era or Zettabyte Zone[1] is a period of human and computer science history that started in the mid-2010s. The precise starting date depends on whether it is defined as when the global IP traffic first exceeded one zettabyte, which happened in 2016, or when the amount of digital data in the world first exceeded a zettabyte, which happened in 2012. A zettabyte is a multiple of the unit byte that measures digital storage, and it is equivalent to 1,000,000,000,000,000,000,000 (1021) bytes.
> According to Cisco Systems, an American multinational technology conglomerate, the global IP traffic achieved an estimated 1.2 zettabytes (an average of 96 exabytes (EB) per month) in 2016. Global IP traffic refers to all digital data that passes over an IP network which includes, but is not limited to, the public Internet. The largest contributing factor to the growth of IP traffic comes from video traffic (including online streaming services like Netflix and YouTube).
But for a city wide or even some state wide institutions, it (300rps) is really a big number.
Well, I laughed. It is common for monitoring infrastructure to poll at the minute-range, and to have dozens of probes per server. 300 is really few for a whole companyI know of an early stage YC startup that has a 6TB Postgres DB. Would it be fair to say that the DB hosting (neglecting replica, engineering time) can be done at $150/month?
Area (square feet) = 600 square miles × 5,280 square feet/square mile = 3,160,000,000 square feet
Next, we need to assume a height at which the population can comfortably stand. Let's assume an average height of 5 feet.
Volume = Area (square feet) × Height (feet) Volume = 3,160,000,000 square feet × 5 feet Volume = 15,800,000,000 cubic feet
Now, we need to convert the volume from cubic feet to cubic miles. There are 1,476,333,333 cubic feet in a cubic mile, so:
Volume (cubic miles) = Volume (cubic feet) ÷ 1,476,333,333 Volume (cubic miles) = 15,800,000,000 ÷ 1,476,333,333 Volume (cubic miles) = 10.7 cubic miles
Therefore, to contain the entire world population, we would need a volume of approximately 10.7 cubic miles, assuming an average height of 5 feet and a density similar to that of a solid object.
Please note that this calculation is purely theoretical and doesn't take into account factors like personal space, comfort, and actual population density.
Also, even taking the assumptions at face value, your model's calculations are wrong by several orders of magnitude. There are not 5,280 square feet in one square mile and 600*5280 is not 3,160,000,000.