S3 is an absolutely terrible financial choice for systems that need to store a vast number of tiny files.
S3 is an absolutely terrible financial choice for systems that need to store a vast number of tiny files.
A key-value store would probably work well, depending on how well its storage layer is architected.
There's a big caveat to any NoSQL database and that's how you handle aggregates/roll-ups. With a standard database it's easy to write these queries. If you do it without thinking on a NoSQL system it'll cost you in performance and where billed per access, money. There's a few ways to address this;
- batch ala map-reduce.
- streaming ala Apache Beam, Spark, etc.
- in query counting (aka sharded counters).
- Up to 1M writes per minute, or ~16k/second. Obviously they don't operate at full capacity 24x7 but let's assume they do for worst-case planning. That means they need 16k WCU, or $7500/mo.
- Their RCU would obviously be significantly lower, let's say just $500/mo for analysis.
- Storage is going to cost $0.25 per GB-month. Let's say a very small payload of 100 bytes per object + 100 bytes for Indexed overhead. 200 bytes * 1M/minute * 1440 minutes/day = +288 GB/day. By the end of month #1, you're looking at $2,160/mo. By year #1 (105TB) you're spending $26k/mo.
At the end of the day, they're going to need to massively scale down their write throughput and constantly perform rollups/aggregations or they'll go broke if they ever want to simply store the raw/original data using S3 or DynamoDB.
Cost analyses like these make me very hesitant to use S3 or DynamoDB for an operation like in this article. Sure, individual operations are a fraction of a penny, but it grows very quickly.
It seems like it'd be far more sensible to just get a few EC2 instances and run something (cassanda, sqlite, spark, hadoop) locally.
1. desired time to market.
2. operating cost.
3. people cost.
If you're dealing at that level of traffic you're not going to go broke unless you lack a decent business model. One of the companies I worked at was spending $100,000+ per month on AWS (more than they should've for the incoming traffic level but they survived for a number of years). A friend was working at another company that was spending $3,000,000+/month. The motivation for using DynamoDB/Datastore over say S3 is that you can do range queries, field sub-selection, offloading of systems administration, some elements of security, etc.
Most traffic patterns have windows and tend to be diurnal in nature. This depends on the number of timezones your service covers. Obvious exception being IoT type time-series. So you might experience peaks of 16k/s but that will be for perhaps a few windows in the day. I also wouldn't be treating DynamoDB as permanent "long-term" storage for granular data. I would be using a sliding window of whatever I consider to be a valuable retention period (e.g. sliding window of 30 days). If you don't you'd need to scale C* as well.
In order to run Cassandra and get streaming aggregation via Spark you need to run "2 regions". One as the primary online region for serving requests and a second region for stream processing using Spark. Minimum recommended memory requirement is 16GB for a C* node. For a C* + Spark node minimum is 32GB. The more the better for both. If you've selected your keys well C* will scale almost linearly with the addition of nodes. Using Netflix 1,000,000 writes/second as a model you could roughly state that dividing nodes (285) by 60 should yield an approximately similar throughput requirement to the articles. So about 5-6 t2.xlarge nodes could potentially work for the FE cluster, and whatever gives you adequate streaming performance for the streaming cluster, as a minimum 3-5 x t2.2xlarge. Let's say ~$2,000/month.
Another thing to consider is that the management and maintenance of a C* cluster doesn't come for free either. So figure you'll need a minimum of 2 sys-admins to administer the cluster. Depending on where you are that'll likely cost the company upwards of $400,000 per year or $33,000k/month once you combine base salary, pension contributions, taxes, etc. You could say 1 sysadmin but then you're going to burn out your sysadmins if 1 guy is always on-call.
So assuming a sliding window of 30 days that's $2,160/day for DynamoDB. With a C* cluster you're looking at a minimum of $1,100/day. More depending on how well it performs for what your system is doing. If you're not using a sliding window you'd also need to scale you're C* cluster as well. In order to configure the C* cluster in a way that provides back-ups, auto-recovery, etc you're probably looking at about 2w-2m depending on experience. Whereas DynamoDB time to market will be probably a 1-2w.