From relational DB to a single DynamoDB table
trek10.com
trek10.com
Document storage freezes access patterns in a way that creates data locality for preferred data access patterns. The major downside is that you might not know all the access patterns and even if you do new access patterns can arise. Either way, it can be extremely difficult to accommodate unexpected access patterns.
I have seen document storage work well for aggregating data from external APIs. The whole point is that you are making preferred access patterns for your data rather than use an external API. And you are probably ready to rebuild your database if your access pattern changes.
I would also expect document storage to work well in an architecture of microservices under the conditions that:
* The schema is small so there aren't many possible access patterns
* Data is shipped to a SQL (warehouse) database for more complex processing anyways.
> * The schema is small so there aren't many possible access patterns
> * Data is shipped to a SQL (warehouse) database for more complex processing anyways.
That's pretty much exactly how I've seen it used.
edit: formatting
I've worked on a team of 8 engineers that spent 6 months building a streaming replication system from MySQL to Redshift. Its crazy how simple this would have been had the originating database been Dynamo; its very possible one engineer could have done it in a fraction of the time. It also seems likely to me that an open source solution might already exist.
Wrapping your head around this takes some pretty serious time along with trial and error. Work and rework your data model multiple times. Turning to spreadsheets actually tremendously helps with being able to define and move data around to play that game of Tetris mentioned in the article.
In addition, you'll probably want to avoid any of the common ORM type clients that exist for DynamoDB, they tend to have made decisions that make using a single table design an uphill battle. Write your own little wrapper for the native AWS DynamoDB client in your language of choice and you'll be much happier long term.
That seems to be the secret with a lot 3rd party components. Write your own wrapper that encapsulates the features you use but don't let that component dominate your codebase.
You also very likely need to use DynamoDB transactions (not really mentioned in the article) or else you're taking big integrity risks with partially incoherent records and so you're paying more for their transaction bookkeeping as well.
Doesn't this seem to map badly to DynamoDB's cost model?
One fun thing: this was my first time playing with the new DynamoDB on-demand capacity features. If you do run the code, check out the CloudWatch metrics on the table. It's pretty cool to see the WCUs shoot up to 150 and back down to 0 with no throttling whatsoever.
They initially implemented capacity units to be spread evenly at the partition level and now they have finally implemented it to work at the table level (which is how they sell it and how they talk about it).
Big whoop, you still have to ask daddy if you want to write at more than 80MB/s per account.
Normally, you just write to a data store until it's done. The faster the disks the less time it takes.
Now, writes can "fail" so you need to wait/retry or put it back on a queue to reprocess when you have write units available (both of those and the associated logging will cost you $ on AWS).
I have a dynamo addon for Heroku apps and need to keep cost under control. I use AWS Budgets and put a fixed price monthly budget on every table. If users go over the budgets access keys get disabled.
IIRC, you can setup events based on dynamo records that can trigger a lambda function to update the search db.
And always define reporting requirements up front. The last thing you want to find out is that you need a real-time dashboard and you don’t have the services or storage to support it.
The general solution if you think you are going to get to volume where that is not sufficient is to do something like add an additional integer (or other known symbol) to the end of the key. Ex: `key.1`, `key.2`, `key.3`. When doing operations you would have to run N operations across the key + symbol space. Not as elegant and you'll pay the additional computation for those queries, but since they can all be run in parallel it's not too bad in terms of latency.
Now, if all you are doing is hammering away at a few rows, you are stuck for sure. DAX [1] or some other caching layer may be your only way out of that problem.
This should add literally pennies, if that, to the total cost and is relatively effective. Of course, it isn't "scale to zero", its just never letting it go to zero. But the practical and financial difference is almost nothing.
Many various GCP databases (most notably Firestore[0], and Datastore[1] before it) do this today. No provisioning, pre-warming, or capacity planning needed.
[0] https://firebase.google.com/docs/firestore/ [1] https://cloud.google.com/datastore/
Then it should also wrap your new table up in an ORM to protect you from yourself.
Could also support generating migrations to support adding new access patterns, at the cost of modifying the entire database.
What I don't like about the auto scaling is it seems to be limited to either scaling reads or writes automatically. That seems short sighted. I don't understand the technical reasoning.
Or you can turn on the very recent “on-demand pricing” and not worry about capacity settings.
I mean, all customers have something like pk=customer1 sk=customer so all customers are in the same partition of the GSI. I know it’s a big deal To distribute the table partition key but doesn’t that also apply to the GSI?
What happens when they change?
If say you're doing dotabuff type stuff then just chunking the records you want out by player or even just an S3 Blob that is updated might be better. Then the rest of your tables are literally just indexes to the large amounts of data.
It totes depends on what you're doing.
The big issue with relational really is that you hit a point where transactions just can't keep up with locking across the cluster. If your system doesn't do this, then your next problem is just sharding on tons of data. If you don't have enough data then relational is fine. If you do have that much data then you're going to have to figure out a sane sharding and joining strategy across a cluster, which is more work. At some point in size of data you either have each user on their own shards of boxes, or you are keeping only indexes in the databases to retrieve other data. Both take a bunch of work.
Much of the motivation behind that talk and amazon's drive away from SQL is really just the business risk of hitting that scaling limit while you have a business that is doubling in revenue. Your options when you hit a limit on sql (esp with transactions) is to heavily re-architect it. And that takes a while, so you have to pause growth during that time. At which point a second mover can come roaring past you.
If I was building somthing like dotabuff I'd go with relational database sharded by users or some type of columnar database for the ad-hoc querying. Tho really for dotabuff you could build almost the entire site to be statically served off of cloudfront and just re-crunch data when a new match comes in.
However I'd have dynamo as the front end data ingest record keeper. Things like "did I see this game data already" to avoid using things like FIFO queues.
Happy to talk about it more if you provide more examples of what you're trying to map.
And again, if you don't have these problems, relational is just totally fine. You can scale quite high on relational, it's just that you run the risk of hitting that asymptotal tipping point where the system just goes down and you can't get bigger hardware. If you're not going to hit it, the tool you know will serve you well.
Another random thought, I find that people who thing in SQL forget how damn awkward it is... For instance, ask someone who uses Relational regularly how to create tables for users's shipping address. They'll spout a denormalized form that works in many many places. Ask a college hire the same question, you might get more than one table. The thing that is obvious to the regular user will be non-obvious and confusing to the college hire. And I would argue that the table layout is somthing that evolved over 30+ years of the industry doing it over and over.
The only document store I've used "in anger" has been DynamoDB, does this technique apply well to other document stores?
I think this nicely illustrates why most projects shouldn't use a high-scalability NoSQL database: they benefit much more from flexible data models and efficient query.
All that being said, there are other factors when choosing a database aside from those three. For some of our microservices, my team doesn't need scalability or flexible queries, just efficiency. However, we've found that it is cheaper/easier for use to use DynamoDB than RDS.