A minimalistic Python wrapper to AWS DynamoDB
github.com
github.com
https://github.com/dineshsonachalam/Lucid-Dynamodb/blob/mast...
On top of that they have mutable keyword arguments and generally quite terrible code: no exceptions thrown, just logging.error(), `if len(x)` rather than `if x`, etc etc.
Isn’t that always the case?
(I especially like how they market any schemaless database capable of strong consistency in the CAP sense as fully ACID compliant RDMS alternative, because if you can’t specify a rule making database states invalid than the consistency in the ACID sense is trivial because every database state is valid.)
It looks like a good place for quick key-values lookups. It's a "table" as a service, with 1 (or 2) columns that you can index on. Anything else and you need an RDBMS.
It's useful for simple storage, without the overhead of setting up an RDS.
If you have a team of people that know how to use it well, I’m sure it’s an excellent tool. But it shouldn’t be casually thrown into an architecture, because there is nearly always a better solution.
Even the documentation of DynamoDB says so:
https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
This means that it's a good use case for shopping cart of Amazon since there's limited innovation possible there and it fits the bill nicely, but for a new application with relational data where the requirements change all the time - it's a terrible, terrible choice for which you will pay dearly.
Dynamodb has some really great features and high performance when used correctly but it also has a steep learning curve especially for anyone coming from sql and other relational databases.
AWS tries to push the concept of single table design which allows you to store an entire application's data in a single dynamodb table and using a bunch of complicated indexing and overloading columns with multiple types of data to sort of get relational data without any joins.
This is really easy to get wrong and you need to have all your access patterns figured out ahead of time so that you can create the indices you'll need when you create the table. You can't create new local indices later... only global secondary indices which duplicate the entire table (and double your storage costs).
Dynamodb streams are really cool - great way to setup webhooks for data changes.
Dynamodb pay per request is super cheap and fast for a serverless db. Aurora serverless isn't pay per request and has slow cold starts.
My recommendations if you are using dynamodb because you want something serverless and NOT because nosql is actually the best storage option:
* don't do single table design - just make more requests to different tables and join your data in code
* if joining in code is too slow then use a sql database...
* abstract all your data access calls so you can eventually migrate to a sql database
Would love to chat more about this sometime if you were interested!
I use that term because my system only allows 1 true level of depth before you have to query again. You also have to predefine the data model, but that's so that it can automatically choose the right entity type (to format the PK/SK correctly) and so that it supports automatic Global Secondary Index selection as it dynamically generates the query. It also means that each entity can only have up to 5 extra indexable key pairs.
Well, I was wrong. I just looked it up and you can now have 20 GSI. It used to be 5 haha.
Anyway, it would be cool to talk about your project. I'll check it out this week.
It is best in extreme cases. Some types of high scale use and super super cheap serverless/bootstrapping startup smaller scale.
Low-scale - it is easy to create a table and just start writing to it. The pay per request pricing is simple and plays well with IAM, other AWS services and the serverless framework.
High-scale - scales easily up and down, global replication is easy, backups and high availability are easy - it is all self-managed so a lot of this is baked in or very easy to toggle on. Depending on your ops resources, the easy tooling may be worth the extra costs or maybe your data lends it self particularly well to how dynamodb works.
Some of our use cases ....
- user profiles to supplement cognito user attributes
- chat/chat groups/chat group memberships/chat websocket watchers/chat messages (this one was not good. I would do it again with just messages and the websocket watchers)
- tenant data with lots of custom attributes
- forms and form submissions
- event webhooks from 3rd party services - catch/store/process using dynamodb streams
- step function states and logs - it has some nice versioning write patterns
- web scraping data from ebay
- web scraping data from golf tee time sites
When you need search or flexible querying you can put elasticsearch in front of it and use dynamodb streams to push items into the index. If you need analytics you can use another layer like redshift.
This book is quite good. Used it to get a deeper understanding of Dynamodb.
Thanks for the laugh. Dynamodb is stupid expensive under any decent workload.
The free tier isn't limited to one year and low-usage is inexpensive.
db.get('Pk.sk') or db.get('Pk.sk', ingore=['one'])
db.get('Pk.sk', model=pydantic_model)
db.update(... db.put(...
Ect
Leave a comment if people are interested in me releasing this project.