Problems with AWS Amplify and its use of DynamoDB
samthor.au
samthor.au
The biggest problem with Amplify is that it hides complexity that the developer really needs to have a firm grasp of when designing their application. You really need to think a lot about your access patterns before using DynamoDB. DynamoDB is OLTP and not OLAP. You aren't going to have a ton of ability to do ad-hoc queries with DynamoDB. that's just not a use case it's well suited for.
HOWEVER - AWS Amplify is marketed without that warning - people use it without really understanding that. Then they get stuck when they start trying to treat it like a traditional RDBMS datastore, which it is not.
Amplify "dumbs down" something that really should not be dumbed down. That's it's biggest problem and the reason I also stopped using Amplify some time ago.
Products like this are great. But if you already know the technology, know what is happening and you are using it to basically eliminate a bunch of coding or setup tasks that you would have to do.
The problem this is only great for a certain type of knowledgeable developer who is also setting up the system.
The next person that comes or a junior developer who does not have grasp of the underlying tech will just see a bunch of magic. But he will be required to be productive with it.
So if the original devs leave and don't get people with grasp of the underlying tech, the new people will start accumulating technical debt at an ever faster rate without ever being able to pay it back. Because with so little knowledge and expectation of performance, whenever something small happens that requires you to learn stuff you will skip learning and just move stuff around hoping some random change will fix it.
So I will amend my original statement:
"Products like this are great." But only if you are working on the project alone and don't expect anybody else to join it in the future.
This was a trap that I fell into. While I didn't go with Amplify, I went with DynamoDB assuming that "it's a high performance database, surely I'll be able to bend it do what I need it too, right?". I feel like AWS should be more upfront about the limitations of DDB, it's a great database but they sell it like the solution to everything, and it simply isn't.
Nonsense. DynamoDB is a pretty good KV store, and a killer feature as part of AWS' free tier. I dare you to present a cheaper and more performance KV store.
Really nothing that could earn my heart over traditional RDMS offerings.
No proper REPL experience.
The barebones SDK.
That is how much I can still remember from 2018's experience using it.
Dynamo has its use cases. Imagine a database with 100K+ of devices sending different metrics out every minute or so. You need to retrieve metrics by device ID:metric and date/time range. I migrated a similar DB from Cassandra to Dynamo, and it worked really well (for that use case!)
In the years since, I've worked at other companies and seen Dynamo used where it was a really poor fit: tons of filtering/scanning, relatively tiny datasets that require many secondary indexes. I even know of one app that starts up and loads an entire DDB table into memory! If you're not absolutely sure Dynamo is a good fit, it isn't. Use a relational database.
I never heard anyone complain that they had to explicitly create primary keys when creating tables in RDBMS.
> or query types for what one can search for.
I never heard anyone complain that they had to put together SQL queries to get data out of a RDBMS.
It seems your only complain boils down to pointing out nosql databases aren't SQL databases, and you didn't bothered to onboard onto the tool you knew nothing about. That's hardly insightful.
I have heard a lot of complaints on it not doing as much. But I swear every time I've taken a dive on those complaints, it has been by folks that don't pay attention to the costs of what they are wanting. :(
It's just odd to me that I had to make this change in the first place. I wonder how many millions of dollars AWS is raking in just because devs are unaware of things like that. Also makes me wonder how much energy is being wasted on inefficient DB operations..
A lot of this is on the author IMO
Fixed that for you.
AWS documentation very clearly outlines the behavior of a query and a scan. The internet is full of guides [1][2] on how to model your data so you utilize queries vs. scans.
If the developers at your company misuse a product, and aren't aware of these features when does personal responsibility come in? How is AWS as a platform suppose to discern that the trade-off with a GSI is what you want, and how is this scenario not generalized to every product? People misuse SQL databases all the time, for example.
1 - https://www.alexdebrie.com/posts/dynamodb-single-table/
2 - https://www.sensedeep.com/blog/posts/2021/dynamodb-singletab...
I'm not going to assume that DynamoDB was forced upon them. I don't know. What I'm more sure about:
Developers have a choice in if they want to invest the time in understanding the tools they use, and learning about the tool's capabilities. In this specific case, we're not talking about a lot of time. 2-3 hours to read the documentation for a new database you're going to be building your business on top of.
What is your argument exactly?
For example: for years, MongoDB took a lot of shit because by default it spins up unauthenticated. Hypothetically; devs take on the responsibility of adding a password. Realistically, shodan was able to index hundreds, maybe thousands, of databases across the internet with no authentication, some containing highly sensitive data.
Firestore is a great analogue to DynamoDB. Not perfect, but they solve a lot of the same problems in a lot of the same ways. Firestore automatically indexes every field in every document in every collection. Flat out. Everything is on an index. That single decision has reverberated to save tons of hours in compute time across all of their customers.
Is this expensive to do? I have no doubt; and its reflected in the cost of the two services. But believe it or not: Firebase isn't that much more expensive than DynamoDB ($0.25/million reads, versus $0.60/million).
I don't disagree, but I think there's consensus on what we consider footguns and what we consider features. Default on security footguns (like unauthenticated access to S3 buckets) should be changed. AWS has changed this. Default on performance features that cost customers money, and may be unused? We should limit those.
> That single decision has reverberated to save tons of hours in compute time across all of their customers.
It's hard to make a statement like this without data. How many of those fields are wasted disk space on a server somewhere under utilized? How much compute does indexing cost on those under utilized fields that should have been omitted?
You make the trade-offs on if you use Firestore or DynamoDB. They're different products, with different features.
When you try to do so, it spits out an error that tells you this operation requires an index and then provides a link to create the specific index.
I know it is a different database paradigm, but I think this is great "UX" design.
Firebase has a bunch of additional great features in this vein. It allows definition of min/max instances for Functions and if you set the min instances in the code, it will force a confirmation of price changes to prevent any unexpected charges. (GitHub Actions will fail, for example, until the price increase is confirmed via CLI.)
So i get OP. Yes, it is the responsibility of the engineer to know these things and make decisions, but the vendor can provide much better UX around these constraints.
Caveat: I used to work at Amazon (not in AWS), but don't anymore.
Sort of, though I'd say scan operations are typically more expensive in any RDMS.
The probably becomes when you want/need something more than the secondary index, then you gotta get more creative.
The cloud service is crappy architected, developed Java application but instead of fixing it , solution is to double/triple or 10 times the instance using Kubernetes Autoscaler. One thing they have going is ever larger budget since it is next gen cloud service, whereas my serivce being plain old efficient Java app is legacy to be deprecated and replaced with new cloud version.
To a lot of engineers there's always bigger fish to fry.
In any event, while I find the structure and docs pretty confounding overall, they do make that behavior pretty clear. The libraries even separate out scans from queries, with the latter taking only a key.
I'd put the blame on the implementer in this case, not AWS.
Part of me does wonder if there would be a benefit in DynamoDB providing a function that under the hood creates secondary indices and uses queries where that would be the most efficient choice automatically. But maybe that would be a better fit for an ORM of some sort.
1 - In a collection, I could find no way to easily do 'give me the latest record.', which is super easy in rdb. I ended up creating a key something static like 'x', with date as the range key to accomplish this. If anyone has a cleaner way, please share.
I can't even claim to be immune from this, though I do try to know what's available and how to use it. I'm damn near 40 but it's not unusual for me to still, today, stumble on some tool older than me that does exactly what some New Shiny crap does—but, not infrequently, better, at least for some common use-cases, plus it'll have all the kinks ironed out, or at least documented. Good luck promoting any of that in the workplace, though. For all we go on and on about the importance of re-use and avoiding NIH and all that, we sure are quick to disregard solutions that we can't npm-install.
No clue what the fix for this is, or if it even makes sense, all things considered, to try to fix it.
I'd rather have that than an industry where any tech newer than 10 years is radical and scary.
Us old farts can remind the kids that old solutions exist when appropriate. It is much harder to try introducing something new into an org that doesn't want it, but needs it.
Engineers often see a database simply as a place to store things, it's just a box.. they work on a product for awhile that's using mongo, they are more likely to choose it later when they are developing a solution because they've used it before and know it 'works' and more importantly they know something about it. Databases as a topic of study are complicated, there are tons of tradeoffs happening all over the place and until you hit any given pothole you may not really understand those tradeoffs.
I guess what I'm saying is that in my life I'm lucky enough to have an old friend that is REALLY into databses. You should find this friend or become this friend.
Honestly at the end of the day this is all human nature. The hammer worked for you before why do the research to find out that there is a better hammer for you to use that would get your job done quicker and easier. You've gotta get these boards nailed down RIGHT FUCKING NOW or your bosses are going to be pissed.
Oh man, I had an edit right after posting where I almost attached an addendum about this specifically, but then decided against it to avoid distracting from my core point. There's a totally bizarre allergy to actually using database features, for fear of "lock-in" or something, yet most programs are more likely to be re-written on top of the same database than to have the database swapped out from under it, unless someone made a really god-awful choice of DB, or load scales tremendously (which falls under "good problem to have, and you can afford the migration"). The result is that we as an industry waste an awful lot of time and money creating often-buggy application-layer solutions to things that could have been solved by leveraging features of, for example, PostgreSQL, with results that'd be cheaper, less-buggy, and better-performing. Drives me absolutely nuts.
Or, like, we have nginx in our stack and it could do something we need done, with a config change or a little Lua... but no, we'll spend 10-100x as long to make a worse-performing custom solution to this problem that a thing we're already using can trivially solve for us, but we'll do it in Javascript or Ruby or Python or whatever, as part of our "app". Ugh.
And don't get me started on system-level tools and daemons. Like, we're already using Linux or FreeBSD or whatever, which come with a stupid-large set of great, proven tools, so how about we actually use it rather than just using it as a fancy DOS and paying for SaaS to do things that our OS or distro can already do, probably more reliably? But no, we don't even bother to configure a recipient for root alert emails more often than not, and we just treat the whole thing as a dumb application runner that needs some complex external support system to keep it working right.
As you said, it's just humans, really. How many self-help books about organisation are out there? Organising your files, organising your clothes, organising your tools -- it's not a solved problem, there are always tradeoffs, and it's very easy to just choose the same solution you've used before, because you already understand how to implement it.
TBH, I think oftentimes it's the right choice to just use any old tool as a hammer[0]. In many (perhaps not most?) cases, you're probably going to save more time implementing the organisational strategy you already know than you are implementing the technically-correct strategy. Though, you probably still need to do some quick maths on that tradeoff -- what's the cost of choosing the solution you know?
By way of example, when we were rebuilding our incident management stack at Dropbox, we spent a lot of time considering our database structure, because one of our key design goals was to reduce friction. Our existing tool had appalling load times, which was (in part) causing users to avoid declaring incidents. We knew we had to create a very low-friction experience to encourage users to declare incidents, and that this had long-term ramifications. So, myself and the other senior engineer on the team invested several days between us determining the correct design for our database. Conversely, when I was building a tool to do some trouble ticket automation, I just implemented things in a simple manner with questionable performance -- because the implementation only cost me 30 minutes, and the impact of a 5-second scan vs. a 0.5-second query was essentially nil (the absolute worst-case scenario was an automated process took a few extra seconds and cost a couple cents a day). No reason to overoptimise.
0. Adam Savage even titled his book "Every Tool's A Hammer"!
What "things like that"? Do you mean the basics of querying data with DynamoDB?
https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
This is literally the second topic in any intro to DynamoDB course, that comes right after creating your first DynamoDB table. Even that covers explicitly the basic requirements of running queries.
> What should AWS do -> provide Postgres or similar as a database choice
If you're interested a postgres option you could check out Supabase GraphQL[1] which is based on pg_graphql [2], a native PostgreSQL extension.
You define your schema in SQL (including any indexes you'd like) and it reflects a full GraphQL API. With that stack, the example you gave about filtering for security would use be solved using a Row Level Security [3] policy where the work all occurs on the DB.
While not a part of pg_graphql, using Supabase as a backend also handles Auth [4] & Storage [5] (both OSS) which covers most of the Amplify bases
[1] https://supabase.com/docs/guides/api#graphql-api-overview [2] https://github.com/supabase/pg_graphql [3] https://www.postgresql.org/docs/current/sql-createpolicy.htm... [5] https://supabase.com/docs/guides/auth [5] https://supabase.com/docs/guides/storage
Even if the startup is starting web only, they're kneecapping their future so its not an option
Amplify has amazing mobile sdks (esp ios).
What if they make a mobile SDK too ? Seems like your comment is an hyperbole.
Swift SDK is here: https://supabase.com/docs/reference/swift/introduction
Amplify's core architectural presumption is to only use serverless options to try to keep costs low. Serverless SQL databases are currently a big hole in AWS's offering. Yes, there's Aurora Serverless, but you need to run it in a VPC, you can't scale to zero (so, for some definitions, not really serverless), v2 threw out their Data API, and v1's Data API had severe scaling issues anyway. v2 is so expensive that you have to very seriously think through the numbers on the small RDS databases and ask what value you're actually going to get out of v2. So what OP proposes can't really be done by the Amplify team.
That said, the fact that Planetscale and Neon (which also have branching!) continue without serious competition from AWS is a serious headscratcher.
You have another potential problem in your stack, but it can be done.
It's just a hassle that's best avoided unless you really don't have any other choice or the financial math actually works out for doing so.
From experience, AWS Amplify auth for interacting with aws cognito is a perfectly acceptable high value (low cost) solution for many use cases. I can't speak to the other technologies.
https://www.passportjs.org/ is this an alternative or not really because you still have to "build it yourself"?
And then goes on to say:
> But for most people, it's probably going to literally just be Postgres. You don't need more than Postgres.
Those statements are basically contradictory. Maybe the author only works on applications where NoSQL is a good fit. Otherwise, it's not clear how the author is different from "most people".
This idea that SQL databases are not "scalable" has become a meme and is difficult to eradicate. Sure, sharding/replicating Postgres (if you need to) can be more difficult than just increasing a field in some auto-scaling group. But a well-tuned Postgres can service a ridiculous number of requests. If you are doing microservices, then it's even better, since you'll have multiple smaller databases - some of these can even be NoSQL if it fits the application.
What 'scalable' NoSQL databases usually means is that you can just throw money at the problem and increase the number of instances, without spending time doing silly things like 'optimizing' queries.
Of course, for some applications, NoSQL is a perfect fit. But far too often I see people reaching out to NoSQL based on some future (usually unspecified) "scale" needs. Heck, I've even seen that mindset in internal, corporate network only apps.
It feels like the team is trying to do way too much, all at once, and none of it well.
But this article has me thinking about exploring other options. To try to get ahead of the seemingly inevitable wave of tech debt. Does anyone have advice or can point to any resources for embarking on that particular journey? Are there similar technologies that do it better than Amplify? Anything in between Amplify and learning all the disparate AWS tools from scratch?
It depends how much of the backend you ultimately would like to avoid (and related to that, how complex your service ecosystem will be for your app).
Without knowing much more about your app, I can throw some darts. You could use RDS to get you postgres/maria/etc and then just run an app server on EC2 for your backend REST/SOAP/etc API in front of that. If the latter body of work isn't comfortable for you, there is a swath of low code tooling out there you could use as well to put an API in front of your SQL for you (nocodb comes to mind, or just postgrest if all you need is CRUD ops).
Custom authorizer in front of aws lambda with serverless framework works great with amplify’s frontend api client.
We also use dynamodb serverless but not with app sync and no graphQL. Would I rather have a good and easy to use serverless Postgres db that played well with aws lambda. Yes. Am I going to switch away because I have to plan my access patterns and indices out carefully… no.
Also he mentions the race condition around clients connecting to the backend between queries and mutations, then in the same breath says the DataStore client handles this but also it's mistake is being reliant on Subscriptions? This contradiction doesn't make sense.
There are some legit bugs in this article though that I've seen, like that Analytics one. It's really a Pinpoint issue though when you dig under the covers that I've been waiting for them to solve for years.
How does this compare to say... a $10/mo DigitalOcean Linux VPS that you deploy stuff to with like... Ansible?
So that's one difference. AWS is expensive, and its performance tiers for stuff like networking can be frustratingly opaque, but when you pay enough, you do get more-or-less what you paid for. And fighting DO's networking can get expensive in a hurry, in lost business or technician hours.
As for Amplify vs. Ansible, yeah, dunno.
[EDIT] Incidentally, this isn't the first time I've seen hosting of this general sort fall apart at the networking level—I ran into very similar problems with RackSpace over a decade ago, there was simply no way to come anywhere near taxing the hardware we were paying for before the pipe got saturated, like, there was an order-of-magnitude mismatch between the two, requests timing out en masse while the hardware ("what a good bargain this price is!") was practically idle.
[EDIT EDIT] It's probably no coincidence that it's easy to list and understand hardware specs, but difficult to communicate or evaluate real-world Internet-connected network performance without trying it—so it's probably tempting as hell to scrimp on networking while selling based on bang-for-the-buck in hardware allocation.
With DO we'll have some clients seeing great performance, and others who may as well be on a spotty dial-up connection at the exact same time. Frustrating, and there's not really anything you can do about it except switch to another host. I'd love to know what the story is with 'em.
Parts I liked:
- Callouts of the subscription model are good. I have used Firebase rtdb and firestore quite a bit, and agree it's much easier with those tools to just get an up-to-date view. With amplify, I often avoid subscriptions and just re-call list methods after updates (supported well in Apollo, but still unideal)
- The discussions on Amplify/NoSQL/Dynamo being difficult to pivot match my experience, especially when I was less experienced with NoSQL patterns
> Amplify's approach in using DynamoDB is literally called out as something you should avoid by AWS
Alex DeBrie is a wonderful resource, and I love his books, but I wouldn't go as far as to say that all of AWS calls out using multiple tables. There is lots of active debate on the single-table vs multi-table designs
> the fundamental mistake that's made here is that AWS Amplify puts your data into DynamoDB, which is not a general-purpose database
Citation needed. It's a fairly standard NoSQL database and can be used generally.
> But DynamoDB Scales?
I was a bit confused on this article between the frequent mentions of DynamoDB only being worth it at a large scale ("If you're storing that much plain text data, good for you—maybe DynamoDB can help you!", "But you should only use it when a traditional database won't cut it", etc.) while also claiming more broadly that Amplify will not be useful for large apps because it doesn't scale. It seems to contradict itself there?
I do think this section brings up a great point on the inflexibility of Amplify and trying to pivot a data model, which can absolutely be tricky! But the core message of this section doesn't seem to support the thesis.
> If I was AWS, I would… provide Postgres or similar as a database choice
I use RDS with Aurora/postgres with Amplify by using lambda resolvers and it works well. You can't use VTL, but in my mind that's a feature, as VTL is not my favorite for a few reasons like being hard to test.
In general, this was one of my larger frustrations with this article: it doesn't mention the flexibility of Amplify. Amplify comes with highly opinionated built-ins, but you don't have to use almost any of it, and it's quite easy to plug-and-play other tools if you want. Lambda resolvers are a good example because they're extremely flexible, you can use many different languages/tools/even docker to run your resolvers. But also, I opt to not use Amplify hosting and it's fine. And you can also use your own auth without Cognito if you'd like (especially if you're already using custom resolvers).
It sounds like the author inherited some parts of the codebase in question, so if some of these things were already setup I could see it being time consuming to pivot.
> rethink whether GraphQL is actually fit for purpose: It's just not very good
Again, while I'm not a huge GraphQL fan, this is very much an opinion and a highly debatable one. For many use cases, GraphQL can work well, and I think from a technical perspective it is at minimum arguable that something like GraphQL makes sense for a tool like Amplify where one of its core use cases is multi-platform usage across web/mobile.
Mongo has a lot of performance gotchas with things like pipelines / aggregations that don’t scale the same way they do in Postgres & similar
Took a while but we eventually migrated to a plain rest api with a react front end. Still used mongo
Meteor was fantastic for building highly reactive apps with a small team. I think it got a lot of bad run for scaling up because it was dependent (at least when I used it) on tailing MongoDB's op log for reactivity. That lead to having to make very different architecture choices to make it work well, and I seem to recall each server instance getting bogged when they had 70-100 concurrent users , so you ended up using lots of containers for meteor application servers. Not sure where it all ended because I moved on to a different platform three yeas ago.