DynamoDB 10 years later
amazon.science
amazon.science
Engineering, however, was a disaster story. Code is horribly written and very few tests are maintained to make sure deployments go without issues. There was too much emphasis on deployment and getting fixes/features out over making sure it won't break anything else. It was a common scenario to release a new feature and put duct tape all around it to make sure it "works". And way too many operational issues. There are a lot of ways to break DynamoDB :)
Overall, though, the product is very solid and it's one of the few database that you can say "just works" when it comes to scalability and reliability (as most AWS services are)
I worked at DynamoDB for over 2 years.
>Overall, though, the product is very solid and it's one of the few database that you can say "just works" when it comes to scalability and reliability (as most AWS services are)
How those two can coexist?
Alternative would be to attempt a near 'perfect' solution for the product requirements and that may either hit an impossibility wall or may require substantial long term effort that would impede product development cycles. So likely the former approach is the smarter choice.
source: currently being burned out on an adjacent aws team..
if you want to know why capitalism causes this, start a startup and prioritize quality, do not get to market, do not raise money, do not pass go, watch dumpster fires with millions of betrayed and angry users raise their series d
I've seen this kind of thing mentioned many times, pretty baffling TBH based on Dynamo's pretty good reputation in industry. Are these mostly to the stateless components of the product, or do they see data loss?
There are times when bad deployments happen and customers were impacted.
Enjoy the sausage, but if you have a weak stomach, don’t watch how it’s made.
(I work for AWS but not on the DynamoDB team and I have no first-hand knowledge of the above claim. Opinions are my own and not those of my employer.)
This is true though there's only so much technical debt and internal process chaos you can create before it affects the outcome. It's a leading indicator, so by the time customers are feeling that pain you've got a lot of work in front of you before you can turn it around, if at all, and customers are not going to be happy for that duration.
Technical debt is not something to completely defeat or completely ignore, instead it's a tradeoff to manage.
One concrete problem with technical debt the article highlights is it that negatively impacts the time to deliver new features. Customers today usually expect not only a great initial feature set from a product, but also a steady stream of improvements and growth, along with responsiveness to feedback and pain points.
Additionally, the business cares about the outcome, not the internal process.
Ostensibly, the business should care about process but it actually doesn't matter as long as the product is just good enough to obtain/retain customers, and the people spending the money (managers) aren't incentivized to make costs any lower than previously promised (status quo).
When this is the case it’s often nice to state this conflict of interest, so others can take your appraisal in the appropriate context.
I’m not implying anything about the post, just stating what I assume to be the reason for the disclosure.
This is designed to reduce the chances of eager employees going out and astro-turfing or otherwise acting in trust-damaging ways while thinking they're "helping".
To me the takeaway is large/interesting/challenging engineering projects are pretty close to disasters generally. Some time they do become disaster actually.
On the other hand if a project looks like straight up designed, neatly put into JIRA stories, and developers deliver code consistently week after week then it may be a successfully planned and delivered project. But it would mostly be doing stuff that has already been many times over and likely by same people on team.
At least this has been my experience while working on standardized / templated projects vs something new.
Some projects are run like the Doomsday Clock, and nobody can get anything done. Other ones increase and decrease on complexity all the time, and those tend to catch-up to the first set quite quickly.
Cassandra is too opinionated and its CAP behavior wasn't great for a service like this, so they built on top of Riak. (This also eliminated any thoughts I had about Erlang being some uber-language for distributed systems, as there were (are?) tons of bugs and missing edge cases in Riak)
Not Invented Here can run very deep in some branches of an organization. Depending on how engineering performance evaluations work, writing a homebrew database could totally be something that aligns with the company incentives. It might not make a single bit of sense from a business standpoint but hey, if the company rewards such behavior don't be surprised when engineers flush millions down the tube "innovating" a brand new wheel.
Questions were handwaved away, and the usual Amazon black box non-answers which always smells like they are hiding problems.
Any ideas how this is working? It seems bolt-on and not well thought out, and I doubt they'll ever pay for Aphyr to put it through his torture tests.
Consistency and conflict resolution
Any changes made to any item in any replica table are replicated to all the other replicas within the same global table. In a global table, a newly written item is usually propagated to all replica tables within a second. With a global table, each replica table stores the same set of data items. DynamoDB does not support partial replication of only some of the items. If applications update the same item in different Regions at about the same time, conflicts can arise. To help ensure eventual consistency, DynamoDB global tables use a last-writer-wins reconciliation between concurrent updates, in which DynamoDB makes a best effort to determine the last writer. With this conflict resolution mechanism, all replicas agree on the latest update and converge toward a state in which they all have identical data.
If the customers don't have any feedback or missed feature asks at launch, you waited too long to ship.
You know who has great internal code and test quality? Google. Which is why Google doesn't ship. They're a wealth distribution charity for talented engineers. And their competitive advantage is that they lure talented people away from other companies where they might actually ship something and compete with Google, to instead park them, distract them with toys, beer kegs, readability reviews, and monorepo upgrades.
When looking at DynamoDB I noticed that there was a surprising amount of discussion around the requirement for provisioning, considering node read/write ratios, data characteristics, etc. Basically, worrying about all the stuff you'd have to worry about with a traditional database.
To be honest, I'd hoped that it could be a bit more 'magic', like S3, and it AWS would take care of provisioning, scaling, sharding etc. But it seemed disappointingly that you'd have to focus on proactively worrying about operations and provisioning.
Is that sense correct? Is the dream of a self-managing, fire-and-forget key value database completely naive?
For example, you used to have to manually set RCU/WCU to a high number when you expected a spike in traffic, since the ramp-up for on-demand scaling was pretty slow (could take up to 30 minutes). But these days, on-demand can handle spikes from 10s of requests a minute to 100s/1000s per second gracefully.
The downside of on-demand is the pricing - it's more expensive if you have continuous load. But it can easily become _much_ cheaper if you have naturally spiky load patterns.
Example: https://aws.amazon.com/blogs/database/running-spiky-workload...
We've some regular jobs that require scaling up dynamodb in advance few times per day, but then dynamo is only able to scale down 4x per day, so we're probably paying for over capacity unnecessarily (10x or more) for a couple hours a day
Now we just moved ondemand and let them handle it, works fine
True, although you don't have to make that choice permanently. You can switch from provisioned to on demand once every 24 hours.
And you can also set up application autoscaling in provisioned mode, which'll allow you to set parameters under which it'll scale your provisioned capacity up or down for you. This doesn't require any code and works pretty well if you can accept autoscaling adjustments being made in the timeframe of a minute or two.
There could exist a fantasy database where you still tell it your hash and range keys, which are roughly how you tell the database which data isn't closely related to each other and which data is (and which you may want to scan) but instead of hard provisioning shard capacity it automagically splits shards when they hotspot and doesn't rely consistent hashing so that every shard can be sized differently depending on how hot it is.
Right now such a database doesn't exist AFAICT as most places that need something the scales big enough also generally have the skill to avoid most of the pitfalls that cause problems on simple databases like Dynamo.
They now have a dynamic provisioning scheme, you simply don't care but it is more expensive so if you have predictible requirements it is still better to use static capacity provisioning. There is an option though.
DynamoDB also requires the developer to know about its data storage model. While this is generally a good practice for any data storage solution, I feel like Dynamo requires a lot more careful planning.
I also think that most of the best practices, articles etc apply to giant datasets with huge scale issues etc. If you are running a moderately active app, you probably can get away with a lot of stupid design decisions.
https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
Beyond that, though, it's really not designed for that kind of use case.
I read "the dynamodb book" and almost got a stroke. So much idiosyncrasies, for what?!
Azure has CosmosDB, GCP has Cloud Datastore/Firestore, and there are many DB vendors like Planetscale (mysql), CockroachDB (postgres), FaunaDB (custom document/relational) that have "serverless" options.
DynamoDB is almost simple enough to learn in a day. And if you're doing nothing with it, you're only really paying for storage. Good luck with your decisions.
Everything has limits, but S3 is remarkably hard to break if used right.
This inherently leads to a complexity debt explosion, fragmentation in the experience, and an operationally brittle posture that becomes very difficult to dig out of (this is probably why AWS loves the paradigm).
Those databases build most of that in, and it's all one fairly excellent distributed monolith.
I am working with a company that is redesigning an enterprise transactional system, currently backed by an Oracle database with 3000 tables. It’s B2B so loads are predictable and are expected to grow no more than 10% per year.
They want to use DynamoDB as their primary data store, with Postgres for edge cases it seems to me the opposite would be more beneficial.
At what point does DynamoDB become a better choice than Postgres? I know that at certain scales Postgres breaks down, but what are those thresholds?
https://aws.amazon.com/rds/aurora/?aurora-whats-new.sort-by=...
Certain access patterns can do pretty well with 3,000 relational tables denormalized to a single DynamoDB table.
I've found also that in Postgres the query performance does not keep up with bursts of traffic -- you need to overprovision your db servers to cope with the highest traffic days. DynamoDB, in contrast, scales instantly. (It's a bit more complicated that that, but the effect of it is nearly instantaneous.) And what's really great about DynamoDB is after the traffic levels go down, it does not scale down your table and maintains it at the same capacity at no additional cost to you, so if you receive a burst of traffic at the same throughput, you can handle it even faster.
DynamoDB does a lot of magic under the hood, as well. My favorite is auto-sharding, i.e. it automatically moves your hot keys around so the demand is evenly distributed across your table.
So DynamoDB is pretty great. But to get the the best experience from DynamoDB, you need to have a stable codebase, and design your tables around your access patterns. Because joining two tables isn't fun.
More than just joining--you're in the unenviable place of reinventing (in most environments, anyway) a lot of what are just online problems in the SQL universe. Stuff you'd do with a case statement in Postgres becomes some on-the-worker shenanigans, stuff you'd do with a materialized view in Postgres becomes a batch process that itself has to be babysat and managed and introduces new and exciting flavors of contention.
There are really good reasons to use DynamoDB out there, but there are also an absolute ton of land mines. If your data model isn't trivial, DynamoDB's best use case is in making faster subsets of your data model that you can make trivial.
This made my head explode. Why would you explicitly join two systems made to solve different issues together? This sounds rather like a lack of architectural vision. Postgres's zero access-design inherently clashes with DynamoDB's; same goes with ElasticSearch scenario: DynamoDB's was not made to query everything, it's made to query specifically what you designed to be queried and nothing else. Redis sort-of make sense to gain a bit of speed for some particular access, but you still lack collection level querying with it.
In my experience, leave DynamoDB alone and it will work great. Automatic scaling is cheaper eventually if you've done your homework about knowing your traffic.
My experience agrees with yours and I'm likewise puzzled by the grandparent comment. But just a shout out to DAX (DyanmoDB Accelerator) which makes it scale through the roof:
Judging a consistency model as "terrible" implies that it does not fit any use case and therefore is objectively bad.
On the contrary, there are plenty of use cases where "eventually consistent writes" is the perfect use case. To judge this as true, you only have to look and see that every major database server offers this as an option - just one example:
https://www.compose.com/articles/postgresql-and-per-connecti...
I have a theory it would be better to have multiple table-replicas for read access. At application level, you randomize access to those tables according to your read scale needs.
Use main table streams and lambda to keep replicas in sync.
Depending on your traffic, this might end more expensive than DAX, but you remain fully serverless, using the exact same technology model, and have control over the consistency model.
Haven't had the chance to test this in practice, though.
And DynamoDB is worse than most.
My prediction is that the future is in scalable SQL; CockroachDB or Yugabase or similar.
NoSQL actually causes more problems than it solves, in my experience.
Which I don't. I'd rather see reliable operation than "predictable except for when it fails outright" in almost every situation.
If you've encountered that other situation, where failures are fine? Then great. But I still assert that's a tiny minority of real-life DB use cases.
HN entrepreneurs take note, this also suggests to me that there may be a market for a database (or a "metadatabase") that takes care of this for you. I'd love to be able to have a "relational database" that is also some "NoSQL" databases (since there's a few major useful paradigms there) that just takes care of this for me. I imagine I'd have to declare my schemas, but I'd love it if that's all I had to do and then the DB handled keeping sync and such. Bonus points if you can give me cross-paradigm transactionality, especially in terms of coherent insert sets (so "today's load of data" appears in one lump instantly from clients point of view and they don't see the load in progress).
At least at first, this wouldn't have to be best-of-breed necessarily at anything. I'd need good SQL joining support, but I think I wouldn't need every last feature Postgres has ever had out of the box.
If such a product exists, I'm all ears. Though I am thinking of this as a unified database, not a collection of databases and products that merely manages data migrations and such. I'm looking to run "CREATE CASSANDRA-LIKE VIEW gotta_go_fast ON SELECT a.x, a.y, b.z FROM ...", maybe it takes some time of course but that's all I really have to do to keep things in sync. (Barring resource overconsumption.)
The concept didn't appeal to me very much then, so I never looked into it further.
---
To address your larger point, I think Postgres has a better chance of absorbing other datastores (via FDW and/or custom index types) and updating them in sync with it's own transactions (as far as those databases support some sort of atomic swap operation) than a new contender has of getting near Postgres' level of reliability and feature richness.
You might be interested in what we're building [0]
It synchronizes your data systems so that, for example, you can CDC tables from your Postgres DB, transform them in interesting ways, and then materialize the result in a view within Elastic or DynamoDB that updates continuously and with millisecond latency.
It will even propagate your sourced SQL schemas into JSON schemas, and from there to, say, equivalent Elastic Search schema.
It's been in preview forever though, not sure when it's going to officially launch.
The amount of complexity to guarantee data integrity while covering all possible use cases will be just unmanageable.
I'd be extremely happy to be proven wrong, though...
Almost every single team at Amazon that I can think of off the top of my head uses DynamoDB (or DDB + S3) as its sole data store. I know that there are teams out there using relational DBs as well (especially in analytics), but in my day-to-day working with a constantly changing variety of teams that run customer-facing apps, I haven't seen RDS/Redis/etc being used in months.
Which says nothing of DDB. It's an god-tier tool if what you need matches what it's selling. However, I see too many teams reach for it by default without doing any actual analysis (including young me!), thus leading to the "oh shit, how will we...?" soup of ad-hoc supporting infra. Big machines look great on the promo-doc tho. So, I don't expect it to stop.
At one company, someone accidentally set the write rate rate high to transfer data into the db. This had the effect of permanently increasing the shard count to a huge number, basically making the DB useless.
It is a resource that can often be the right tool for the job but you really have to understand what the job is and carefully measure Dynamo up for what you are doing.
It is _easy_ to misunderstand or miss something that would make Dynamo hideously expensive for your use case.
Bulk loading data is the other gotcha I've run into. Had a beautiful use case for steady read performance of a batch dataset that was incredibly economical on Dynamo but the cost/time for loading the dataset into Dynamo was totally prohibitive.
Basically Dynamo is great for constant read/write of very small, randomly distributed documents. Once you are out of thay zone things can hey dicey fast.
I'd say requiring scans or filters as opposed to queries is one of the biggest issues that can bite your pocket.
Think carefully about how you'll access your data later. You won't be able to change it drastically and cheaply later.
If possible, put the json in Workers KV, and access it through Cloudflare Workers. You can also optionally cache reads from Workers KV into Cloudflare's zonal caches.
> To be honest, I'd hoped that it could be a bit more 'magic', like S3
You could opt to use the slightly more expensive DynamoDB On-Demand, or the free DynamoDB Auto-Scaling modes, which are relatively no-config. For a very ready-heavy workload, you'd probably want to add DynamoDB Accelerator (an write-through in-memory cache) in front of your tables. Or, use S3 itself (but a S3 bucket doesn't really like when you load it with a tonne of small files) accelerated by CloudFront (which is what AWS Hyperplane, tech underpinning ALB and NLB, does: https://aws.amazon.com/builders-library/reliability-and-cons...)
S3, much like DynamoDB, is a KV store: https://news.ycombinator.com/item?id=11161667 and https://www.allthingsdistributed.com/2009/03/keeping_your_da...
I’d urge you to start writing a prototype, a lot of your assumptions might get thrown out the window. Dynamo is not necessarily good for reading high volume. You’ll end up needing to use a parallel scan approach which is not fast.
Our entire solution is basically based on top of lambda and dynamodb tables and it works really as long as you don't threat the tables like SQL.
Put it on on-demand pricing (it'll be better and cheaper for you most likely), and it will handle any load you throw at it. Can you get it to throttle? Sure, if you absolutely blast it without ever having had that high of a need before (and it can actually be avoided[0]).
You will need to understand how to model things for the NoSQL paradigm that DynamoDB uses, but that's a question of familiarity and not much else (you didn't magically know SQL either).
My experience comes from scaling DynamoDB in production for several years, handling both massive IoT data ingestion in it as well as the user data as well. We were able to replace all things we thought we would need a relational database for, completely.
My comparison between a traditional RDS setup: - DynamoDB issues? 0. Seriously. Only thing you need to monitor is billing. - RDS? Oh boy, need to provision for peak capacity, need to monitor replica lags, need to monitor the Replicas themselves, constant monitoring and scaling of IOPS, suddenly queries get slow as data increases, worrying about indexes and the data size, and much more...
[0]: https://theburningmonk.com/2019/03/understanding-the-scaling...
Lets say that you operate in an AWS-less environment, with everything bare metal, in a datacenter. Your GOOD infra team has to do the following:
Hardware:
- make sure there is a channel to get new hardware, both for capacity increase and spares. What are you going to do? Buy 1 server and 2 spares? If one of the servers has an issue, isn't it quite likely that the other servers, from the same batch, to have the same issue? Is this affecting you, or not? Where do you store the spares? In a warehouse somewhere, making it harder to deploy? In the rack with the one in use, wasting rackspace/switch space? Are you going to rely on the datacenter to provide you with the hardware? What if you are one of their smaller customers and your requests get pushed back because some larger customer requests get higher priority?
- make sure there is a way to deploy said hardware. You don't want to not be able to deploy a new server because there is no space in the rack, or no space in the switch. Where are your spares? In a warehouse miles away from the datacenter? Do you have access to said warehouse at midnight, on Thanksgiving? Oh shit, someone lost the key to your rack! Oh noes, we don't have any spare network cable/connectors/screws...
Software:
- did you patch your servers? did you patch your switches?
- new server, we need to install the os. And a base set of software, including the agent we use to remote manage the server.
- oh, we also need to run and maintain the management infra, say the control plane for k8.
- oh, we want some read replicas for this db, not only we need the hardware to run the replicas on (and see above for what that means), now you need to add a bunch of monitoring and have plans in place to handle things like: replicas lagging, network links between master and replicas being full, failover for the above, master crapping out yada yada.
I bet there are many other aspects I'm missing.
Choices:
Your GOOD infra team will have to decide things like: how many spares do we need, is the capacity we have atm enough for the launch of our next world-changing feature that half the internet wants to use? Are we lucky enough to survive a few months without spares or should we get estra capacity in another datacenter? Do we want to have replicas on the west coast or is the latency acceptable?
These are the main areas of what an infra team is supposed to do: Hardware, Software and Choices. AWS (and most other cloud providers) is making the first 2 points non issues. For the last area you can do 2 things: get an infra team (could be a full fledged team, could be 1 person, you could do it) and teoretically you will get choices tailored to what your business needs OR let AWS do it for you. *AWS might make these choices based on a metric you disagree with and this is the main reason people complain*.
Happy to talk more. We're actively moving a bunch of workloads away from DynamoDB and to Aurora so this is fresh on our minds.
But data at scale is about:
1) knowing your queries ahead of time (since you've presumably reached the limit of PG/maybesql/o-rackle.
2) dealing with CAP at the application level: distributed transactions, eventual consistency, network partitions.
3) dealing with a lot more operational complexity, not less.
So if the snake oil salesmen say it will be seamless, they are very very very much lying. Either that, or you are paying a LOT of money for other people to do the hard work.
Which is what happens with managing your own NoSQL vs DynamoDB. You'll pay through the roof for DynamoDB at true big data scales.
It's not, if you plan it right. Learn about single table design for DynamoDB before you start. There are a lot of good resources from Amazon and the community.
Here is a very accessible video from the community:
https://www.youtube.com/watch?v=BnDKD_Zv0og
Here is a video from Rick Houlihan, a senior leader from AWS who basically helps companies convert to single table design:
https://www.youtube.com/watch?v=KYy8X8t4MB8
And a good book on the topic:
If you use single table design, you can turn on all of the auto-tuning features of DynamoDB and they will work as expected and get better and more efficient with more data.
Some people worry that this breaks the cardinal rule of microservices: One database per service. But the actual rule is never have one service directly access the data of another, always use the API. So as long as your services use different keyspaces and never access each other's data, it can still work (but does require extra discipline).
Yes, you have to learn about all these things upfront. But once you figure it out, test it, and configure it - it will work as you expect. No surprises.
Whereas Relational Databases work until they don't. A developer makes a tiny (even a no-op) change to a query or stored procedure, a different SQL plan gets chosen, and suddenly your performance/latency dramatically reduces, and you have no easy way to roll it back through source control/deployment pipelines. You have to page a DBA who has to go pull up the hood.
With services like DDB, you maintain control.
The partitioning scheme came off as confusing and opaque but I think that says more about Amazon's documentation than the scheme itself.
I do not like that there's no really third party tooling integration to be able to query. Their UI in the console is _so freaking terrible_ yet you have no other way than code to query it. This problem is so bad that I will avoid using it where I can despite it being a good option, performance-wise.
- DynamoDB Workbench (free, AWS official): https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
- Dynobase (paid, third-party): https://dynobase.dev/
"Deploying DynamoDB Locally on Your Computer" https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
But i like the python boto3 library https://boto3.amazonaws.com/v1/documentation/api/latest/guid...
Build yourself a few wrappers to make querying more convenient, and i query straight from a python repl pretty effectively
* No way to "DELETE * FROM T"
Dynamo was a key factor to us when we were releasing the MVP of our News API [0]. We used Dynamo, ElasticSearch, Lambda and could make it running in 60 days while being full-time employed.
Also, the best tech talk I saw was given by Rick Houlihan on re:Invent [1]
I highly recommend every engineer to watch it: it's a great overview of SQL vs NoSQL
[0] https://newscatcherapi.com/blog/how-we-built-a-news-api-beta...
https://twitter.com/houlihan_rick/status/1472969503575265283
On that thread he criticizes AWS regarding DynamoDB openly.
> I will always love DynamoDB, but the fact is it is losing ground fast because AWS focuses most of their resources on the half baked #builtfornopurpose database strategy. I always hated that idea, I just bit my tongue instead of saying it.
> The problem is the other half-baked database services that all compete for the same business. DocumentDB, Keyspaces, Timestream, Neptune, etc. Databases take decades to optimize, the idea that you can pump them out like web apps is silly.
> I was very tired of explaining over and over again that DynamoDB is actually not the dumbed down Key-Value store that the marketing message implied. When AWS created 6 different NoSQL databases they had to make up reasons for each one and the messaging makes no sense.
> No one uses DynamoDB alone: they bolt it onto Postgres after realizing they have availability or scale needs beyond what a relational database can do, then they bolt on Elasticsearch to enable querying, and then they bolt on Redis to make the disjointed backend feel fast. And I'm just talking operational use cases; ignoring analytics here.
Perhaps MongoDB is prime for a comeback?
I've seen it used in many places over the years.
Today I would choose JSON in Postgres before I would just jump to Monogo but it certainly serves a purpose for many shops and it is still widely used AFAIK.
I _really_ miss RethinkDB.
Hat tip for compass, very nice tool I was just losing last week.
We use Atlas and it "just works" so no comment on administering mongo vs rethink haha.
If Rick's vouching for it maybe it's time to give it a try. It must be pretty mature by now.
Lots of good info in the response in this SO post
https://stackoverflow.com/questions/10560834/to-what-extent-...
And Doug Terry talking through the details of how DynamoDB's transaction protocol works: https://www.usenix.org/conference/fast19/presentation/terry
If we did publish more about the internals of DDB, what would you be looking to learn? Architecture? Operational experience? Developer experience? There's a lot of material we could share, and it's useful to hear where people would like us to focus.
1. DynamoDB is not as convenient. There are a bit too many dials to turn.
2. DynamoDB does not have a SQL facade on top.
3. DynamoDB is proprietary, I believe there's no OSS API equivalent if you want to migrate out.
4. DynamoDB was kind of expensive. But it has been a while since I last check the pricing page.
It's simply much better to start with PostgreSQL Aurora and move to a more scalable storage based on specific uses-cases later. For example: Cassandra, Elastic, Druid, or CockroachDB.
I've been in a couple startups that went Dynamo first and development velocity was a pale shadow of velocity with Postgres. When one of those startups dumped dynamo for Postgres velocity multiplied immediately. I'd estimate we were moving at around 1000% and the complete transition took less time than even I expected (about a month). Once the business matures, moving tables onto dynamo and wrapping them in a microservice makes a lot of sense. Dynamo does solve a lot of problems that become increasingly material as the business evolves.
Eventually, SQL's presence declines and transitions into an analytics system as narrower, but easier for ops, options proliferate.
I think this is out of date https://aws.amazon.com/about-aws/whats-new/2020/11/you-now-c...
And then on-demand provisioning was released and it was cheap enough to be worth simplifying our workflows.
I just looked him up (I had not heard of him before seeing his name mentioned on r/aws a few days ago) and he was an L7 TPM/Practice Manager in AWS's sales organization. That's not really a notably high position, and in the grand scheme of Amazon pay scales, isn't that high up. An L7 TPM gets paid about the same as, or sometimes less than, an L6 software dev (L6 is "senior", which is ~5-10 years of experience).
Also, him being in the sales org means he had practically nothing to do with the engineering of the service. AWS Sales is a revolving door of people. I mean no offense towards Rick (again, I didn't know him or even know of him before I read his name in a comment a few days ago), but I would not read anything at all into the fact that an L7 Sales TPM left for another company.
unless he posts here about it we can't really know -- we can only speculate but I think he had a higher amount of influence than his title/rank might suggest. I think Rick's influence with respect to DynamoDB is akin to that of Kelsey Hightower's influence over k8s at Google.
AWS re:Invent 2018: Amazon DynamoDB Deep Dive: Advanced Design Patterns for DynamoDB (DAT401) https://youtu.be/HaEPXoXVf2k
AWS re:Invent 2019: [REPEAT 1] Amazon DynamoDB deep dive: Advanced design patterns (DAT403-R1) https://youtu.be/6yqfmXiZTlM
AWS re:Invent 2020: Amazon DynamoDB advanced design patterns – Part 1 https://youtu.be/MF9a1UNOAQo
AWS re:Invent 2020: Amazon DynamoDB advanced design patterns – Part 2 https://youtu.be/_KNrRdWD25M
AWS re:Invent 2021 - DynamoDB deep dive: Advanced design patterns https://youtu.be/xfxBhvGpoa0
Amazon DynamoDB | Office Hours with Rick Houlihan: Breaking down the design process for NoSQL applications https://www.twitch.tv/videos/761425806
This person might be responsible for the majority of evangelism and revenue for the company. Do you expect the SDEs to know about him?
Again, no shot against against Rick - he is amazing, smart, technical, competent, and a deep owner.
But the average SDE on the team won't know about these or watch these talks. There are too many deep internal engineering challenges to solve.
Hell, what do you think re:Invent is? It's a sales conference.
In any company you have two groups of people: Those that build the product, and those that sell it. Ultimately, solutions architects and developer advocates are there to help sell the product.
Of course Amazon is customer obsessed. And genuinely interested in ensuring customers have a good experience, and their technical needs are met - through education, support, and architectural guidance. But ultimately, that's what it is.
Rich was in the sales org. His primary job was sales. Reinvent is a sales conference. Speaking at reinvent is a sales pitch. He was a salesperson. I'm not sure why you're so offended by that. Being a salesperson isn't bad, it's just an explanation for why engineers wouldn't have heard of him.
DDB is a steady ship. The explanation on https://news.ycombinator.com/item?id=30009611 is likely the best explanation. L7 TPMs make the same money as L6 SDEs.
Getting promoted to L8 - director - is a monumental effort and likely seemed much harder than pursuing a comprable position at MongoDB.
Good for him for doing it, and for making Amazon take a long hard look at every way they failed in not keeping him.
How? Rick wasn't part of the DynamoDB service team. He wasn't an engineer, nor a manager on the team, nor even a product manager. He was a salesperson that specialized in DDB. He most likely had very little interactions, if any, with the engineering team. I don't see how him leaving speaks at all to anything about the inner workings of the engineering teams.
Rick seems cool, and after skimming some of his chats he seems really knowledgeable about the customer-facing side of DDB, and I mean absolutely no disrespect to him. But I think you're making way too many assumptions about his "rank" and "influence" within the company.
>At the same time you are able to this internal lookups?
I looked him up on LinkedIn. Nothing internal about it.
Compensation was a minor issue. I was an org chart aberration already and AWS pulled out all the stops to retain me. I will always appreciate the opportunity that AWS provided me and my time at DynamoDB will always hold a special place in my heart. I really do believe that MongoDB is poised to do great things and my decision had more to do with being a part of that than anything else.
Once you understand how to properly use dynamo and what it’s good at it you get so much power; at a fraction of the cost depending on your workload.
Was a breeze to setup multi region applications that utilize the single table design strategy.
If you ever need complex queries just use dynamo streams to power whatever search solution you are comfortable with.
Presumably create more value. I know it's a marketing post but still - storing and retrieving things is pretty valuable in and of itself.
We really struggled with implementing adhoc queries/search. For e.g:- select * from employees where name = X and city = Y.
Any improvements in DynamoDB that make it easier to implement such queries?
Data should be stored in the fashion you wish for it to be read, and storing the same data in more than one configuration is acceptable.
Good resource: https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
Dynamo is primarily designed for high volume storage/querying on well understood data sets with a few query patterns. If you want to be able to query information on employees based on their name and city you'll need to build another index keyed on name and city (in practice Dynamo makes that reasonably simple by adding a secondary index).
This is often easier said than done, but it can be far less expensive and more performant than adding an index for each search.
If what you are trying to do looks more like "Give me all the customers that live in Cuba and have spent more than $10 and have green eyes", Dynamo isn't for you. You can query that way but after you put all the work in to get it up and running, you'd probably be better off with Postgres.
https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
I haven't used DynamoDB in a couple of years, so I'd be curious to know how querying compares if anyone can share some light that has used both Cosmos and Dynamo recently.
If you want to search by parameters that aren't keys then you need to store your data that way. Most of these systems have secondary indexes now, and that's basically what they do for you automatically in the backend, storing another copy of your records using a different key.
If you need adhoc relational queries then you should use a relational database.
Not that I recommend it, but by using space-filling curves, one could to index multiple dimensions onto DynamoDB's bi-dimensional (hash-key, range-key) primary-index: https://aws.amazon.com/blogs/database/z-order-indexing-for-m... and https://web.archive.org/web/20220120151929/https://citeseerx...
However, so long as you add a global secondary index (GSI) with name, city as the key, you can certainly do such things. But be aware for large-scale solutions:
1. There's a limit of 20 GSIs per table. You can increase with a call to AWS support.
2. GSIs are latently updated; read-after write is not guaranteed, and there is no "consistent read" option on a GSI like there is with tables.
3. WCUs on GSIs should match (or surpass) the WCUs on the original table, else throughput limit exceeded exceptions will occur. So, 3 GSIs on a table means you pay 4x+ in WCU costs.
4. The keys of the GSI should be evenly distributed, just like the PK on a main table. If not, there is additional opportunity for hot partitions on write.
Ref: https://aws.amazon.com/premiumsupport/knowledge-center/dynam...