With only 4 years of professional experience I have never worked on a MongoDB project in which MongoDB somehow wasn't an issue. The proposed solutions merely being "improve the indices" or "scale the cluster", often without clearly defining what's going on.
So granted I have a very limited experience, it's baffling to me mainly in comparison to the PostgreSQL projects I worked on. They definitely also had problems but these problems were clearly defined even if resolution wasn't quick or easy. It was usually an out-of-date or generally messy schema that was causing issues and folks usually were able to clearly define the schema problems.
Why is DynamoDB the bee's knees but MongoDB is a thing to be despised?
2. DynamoDB is pushed in slightly more balanced ways than Mongo was, at the start Mongo was supposed to be the second coming of the (database) Messiah.
3. I don't believe DynamoDB has defaults that lose your data.
There's limits on how big your data can be, which is a little annoying if you want to use it to store a couple larger things alongside all your small data.
The C# API is absolutely painful IMO, and makes it easy for new developers to use it in a way that you would have been better off grabbing a different technology.
I suppose one could argue S3 would be the better choice for large payloads, OTOH I'm a fan of minimizing potential points of failure so if we already are using DynamoDB I'd rather not toss in an additional S3 integration that could break or need maintenance later.
DynamoDB is a managed service that scales far beyond Mongo and truthfully far beyond what any of us here would be seriously discussing.
With regard to complex schemas and related data, both DynamoDB and Mongo are harder to deal with once you've past a trivial size. If all you need is basic CRUD with limited or no joins, Mongo could suffice to a large deployment, but if you're mostly accessing by primary key, DynamoDB will soundly eat its lunch performance-wise and for cheaper.
If you need to join related data (or just find them convenient), relational models work famously well, and usually up the point where you have 10 million simultaneous users.
Mongo more and more is being relegated to niches where the problem just happens to exactly fit Mongo's feature list, which unfortunately for Mongo encompasses ever-narrowing gaps in the other options' offerings.
Mongo can't match the scale or economy of DynamoDB and can't match the flexibility of relational. That's why fewer and fewer treat it as their go-to database nowadays.
https://db-engines.com/en/ranking_trend/system/MongoDB
https://www.macrotrends.net/stocks/charts/MDB/mongodb/revenu...
https://seekingalpha.com/article/4473768-mongodb-mdb-stock-s...
2. DynamoDB is a 100% managed service. No instances to wrangle. No manual partitioning. Just define your table's name, the table's partition key, and an optional sort key, and you're off to the races! It just works.
If you have a serverless environment like lambdas with API Gateway or AppSync, it can scale pretty much to the extent of your business model rather than some fixed limit.
BUT for data schemas beyond the most trivial, it can easily be more complex to deal with than a relational database. Whereas a relational database usually aims for normalization where no data is duplicated and foreign keys keep things straight, DynamoDB works best with a denormalized data set. No joins. Ever. Schema integrity is your problem, not the database's. In the deal though, you get a database engine that can scale effectively infinitely.
In other words, storage is cheap, but access is expensive, so data duplication is pretty much encouraged in DynamoDB for the sake of speed. You aim for getting everything you need in a single entry or sequential row iteration.
When we use a relational database and run into problems, we run EXPLAIN and EXPLAIN ANALYZE to figure out the query plan, so we can optimize. SQL is a 4th generation, declarative language that describes WHAT data you want, not HOW you get it.
DynamoDB in the larger sense is at the level of EXPLAIN output. It is 100% HOW to get data. The WHAT is at application level and implemented by you in code.
If EXPLAIN output makes no sense to you, then DynamoDB probably isn't for you either unless it's a trivial app/data set.
But then again, if it's a trivial app/data set, literally anything can work. An O(n!) algorithm is perfectly reasonable given a small/simple enough data corpus and large enough computing resources. It's when the data set gets slightly larger that decisions become important.
But when things are very small/simple, it's often hard to argue with fast+free. Those are the sweet spots for DynamoDB: very small/simple and the mind-bogglingly humongous. For everything in the middle, relational databases work wonderfully and are much easier to work with, especially for non-trivial data sets.
I've learned this by doing it the opposite way for years. however, website schemas can frequently change, which ruins the database schema if you're immediately parsing into a normalized structure. In other words, it's brittle.
noSQL (or JSON fields) allow one to store unstructured data which is more forgiving when a schema changes (I.e. change the spelling of a dictionary key, etc.).
Then yes, you'd have to parse it yourself, but you'd be doing that anyway with JSON fields, more or less.
I thought I was being such a good little programmer by normalizing my data as soon as possible. All it did was provide me job security (at the cost of headaches) every time the schema changed.
The problem is when you need to update your schema to support data/relations not available in your own schema.
The problem with SQL is that schemas are so sticky and hard to change. People are always hesitant to change their schemas because schema changes themselves are difficult and then you have to update a bunch of backend code, and then your frontend code maybe does strange denormalized things, and then everything breaks.
I've been thinking about this for a while, and I think what is needed is a visual tool that can show all dependencies of a data schema element (backend, frontend), so that schema changes are easier to make.
All the layers (db access, api, cache, frontend data stores/caches/frameworks) in modern architectures make it virtually impossible to modify a schema without causing chaos. The solution is keeping the data model and query interface as close as possible throughout the entire stack. For example you should never write any manual data manipulation code (e.e. `people.map(p => p.full_name = p.firstName + p.lastName)` unless this is able to be traced through the entire system. A monorepo, typed ORM, refactoring tooling (e.g. IDE) can help, but its usually never setup well enough or integrates close enough with the db.
IMO, this is the main reason we have so many database technologies: they make important performance tradeoffs.
Most of these high performance use cases actually fit this model well (well, the schema management being awesome is the wildly varying part).
But for everything else, there's Mastercard and RDBMS.
For enterprise stuff with relatively low traffic and high amounts of complex ad-hoc queries, RDBMS is without a doubt the best choice. If you have high traffic web services with strict availability and latency requirements, I would seriously consider avoiding RDBMS's as they tend to be difficult to operate and scale with those requirements, and they let you easily do terrible things in that context (e.g. locking behavior, txid exhaustion, etc.).
We got queries down from 30+ seconds to 5 ms simply by properly indexing, defragmentation, analyzing query plans, SQL Stored Procedures, etc.
I see a lot of complaints from developers claiming they have to join a "million row table to million row table" and reports are slow, and this gets blamed on the DB. These should not be slow outside of how much bandwidth is being pushed over the wire, which is often exactly what the problem is. They just didn't see it.
It is similar IMO to the diversity in programming languages. They are all Turing complete, at the end of the day. The difference is in the patterns they encourage and the constraints they impose.
If anything, devs probably just need to be more familiar with DB internals and have tools setup to analyze queries easier.
Impressive!
Columnar stores and graph DBs have always felt more like "real" tech, something promoted for specific use cases.
Key-value stores I'm not even going to be denigrating at all since we've had them since forever and they do a great job at their tasks (BerkeleyDB, Memcache, Redis, etc.). They don't deserve being included into fads or niches, they're a very valuable universal resource.
At least that's how I see it.
But the term has shifted underneath me. NoSQL means document database.
And - in practice - it means using a document database as the primary store when it's not really the right choice (been there!)