Bye Bye Mongo, Hello Postgres
theguardian.com
theguardian.com
Great read. Well done.
> Digital Blog
> A blog by the Guardian's internal Digital team. We build the Guardian website, mobile apps, Editorial tools, revenue products, support our infrastructure and manage all things tech around the Guardian
Guess it makes sense to reuse the platform that already has the templates than use another platform and reimplement the design.
> Due to editorial requirements, we needed to run the database cluster and OpsManager on our own infrastructure in AWS rather than using Mongo’s managed database offering.
How is running on AWS different than Guardian Cloud in their basement?
Did reports of the Snowden revelations reside on the CMS?
After all, the client computer that connects to the CMS is just as, or more likely to be compromised. I wouldn't be surprised if the coverage (or at least parts of it) were edited on airgapped laptops.
If those were the only two choices, you might be right. But the resources needed for the actual CMS functionality sound modest enough to run independently of the main website.
> the client computer that connects to the CMS is just as, or more likely to be compromised
That's faulty reasoning.
Why? It's an obvious potential point of compromise.
Disclaimer: The voice is my head does not come out of nowhere. I am building a product which addresses this: https://github.com/wallix/datapeps-sdk-js is a API/SDK solution for E2EE. Sample app integration is available at: https://github.com/wallix/notes (you can switch between master and datapeps branches to see the changes of the E2EE integration)
Our software E2EE solution has advantages over HSM though: Cost obviously, and more features and extensibility.
Can I ask what was the total cost of the migration?
If there was software that could do this database migration without downtime, how much would you/Guardian be willing to pay?
We accured lots of downtime due to mongo.
But later versions were rock solid and I've matainer mongo installations at many startup and SMEs once you setup alertd for disk/memory usage, off you go. Works like charm 99% of the times.
Not that that's really a surprise or was unknown, it's just fairly new to see in the open source ecosystem instead of the enterprise one.
SQL has decades of production maturation, and has wider domain knowledge.
Unfortunately the number of my customers who would sign off on just "two nines" is approximately zero...
Part of my duties at work require me to deal with "large" issues. While a solution to them is usually necessary and quick and high quality, I've seen the analyses that come after them vary in quality.
Good writeups tend to stick around in people's memories and become company culture, and drive everyone to do better. Bad writeups are forgotten, and thus the lessons learned from them are forgotten as well.
This particular article stands out for me. English is not my first language, and I've spent most of my life dealing with very fundamental technical details, so most of my writeups aren't the best. I'm going to bookmark this one and come back to it to learn how to write accessible technical narratives.
I don't even.
Well done and congratulations to everyone on the team.
In most early stage startups, that would be an unacceptable loss of time.
So I don't judge them for doing a one-shot migration even if it causes an hour of downtime.
It all depends on the business.
Another team at the guardian did a similar migration but went for a 'bit by bit' approach - so migrating a few bits of the API at a time - which worked out faster, in part because stuff was tested in production more quickly, rather than our approach with the proxy which, whilst imitating production traffic, didn't actually serve Postgres data to the users until 'the big switch' - so not really a continuous delivery migration!
I like the article but it was a bit hard for me to consume with multiple voices in different parts.
https://www.mongodb.com/customers/guardian
https://www.mongodb.com/presentations/mongodb-guardian
https://www.slideshare.net/tackers/why-we-chose-mongodb-for-...
And reupping my previous, three-part series on MongoDB:
On MongoDB
NoSQL databases were the future. MongoDB was the database for "modern" web engineers and used by countless startups. What happened?
It's interesting because I feel like the NoSQL world spent a lot of time reinventing the RDBMS from the "opposite direction". For all its faults, SQL is a fascinating language because it mostly ignores low level details of how the database scales, how it operates under the hood. It's not that SQL is intrinsically hard to scale (certainly at the relational algebra roots it shouldn't be hard, in theory), but it certainly leaves a lot of work to anyone building a database engine to figure out what/when/why/how to scale. I feel like a lot of RDBMS' query analyzers/planners resemble things like HBase a lot more than folks realize.
It's great that NoSQL realized that sometimes those "low level details" in an SQL engine are useful in their own ways, and have increased the spectrum of performance versus power/flexibility trade-off options. But it shouldn't be that big of a surprise that SQL databases remain competitive in that trade-off space, given under the hood they've had to think about a lot of that stuff over many decades.
For code, I think by now we all understand that you should always start with clean, well-factored code, and then optimize only as much as is necessary, which is usually not at all, and always under profiler guidance. It's the same with DBs: You start with a clean, well-normalized schema, and then de-normalize only as much as is necessary, which is usually not at all, and always under profiler guidance.
Also, keep in mind that improvements in compiler technology over time mean that the performance tricks of old can be useless or even actively harmful nowadays. This is true of SQL every bit as much as C.
It is true that for some narrow class of analytical workloads 20-25 years ago (behold BW of 199x) the denormalized case performance was better compare with straight non-optimized running of the same queries over normalized schema. Since then, the exponential availability of RAM and huge increase in streaming speed of HDD with stagnating IOPS (the main mistake in analytical workloads on HDD in the last 10-15 years - using nested loop join with indexed lookup into the large facts tables :) have made denormalization obsolete and harmful. If anything, the emergence of SSD and the huge RAMs moved things even further toward and beyond normalization, by making the "super-normalization", i.e. columnar tables, a viable everyday thing.
The idea that joins are slow is a holdover from the bad old days when everyone used MySQL and MySQL sucked at joins. On a more robust DBMS, a normalized schema will often yield better performance than the denormalized one. Less time will be lost to disk I/O thanks to the more compact format, the working set will be smaller so the DBMS can make more efficient use of its cache, and the optimizer will have more degrees of freedom to work with when trying to work out the most efficient way to execute the query.
(edited to add: If you're having performance problems with a normalized schema, the first place to look isn't denormalization, it's the indexing strategy. And also making sure the queries aren't doing anything to prevent proper index usage. Use the Index, Luke! is a great, DB-agnostic resource for understanding how to go about this stuff: https://use-the-index-luke.com )
JOINS are fast and it all comes down to how much data you're moving. If it's a large table joining to a small set of values, then the joined data is quickly retrieved and applied to the bigger table, with great performance.
If the join is between two large tables where every combination is unique then that's the unique case where the joined table is just adding another hop to get the full set of data for each row and is a perfect candidate for denormalization, although in that case it probably should've been a single table to begin with. Of course there's a spectrum between these two scenarios but it takes a lot before denormalization makes sense on any modern RDBMS.
1. You start from a normalized schema.
2. You denormalize in a structured way (eg dimensional modelling), rather than any old how.
3. You test the change.
Database query planners work better when they can take logically-safe shortcuts in their work. In large part that comes down to a properly-constructed schema.
Denormalizing makes it harder for the query planner. It also means you will probably lose out on future query planner enhancements.
Generally speaking, if someone wants to denormalize, I want to know the actual business value created and that the business risk is properly understood.
Which is why MongoDB is awesome for teaching and building new stuff but horrible for reports, metrics, and scaling.
“The computer industry is the only industry that is more fashion-driven than women’s fashion.” — Larry Ellison
I suspect that in a year I’ll be using what is effectively an Elisp machine with an Apple logo paired with an iPhone.
It's been working for me for quite a while now. And it's depressing, but here we are.
Of course it's entirely possible to screw this up, modern phone-centric design standards are really bad about it for example, but in general you need to consult the manual (or Google) far less often.
Early GUI developers also put effort into making their GUIs at least somewhat intuitive, but they also bundled thick books of documentation on how to use their systems.
To a degree, they are self-documenting, but not because of their GUIness so much as their design choices; menus helped a lot with this. Menus actually are a good subject to touch upon, they are definitely one of the better conventions we developed, and they were based upon and analogous to a restaurant menu. However their very nature as a list of commands you can issue a GUI does also typically but not always, limit the application. If the only options available to you are what is on the menu and there is no other interface, then an application is much more limited. This isn’t even true of most restaurants which will often allow you to issue an order for something not on the menu if they have the ingredients, equipment and expertise to make it.
But as a convention, it is not limited solely to GUIs, you can incorporate menus into any interactive interface.
I think the real innovation wasn’t GUIs, it was interactive software. The innovation beyond that is scriptable software.
What is important isn’t GUIness, but the developer’s intent. If you develop software with the intent to be self-documenting and discoverable, you will end up with an interface that is both of these things provided you did a competent job of it. A GUI might help, relying on platform conventions might help, and using common cultural conventions might help, but these aren’t the necessary ingredients for those qualities.
Emacs has the quality of self-documentation, but it is emphatically not a GUI even though it is interactive.
I would say well built systems, provided you intend to do anything productive and even slightly complex, should have a manual included, or else it isn’t a well built system.
Literal billions of users.
The real driver of the NoSQL movement, I believe, was that everybody wanted to be the next big social network or content aggregation site. Everybody wanted to be the next Facebook, Instagram, Twitter, etc. and that's what people were trying to build. Ginormous sites like these are are one of the applications that strongly favors availability/eventual consistency over guaranteed consistency, whereas most other applications are quite the opposite.
Nobody really cares if your Instagram post shows up 10 minutes later in New York than it does in LA, and certainly not if the comments appear in similarly inconsistent order. It's one step above best-effort delivery. However, your bank, hospital, etc. often care quite a bit that their systems always represent reality as of right now and not as of half an hour ago because there's a network problem in Wichita.
The question is, "If my data store isn't sure about the answer it has, what should it do?" RDBMS says, "Error." NoSQL says, "Meh, just return what you have."
Even that's too simplistic. For most RDBMSes, the answer depends on how you have it configured, and usually isn't "Error". If you're using a serializable transaction isolation level, it usually means, "you might have to wait an extra few milliseconds for your answer, but we'll make sure we get you a good one." Other isolation levels allow varying levels of dirty reads and race conditions, but typically won't flat out fail the query. This is probably the situation most people are working under, since, in the name of performance, very few RDBMSes' default configurations give you full ACID guarantees.
To the "DB in NY knows something different from DB in LA" example, there are RDBMSes such as the nicer versions of MSSQL that allow you to have a geographically distributed database with eventual consistency among nodes. They're admittedly quite expensive, but, given some of the stories I've heard about trying to use many NoSQL offerings at that kind of scale, I wouldn't be surprised if they're still cheaper if you're looking at TCO instead of sticker price.
Some array magic and jsonb_agg and suddenly you can get easy to json decode results instead of having to play with rows in your app. Yes you can also do it with xml_agg but these days people tend to consider anything xml as evil (they're wrong).
Many ATMs will still give you money when they're offline, and things become eventually consistent by comparing the ledger.
Shops also generally want to take orders and payments irregardless of the network availability, so whilst they might generally act as CP systems, they'll be AP in the event of network downtime, but will likely lose access to fraud checks, so may put limits of purchases, etc.
They're probably all CP locally and AP (or switchable from CP) from an entire system perspective.
Speed to market/first version using JSON stores is attractive, especially when you're still prototyping your product and won't have an idea of exact data structures until there's been some real world usage.
One of the biggest and most expensive was using Cassandra to store membership details. Something like 4 years of work, by a team of 40, wasted by stupid decisions.
They included: o Using Cassandra to store 6 million rows of highly structured, mostly readonly data
o hosting it on real tin, expensive tin, in multiple continents (looking at >million quid in costs)
o writing the stack in java, getting bored, re-writing it as a micro service, before actually finishing the original system
o getting bored of writing micro services in java, switching to scala, which only 2/15 devs knew.
o writing mission critical services in elixir, of which only 1 dev knew.
o refusing to use other teams tools
o refusing to use the company wiki, opting for thier own confluence instance, which barred access to anyone else, including the teams they supportedJumping right into writing "mission critical services" in the brand new language that few people at the company know well is asking for trouble.
there were no problems of scale, speed or latency. It was migrating from one terrible system to something that should be smaller, simpler, cheaper and easier to run.
The API is/was supposed to do precisely four things:
o provide an authentication flow for customers
0 provide an authentication flow for enterprises to allow SSO
o handle payment info
o count the number of articles read for billing/freemium/premium
That is all. Its a solved problem. Don't innovate, do.Spend the time instead innovating on the bits that are useful to the business and provide an edge: CRM, pattern recognition and data analytics.
Step one was hire everyone the guy had worked with at his previous company. They were all winforms / excel / SQL / sharepoint / office developers from big finance and had no idea where to go really. None of them had even touched asp.net.
Cue "what's popular". Well that was Ruby on Rails back then on top of MySQL and Linux. 4 people with zero experience pulled this stack in and basically wrote winforms on top of it. Page hits were 5-8 seconds each. Infrastructure was owned by SSH worms and they hadn't even noticed.
I think I lasted two days there before I said "I'm done".
At the time, under the influence of React, the idea was to "build web application sorely based on Functional Programming". Since after years of trying no one could figure out what that meant, the company ditched the CEO and ended up wasting a couple years of work.
Not ephemeral Cloud Stuff rented from some company.
Also, I think MongoDB tried to be everything and failed to be good at anything. It offers neither stellar performance nor scalability and I guess for most projects there is not much advantage over regular SQL database. Certainly nothing to fight over when there is much more technology choices to make.
Another example is "Agile". Everybody is doing it yet I still wait to see a single company that understands what the term means. My current boss is big promoter of "Agile" which in his language is synonym to "Scrum". Yet when asked he has never heard of Pheonix Project, The Goal, theory of constrains or basically any theory at all. So what the people are doing is fighting fires almost 100% of the time with not much project work done for the effort and absolutely no improvements. Yet, because everybody complies to do daily standups and Jira updates we are 100% agile.
Don't even get me started on devops...
IMO, the reason is that newer developers faced the choice of learning SQL or learning to use something with a Javascript API. MongoDB was the natural choice because they excelled at being accessible to devs who were already familiar with Javascript and JSON.
Not only that, their marketing/outreach efforts were also aimed at younger developers. When was the last time you saw a Postgres rep at a college tech event?
> 10gen's key contributions to databases — and to our industry — was their laser focus on four critical things: onboarding, usability, libraries and support. For startup teams, these were important factors in choosing MongoDB — and a key reason for its powerful word of mouth.
Startup Engineers and Our Mistakes with MongoDB
https://www.nemil.com/mongo/2.html
The Marketing Behind MongoDB
Is a Postgres rep a thing?
PostgreSQL looks scary because it is a swiss army knife. It has a million different features and data structures. MongoDB does only one thing.
It's not that learning SQL is hard. It's that people are inherently lazy. "Learn another thing on top of the thing it already took me a couple of years to learn? No thanks."
You seem like the kind of person ready and willing to learn the right tool for the job. From my experience a few years ago on an accredit computing course that covered database admin and programming, this attitude is not representative of most of the software engineering students //unless// there's a specific assignment that requires particular knowledge.
Cs get degrees. And for plenty of developers out there, knowing one language (not even particularly well) gets jobs.
IOW, it's not just laziness, it's a kind of professional conservatism. which is partly what gets older engineers stuck in a particular mindset, but it's also a very effective learned skill. The opposite is being a magpie developer, which results in things like MongoDB taking off :)
You have to do the exact same things with Mongo+JS (e.g. learning when to avoid the JS bits like the plague).
SQL is a skill that rewards investment in it 1000x over, in terms of longevity. It has spanned people’s entire careers! What’s the shelf life of the latest JS framework, 18 months at most...
And that's a big fat mistake. There are so many ways to shoot yourself in the foot with mongo such that simply knowing the language mongo uses for most of its queries while not actually knowing the particulars of how mongo uses that language… well that's just a road to a world of hurt.
For example, when I first inherited a mongo deployment I noticed the queries were painfully slow. Ah hah says me, let's index some shit. Guess what? Creating an index on a running system with that version of mongo = segfault.
After a bunch of hair pulling I got mongo up and running and got the data indexed. But the map reduce job was STILL running so slowly that we couldn't ingest data from a few tens of sensors in real time. So I made sure to set up queues locally on the sensors to buy myself some time.
Even in my little test environment with nothing else hitting the mongo server, mongod was still completely unable to run its map reduce nonsense in a performant manner. Mongo wisdom was: shard it! wait for our magical aggregation framework! Here's the thing: working at a dinky startup we can't afford to throw hardware at it especially that early in the game. Sharding the damn thing would also bring in mongo's inflexible and somewhat magical and unreliable sharding doohickey.
So I thought back to previous experience with time series data. BTDT with MySQL, you're just trading one awful lock (javascript vm) for another (auto increment). So I set up a test rig with postgres. Bam. I was able to ingest the data around 18x faster.
And that's the thing. Mongo appeals to people who are comfortable with javascript and resistant to learning domain specific knowledge. All that appealing javascript goodness comes with a gigantic cost. If you're blindly following the path of least resistance you're in for a bad time.
P.S. plv8 is a thing, and you can script postgres in javascript if you really wanted to.
Document-based storage definitely fits the generalised use-case better than tabular storage.
In practice people who fall into complexity traps are usually asking a lot more of their database engine than any beginner. It's usually not that hard to figure out the approximate cost of a particular query.
Or you have a simple fast query with a lovely plan until the database engine decides that because you now have 75 records in the middle table instead of 74, the indexes are suddenly made of lava and now you're table-scanning the big tables and your plan looks like an eldritch horror.
[Not looking at MySQL in particular, oh no.]
Which brings us back to the original point: data is hard.
If you want to work with databases a domain specific language like SQL really provides a lot of value in solving these hard data problems.
Here's a good article about that: https://blog.jooq.org/2016/07/05/say-no-to-venn-diagrams-whe...
Mongo and Javascript don't solve that either. In fact you get additional problems by virtue of not being able to do a variety of joins. For extra points, you're going to need to go well beyond javascript with mongo if you want performance. 10gen invented this whole "aggregation framework" to sidestep the performance penalty that javascript brings to the table.
On the other side, the postgresql documentation is second to none. SQL isn't necessarily easy but the postgres documentation gives you an excellent starting point.
The idea is, in relational databases, that the vast majority of the time you shouldn't have to do it. Because you're writing your queries in a higher level (nay, functional) language, the query planner can understand a lot more about what you're trying to do and actually choose algorithms and implementations that are appropriate for the shape and size of your data. And in 6 months time when your tables are ten times the size, it is able to automatically make new decisions.
More explicit forms of expressing queries have no hope of being able to do this and any performance optimization you do is appropriate only for now and this current dataset.
ORMs have existed for decades so developers can use a SQL database just fine without knowing the language. So it's definitely not this.
It's more likely because Mongo is (a) is extremely fast, (b) the easiest database to manage and (c) has a flexible schema which aligns better with dynamic languages which are more popular amongst younger developers.
Mongo literally has no upside vs postgres.
The thing I dislike about this type of comment – although I now notice yours doesn't explicitly say this – is the implication that devs don't like SQL because they're lazy or stupid. Well, sometimes that is probably true! But there are some tasks where you need to build the query dynamically at run time, and for those tasks MongoDB's usual query API, or especially its aggregation pipeline API, are genuinely better than stitching together fragments of SQL in the form of text strings. Injection attacks and inserting commas (but not trailing commas) come to mind as obvious difficulties. For anyone not familiar, just look at how close to being a native Python API pymongo is:
pipeline = [
{"$unwind": "$tags"},
{"$group": {"_id": "$tags", "count": {"$sum": 1}}},
{"$sort": SON([("count", -1), ("_id", -1)])}
]
result_cursor = db.things.aggregate(pipeline)
Of course you could write an SQL query that does this particular job and is probably clearer. But if you need to compose a bunch of operations arbitrarily at runtime then using dicts and lists like this is clearly better.Of course pipelines like this will typically be slow as hell because arbitrary queries, by their nature, cannot take advantage of indices. But sometimes that's OK. We do this in one of our products and it works great.
With JSONB and replication enhancements, Postgres is close to wiping out all of MongoDB's advantages. I would love to see a more native-like API like Mongo's aggregation pipeline, even if it's just a wrapper for composing SQL strings. I think that would finish off the job.
You're using the Pymongo library as an example. Someone can just as easily use SQLAlchemy and not have to worry about those things.
For example, if documents in the JSONB column all look roughly like this:
{
"someArrayField": [
{ "key": "steve", "value": 7 },
{ "key": "bob", "value": 15 },
],
"someOtherField": [ "whatever" ]
}
* Can I count the number of entries in someArrayField, summed across all records?* Can I get the per-record mean of the "value" sub-field, summed across all records?
* Can I filter by records that have a "someArrayField" entry where "key" is "steve" and "value" is at least 10 (the above record should NOT match)?
The jsonb_array_elements function is roughly similar to Mongo’s $unwind pipeline op. It explodes a JSON array into a set of rows. From there it’s pretty simple aggregates to achieve what you’re looking for.
I was evaluating Mongo a couple months back to solve roughly the same problems. Eventually discovered Postgres already had what I was looking for.
> Does it do that?
It was supposed to be clear from the context that this meant:
> Does building queries programmatically with SQLAlchemy do that?
Maybe I'm misreading your comment, but you seem to just be talking about writing queries directly in SQL.
If not, could you give an example/link of how to programmically build a query in SQLAlchemy that dynamically makes use of jsonb_array_elements? It would be hugely useful if I could do that.
SQLAlchemy’s Postgres JSONB type allows subscription, so you can do Model.col[‘arrayfield’]. You can also manually invoke the operator with Model.col.op(‘->’)(‘arrayfield’).
So you should be able to do something like:
func.sum(func.jsonb_array_elements(Model.col.op(‘->’)(‘arrayfield’)).op(‘->’)(‘val’))
(Writing on my mobile without reference, so may not be fully accurate)
# Query all rows in the "users" table, filtering for users whose age is > 18, and selecting their name
"users"
|> where([u], u.age > 18)
|> select([u], u.name)
# Build a dynamic query fragment based on some parameters
dynamic = false
dynamic =
if params["is_public"] do
dynamic([p], p.is_public or ^dynamic)
else
dynamic
end
dynamic =
if params["allow_reviewers"] do
dynamic([p, a], a.reviewer == true or ^dynamic)
else
dynamic
end
from "posts", where: ^dynamic
Across all the different means of interacting with a database I have experience with (from full-fledged ORMs like ActiveRecord, to sprocs in ASP.NET), I've found that it offers the best compromise between providing an ergonomic abstraction over the database, and not hiding all of the nitty-gritty details you need to worry about in order to write performant queries or use database-specific features like triggers or window functions.My main point, though, is that you don't need to reach for NoSQL if all you need is a way to compose queries without string interpolation.
Ahh Elixir. My favorite language that really just tries so hard to shoot itself in the foot. I'm currently in the protracted process of trying to upgrade a Phoenix app to the current versions. Currently I'm at the rewrite it in Rust and try out Rocket + Diesel stage.
Diesel is... interesting and makes me long for Ecto (which is often used as an ORM although the model bits got split off into a different project).
Erlang and Elixir have plenty of promise but there simply is no good story for production deployments. Distillery and edeliver approximate capistrano, and that sounds great when it works (although I'd just as soon skip edeliver). But when it doesn't I'd much rather dig into the mess of ruby that is Capistrano than the mess of shell scripts, erlang, and god knows what else goes into a Distillery release.
Elixir is a really interesting language, but Phoenix seems to still be pretty wet behind the ears and very much in flux. Ecto too to a much smaller extent.
1: Some of the distillery scripts can communicate with epmd, some just give up.
Also, one of the benefits of Mongo's API is that it has excellent native implementations in numerous languages (we already use C++ and Python), so a suggestion to switch language entirely is not really equivalent.
Huh? The aggregation framework is a solution to a mongo-only problem. Most other databases are performant, but Mongo suffers wildly from coarse locking and slow performance putting things into and retrieving things from the javascript VM.
> For example, can it unwind an array field, match those subrecords where one field (like a "key") matches a value and another field (like a "value") exceeds an overall value, and then apply this condition to filter the overall rows in the table?
This sounds suspiciously like a SQL view.
Edit: But if you actually need an array in a cell, Postgres has an array type that's also a first-class citizen with plenty of tooling around it.
As far as views are concerned, I don't know what to tell you. Sure, you'll probably have to craft the view itself by hand. The result is that you can then use most abstractions of your choosing on top of it though.
Yes. There will be a subquery and jsonb indexes need to be thought out in order to make it fast
At my company we built a UI on top of Slick that lets users of our web app define complex triggers based on dynamic fields and conditions which are translated to type-safe SQL queries.
Every language I know of has great ORMs which do this for whatever SQL flavors people tend to use on that platform. I write things like this all the time, and it gets turned into SQL for Postgres:
```` Article.where(author_id: 37).order(:modified_date, :desc).where.not(published: false) ````
When using an ORM correctly (and indeed, the less I'm using any of my own bits of SQL the more this is true) I am also protected against injection attacks.
I'm not saying NoSQL has no value, but I believe it to be the wrong tool for data that lends itself to an RDBMS. If you have a bunch of documents who have deeply nested or inconsistent structures and where it makes no sense that you'd want to query by something other than the primary key, sure, it's a no-brainer to use a NoSQL system. For a CMS, which has been implemented thousands of times in RDBMSs, it is madness though. I cringe at realizing that apaprently there are developers out there who have avoided learning SQL entirely in their career out of fear, and as a result have to use Mongo for every application because that's the only thing they know how to do. I'm sure they're out there, but I wouldn't hire one.
Marketing, marketing, and more marketing. Mongo was written by a couple of adtech guys.
> Not only that, their marketing/outreach efforts were also aimed at younger developers. When was the last time you saw a Postgres rep at a college tech event?
I remember being underwhelmed by two things at the one MongoConf I went to earlier this decade:
1.) My immediate boss was an unfathomable creep who was there mostly to pick up women
2.) Mongo was focused on how to work around the problems (e.g. aggregate framework) rather than how to solve them.
I can't recall ever seeing a Postgres rep, but I can recall having worked out a PostGIS bug with a fantastically tight feedback loop. The Postgres documentation and community are nothing short of amazing.
Meanwhile with Mongo I watched as jawdropping bugs languished. IDGAF what the reps say, anyone with even a few years experience should've been able to see through the bullshit that Mongo/10gen was/is selling.
- Misunderstanding by most developers of the relational model (I heard a lot of blathering about 'tabular data', which is missing the point entirely).
- The awkwardness and mismatchiness of object-relational mappers -- and the insistence of most web frameworks on object-oriented modeling.
- The fact that Amazon & Google etc. make/made heavy use of distributed key-value stores with relatively unstructured data in order to scale -- and everyone seemed to think they needed to scale at that level. (Worth pointing out that since then Google & Amazon have been able to roll out data stores that scale but use something closer to the relational model). This despite the fact that many of the hip NoSQL solutions didn't even have a reasonable distribution story.
- Simple trending. NoSQL was cool. Mongo had a 'cool' sheen by nature of the demographic that was working there, the marketing of the company itself.
I remember going to a Mongo meet-up in NYC back in 2010 or so, because some people in the company I was at at the time (ad-tech) were interested in it. We walked away skeptical and convinced it was more cargo-cult than solution.
I'm _very_ glad the pendulum is swinging back and that Postgres (which I've pretty much always been an advocate of in my 15-20 year career) is now seeing something of a surge of use.
I do remember a lot of MongoDB t-shirts, cups and pens around every office I was in around 2011-2013. When I would ask they would tell me that a MongoDB developer flew halfway across the world to give them all a workshop on it.
Ability to store Algebraic Data Types and values with lists without a hassle of creating a ton of tables and JOINs. Postgres added JSON support since, plus there are now things like TimescaleDB, which didn't exist previously.
On the server front that means ever increasing complexity with decoupled microservices and latency issues that play nicely with the classic approaches to those domains (static typing, functional programming).
On the data front sites like HN, Reddit, or Facebook need scability more than consistency, and have oodles of 'uninteresting' data that jives nicely with a schemaless document store.
Text and still imagery based media companies aren't big data providers. NoSql was a bad choice from the start.
They used the version of OpsManager that doesn't manage the deployment - is specifically not a deployment manager. Mongo does offer a managed version of this software, which the author mentions - with a justification for why they couldn't use that offering. However, I think this was the main mistake that The Guardian made. As the author notes: "Database management is important and hard – and we’d rather not be doing it ourselves." They underestimated the complexity of managing database infrastructure. If they had been attempting to set up and manage a large scale redundant PostgreSQL system, they would have spent an enormous engineering effort to do so as well. Using a fully managed solution - like PostgreSQL on RDS from the beginning would have saved them time. Comparing such a fully managed solution to an unmanaged one is an inappropriate comparison.
Full disclosure - I used to work at MongoDB. I have my biases and feelings w.r.t the product & company. In this case I felt that this article didn't actually represent the problem or it's source very accurately.
> Clocks are important – don’t lock down your VPC so much that NTP stops working.
> Automatically generating database indexes on application startup is probably a bad idea.
> Database management is important and hard – and we’d rather not be doing it ourselves.
This is true of any database infrastructure with redundancy/scalability requirements.
What they did was take a technical problem and solve it by buying an off the shelf solution. Which is fine, of course, but I’m a bit surprised by the reaction here on HN.
Re criticism of OpsManager - I think this is fair, given the sheer number of hoops we had to jump through to get a functioning OpsManager system running in AWS - no provided cloudformation, AMIs etc. £40,000 a year felt like a lot for a system that took 2 or more weeks of dev time to install/upgrade. The authentication schema thing was a bit of a pain as well, though we were going from a very nearly EOL version of Mongo (2.4 I think).
That sounds awful. Reading these stories I'm happy I work with small companies without such a huge infrastructure.
Training is exceptionally hard. Databases are hard to manage, and it takes years to learn to diagnose their function on unknown hardware / software.
This all being said, "no provided cloudformation, AMIs etc." is no bueno - not a good experience for the user.
If you haven't used Mongo with WiredTiger, you really haven't used it at it's best.
I work for MongoDB so if you have any questions about storage, feel free to reach out.
Review driven design is likely a large contributer. If you know you will be looking for a new job in 2 years, you better get some experience in $latestTrend instead of $provenTechnology.
For another, it may be everyone is susceptible to the influence of trends. Even engineers. And the more complex the details of a subdomain of software engineering are, the bigger the tradeoff between becoming someone who can make a true engineering assessment in that niche and developing expertise elsewhere...
Further, most of the use cases are fairly shallow, so it doesn't actually matter that much, and you'll be paid more to paper over the cracks.
We need to change the market, but it doesn't like the strong and stable engineers are getting lead roles to pay more strong and stable engineers to build boring tech, despite that actually being better for most businesses.
It should be the default choice (or plain filesystem storage), unless you have a specific requirement for something different. In the latter case, this should be people who are well informed about the database choices they're making.
That means that 18 months after you implemented a feature you've long since forgotten about, you don't have to come back and figure out what the new plan should be. And that's huuuuge.
You can find accounts of the db changing the plan from a good one to a bad one, but I'll go out on not that far of a limb and say that those are < 1% of the cases. Nobody complains about the queries they didn't have to come back and change. And the better the optimizer, the better that trade off will be.
So again, how do these poor choosers make better decisions in the face of changing data when they have to write all the algorithms by hand, and understand how to scale the data correctly? Seems to me that pretty soon they'd just be writing a very crappy database on top of some keystore.
the whole promise of mongo is/was distributed (HA+LB), which was all the rage back then, when AWS AZs dropped like flies every few weeks and scaling was seen as the problem. go fast, break things was the mantra.
and it's still not trivial to do pgsql maintenance without downtime, whereas in a clustered/distributed "solution", you can enjoy certain additional freedoms. this includes the freedom to shoot yourself in the foot (data inconsistency, but you had to fight a bit for that state by promoting a non latest slave to master).
Yes. It is so easy to bring up a mongo cluster and feel like you have HA. Don't worry it won't be proven wrong until you have writes during your cluster degradation and are unlucky.
it's pretty great.
and, most importantly, as far as I know the default replica set config is safe. (durable, consistent, atomic) it handles a lot of failover scenarios for you automatically. and you have to manually intervene to get into a bad state - which might be okay for that particular business case (better than full downtime)
One major change is that in 2010 you couldn't always run your whole DB in RAM, so there were some real performance benefits with MongoDB.
Another difference is that MongoDB was early with great JSON-support. Something that Postgres has since gained.
I think there are pros and cons with both. If I'd chose one today, for most tasks I'd probably chose Postgres.
That said, to understand why people made those choices, you need to 'teleport back' to that time and compare them in that year (given the tradeoffs at that time).
I worked with databases that only had table level locks(not row level) and there were more than enough occasions I cursed the creators.
Instance level(& indeed DB level) is insanity unless your DB is a read only DB.
They have very few write actions; in the Guardian's case hundreds of authors publishing a few articles a day, versus millions of users reading data. Writes would be sporadic and batched.
News products also changes more rapidly than you might expect. A modern news team may be writing and shipping code to better cover a breaking news event. I can imagine why a document store would be appealing to them.
I now use json supported functions in SQL Server and do not have the need for a different type of database. SQL Server handles my small 'documents db' implementation with the infrastructure of a RDMS. Win win for me.
To me Mongo just got popular by mistake way too early. It's like having a celebrity retweet your post because they liked what they saw at the time, exposing you to the world where everyone now thinks you have something important to say. Not surprisingly, you don't!
I don't see how Mongo was built based on Postgres, either?
I might be misremembering, though. And there's a possibility that what I read then was inaccurate.
The relevant data structures and indexing mechanisms (hstore, tsvector, gin, gist) were all available -- if only experimentally -- before Mongo launched.
As an aside, I did consider using SQL Server until I looked at the licensing fees. Why would someone choose SQL Server when options like Postgres or MySQL/MariaDB exist? Is there a specific SQL feature (MSSQL or Oracle SQL) provides which is not available elsewhere and would be a core feature which the companies data storage is built upon? I.e. that feature is so important that the companies product architecture would be fundamentally different without using the proprietary database.
Because they're in an organization where they're already invested in Microsoft technologies, so it's much easier to just use MS's DB instead of something entirely different that doesn't integrate as well.
It's similar to why many people use Apple software products that aren't as good as alternatives: if you already have an Apple platform/device, it's easier and better integrated and likely already installed.
Similarly, if your organization is already a Linux-based one, running Linux servers for everything, adopting SQL Server would be a huge PITA and would likely be laughed out of the room if suggested.
Outside of my professional career, with my personal projects I use SQLite, but if I were to build out something that intended to be larger scale business venture, I'd probably go with Azure SQL Database for multiple reasons. A large one being MS's overall integration, including their full control of the stack and CI/CD with .Net->Azure DevOps->github->Azure Pipelines->Azure PaaS/SQL Database. I admire what they're building, but a lot of people work outside of the MS realm. Companies don't tend to have a tech stack bone to pick as we do on HN, and many are already using part of the MS stack.
In sum, there's just a lot of business realities at play. I wouldn't personally run out and buy a SQL Server license myself, for what I do at home or with my (very) small side business, either.
Dealing with one vendor, M$, is better than dealing with 20. You pay for one support package and get all the benefits that come with one platform.
As much as I love open source, I do see the benefit of working with one vendor on all of your stack needs.
M$ reporting systems of SSRS SSIS and other is probably their bread and butter when it comes to DB space. Very few want to spend 60-70 hours a week building complex command line reports where with SSRS and SSIS most of these tools come with a nice GUI that helps you build these reports. Sure there are others that do the same but most required a third party vendor to add to the functionality. For example Apache. With Java and Apache I've had to deal with literally 15 to 20 different vendors just to do the same thing I can do with one vendor. Jenkins, Zookeeper, Camel, Cassandra, Tomcat all under the banner of Apache, but in reality governed by their own set of standards. I mean take a look for your self: https://www.apache.org/index.html#projects-list
1) MongoDB Atlas would work well here: it’s Mongo hosted on a cloud provider of your choosing managed by the people that make it.
2) a DBA, even a noSQL one*, is worth every penny. (I don’t mean this to put down noSQL DBAs just that it’s good to have someone dedicated to managing them and that a DBA versed in the administration and upkeep of a relational system could retain pretty quickly to help get the most out of a noSQL one)
A few things stood out at me. In no particular order:
- Going with Mongo in the first place cost them dearly. CMS are a weird application for schemaless. Not necessarily wrong, but definitely weird. I wonder if they would benefit from moving to more structured schema and I'm willing to bet a lot of the migration complexity comes from that in the first place.
- God damn that is a long migration. Holy shit. I know tech isn't their core competency but they do seem to have very competent staff. I've worked in places this glacially slow and I get how it can get this bad but it really strikes me as having no right to be. 10 months to migrate .... from the point where they were ready, until they were done. And somehow integration tests got overlooked during all that.
- Close call with dynamodb. Bit of a wtf on waiting for the feature to be implemented for nine months though. I'm sure they have an account manager with aws... I definitely think blocking an internal process on such a fragile externality as a closed process upstream publishing a feature is the wrong move and a red flag. Their migration path would have probably been harder with dynamodb too.
- I feel the pain of their troubleshooting issues on the load testing step. I can completely see how this can happen. That said it also raises a few red flags to me. Letting something as simple as the load testing step get complex enough to require weeks of engineering is ... Eh.
Have more thoughts but I hate typing on mobile. This is a fantastic write up, I love when non-tech companies publish this sort of stuff. And hurrah for postgres.
That depends. If the CMS is managing documents that have flexible structures that are tree-like, it might not be such a horrible idea to model that structure in a document instead of relationally.
C = Actual content to be published
MS = User accounts, user permissions, user authentication options, user authentication logging, meta change logs (user X removed tag Y at time Z), behaviour logs (user X viewed revision Y of article Z at time A)... and that's just scratching the surface of a very, very basic CMS.
I'm being generous and assuming articles, tags, bylines, and attached media (with full change history) is all "C".
I am saddened by many of the comments here though which equate to: "never try anything until you know everything" - sorry but that's just not realistic and it's unfair to the people who - commendably - contribute to these write ups and hold their hands up to mistakes-made, decisions that went badly with hindsight, etc.
Bigger picture: the Guardian appears to be thriving, and is succeeding based on the efforts of the tech team here. So if you read this and come away with a sense of "failure" you're probably missing something important.
It's OK to try new things and fail - just get back up and keep on trying, and use the new wisdom you build!
Plenty of people here like to dish on Mongo and the product seems to have been re-architected a few times since I used it seven years ago. By what metrics can we say the product is one worthy of passing a HN smell test? Passing Jepsen was seemingly not enough. https://www.mongodb.com/jepsen
> This interpretation hinges on interpreting successful sub-majority writes as not necessarily successful: rather, a successful response is merely a suggestion that the write has probably occurred, or might later occur, or perhaps will occur, be visible to some clients, then un-occur, or perhaps nothing will happen whatsoever.
> We note that this remains MongoDB's default level of write safety.
- http://jepsen.io/analyses/mongodb-3-6-4 2018-10-23
The summary of the latest MongoDB report [1] follows,
"In this Jepsen report, we will verify that MongoDB 3.6.4’s sharded clusters offer comparable safety to non-sharded deployments. We’ll also discuss MongoDB’s new support for causal consistency (CC) in version 3.6.4 and 4.0.0-rc1, and show that sessions prevent anomalies so long as user stick to majority reads and writes. However, with MongoDB’s default consistency levels, CC sessions fail to provide the claimed invariants."
That's really not what I want to hear about a place I'd be storing data.
Was it just a prototype back then? (When did it stop being a prototype?) Would the creators be comfortable using version 1 still?
Actually, I think that could be a good test of database stability: would you be willing to run a 5 year old version of your own software?
Common sense and formal education? I'm sorry, I'm aware of how incredibly snarky and arrogant that sounds, but in this case I always struggled to comprehend how MongoDB, or most of "NoSQL" in general, was considered viable to begin with.
"Schemaless" just immediately means that instead of the database keeping consistency, you now essentially have to do all your type and constraint checking in the application; I never understood how that's favorable, and how that remains maintainable in any way. On top of that, NoSQL (and maybe Mongo in particular?) also decided to throw away all guarantees that classical SQL databases have been offering for multiple decades for very good reason.
I think everyone who's had a somewhat theoretical class on databases and learned about, for example, what ACID really means, will have smelled that something doesn't quite add up here.
This is so true. Some component has to maintain integrity of the data. Should it be the application developer? Or the database software team?
I know which one I'd choose, and which one is focused on data integrity and not business logic.
if you're using partitioned (inherited) tables in Postgres, then it's the application developer, as foreign keys on inherited tables don't really work.
the workaround is triggers.
But eventual consistency may be a viable trade off to help you scale, but the thing is you probably need to be at a massive scale before it is worth moving away from traditional SQL, which can have read replication and the app can have caching too. But there is probably some point at which eventually consistent nodes make sense.
In a sense that is what a CDN is, and we like to use those.
Whoa, that was close. I really don't see why anyone would choose DynamoDB as a general purpose data store, unless they enjoy wasting countless hours finding ways around the limitations it imposes about how data should be stored and accessed. At least that was my (admittedly limited) experience. Postgres is a much better choice.
It's an epic shift in thinking to go from a schema of 50-60 tables to 1. Almost every dev I've worked with is very new to this.
The #1 sign a dev has no business using NoSQL: They chat about how flexible NoSQL is. Yeah, you can add attributes on the fly but I've found Dynamo to require much more careful planning then RDBMSes. Most devs I've met can understand when they need to use an index. But, almost all of them have issues predicting if their changes will lead to unbalanced requests against partitions.
Anyhow, since you can even connect the PostgreSQL WAL easily to a log stream like Kafka or Kinesis, I'm not sure why you'd ever start with a NoSQL DB, unless you just had master NoSQL data wonks.
I mean yeah who knew that blocking NTP therefore time drifting would break everything...
For those criticizing MongoDB, Fortnite generates $3B/year and runs on MongoDB, you should tell them it's a mistake and that they should use PG instead.
But I'm also one of the the people that still prefers to run MySQL over PostreSQL just because the tooling is still far superior
https://www.percona.com/doc/percona-monitoring-and-managemen...
https://www.percona.com/software/database-tools/percona-tool...
https://www.percona.com/software/database-tools/percona-moni...
That being said they had a massive outage due to a MongoDB issue.
https://www.epicgames.com/fortnite/en-US/news/postmortem-of-...
https://www.epicgames.com/fortnite/en-US/news/postmortem-of-...
1. Stop. Trying. To. Build. Your. Own. Cloud. A pizza shop doesn't build their own cars to deliver pizzas.
2. There's no such thing as hassle-free anything, unless you are paying someone else to deal with the hassle. Sales teams lie.
3. Justifying an untested idea with "but it's modern technology" is going to backfire. Follow established patterns with good track records.
4. Writing your own in-house behemoth product, of which there are already many kinds available, results in long-term expensive engineering projects necessary to to get around the high costs you didn't know were coming.
5. Don't write business logic, or your primary software product, in a way that talks directly to a database. Just... No.
6. Magic new technologies that remove the problems of old technology also introduce the problems of new technology.
Sure, use an ORM to abstract away the differences between PostgreSQL and MySQL (up until you need to care about them). That's reasonable. But maintaining a magical MaybeSQL layer that's powerful enough to not totally suck is going to totally suck.
(Alternatively, people find ways to use the abstraction layer such that it produces the desired usage pattern. Of course, then the code is no longer truly portable to another DB, because that same pattern is likely to be a perf issue there.)
* getUsers() * saveUser() * listUsers() * deleteUser()
and implementation can implement it in specific way to take advantage of chosen technology. If you have some esoteric use cases they could be handled in special way separate from business code.
Being tasked to come up with an abstraction layer that supports the speed of precomputed, clustered indexes with the flexibility of SQL - if I were in a content creation business and not a database engine writing business - sounds like the kind of project that would make me quit my job and go on a one year silent meditation retreat.
The OP is correct, your app can speak to an internal API without the underlying database infecting your domain code. That in no way implies you can't take advantage of the best of each database.
And it's not just queries. Transactions often have important semantic differences that will be visible on application layer - again, even between different SQL implementations (e.g. MVCC vs locks).
Which is hidden in the query interpreter for said db implementation. Each implementation can break down that abstract query into whatever implementation specific query works best in that database.
There's always some abstract way to represent it that doesn't require vendor specific knowledge nor does it remove the ability to apply vendor specific abilities.
Look, I just don't agree with you, I agree with OP. Db specific stuff should be hidden from the domain layer by an abstract query representation and an abstract transaction representation to be plugged in at a later time.
Stuff like "each implementation can break down that abstract query into whatever implementation specific query works best in that database" is wishful thinking. It's like saying that Java is faster than C++, in theory, because JIT can produce better code. And in theory, it can. In practice, we're not there yet. Same thing with high-level database abstractions - they're all either leaky in subtle ways, or they constrain you to extremely basic operations that can be automatically implemented efficiently on everything (but e.g. forget joins).
Several actually, which is why I know what I'm talking about; I've explored this area extensively. When Fowler first released PEAA I dug and went nuts and spent years coding up and exploring all the possible approaches and figuring out which ones I liked and why and which ones I didn't and why.
> Same thing with high-level database abstractions - they're all either leaky in subtle ways, or they constrain you to extremely basic operations that can be automatically implemented efficiently on everything (but e.g. forget joins).
If you're doing joins in your ORM, frankly, you're doing it wrong. Most ORM's do it wrong, they try and replace what a db does best; the right way to do it is to keep joins in the db. The role of an ORM when used properly is to map tables and views into objects and allow querying over those tables and views with an abstract query syntax. Joins belong in a view, not in code. It's called the object relational impedance mismatch for a reason, you have to draw a line in a reasonable place to get anything reasonable to work well and putting joins into the ORM is crossing that line and is why most ORM's utterly suck. Joins aren't queries, they're projections; put the queries in the code and the projections into the database, this works perfectly and lets each side do what it does best. Queries are easily abstracted, projections are not, projections don't belong in the ORM.
Incorrect. Named tuples will give you nothing back but a result set; the impedance mismatch refers to the mismatch between result sets and a domain model; getting tuples back doesn't remotely address this problem. I'd suggest you don't understand what the objection relational impedance mismatch problem actually is.
It's not a type system problem, it's fundamental mismatch between the relational paradigm and the object oriented paradigm. If a domain model has customers and addresses, and you do a relational query that joins the customer table and address table to return only the customer name and address city, the resulting set (name, city) doesn't map to the domain objects and isn't enough data for the domain model to load either of those objects which may contain various business rules. This is what the impedance mismatch refers to, relational projections of new result sets simply do not map to the OO way of doing things. Joins that create new projections are a relational concept that have no place in the object oriented world view: objects don't do joins, and object queries don't return differently shaped objects.
Hacks like partial loading of domain objects are attempts to mitigate the impedance mistmatch, but they do not solve it; they cannot solve what is a fundamental difference between two different ways of seeing data. Data is primary in the relational model and its shape can change on a per query basis, this is incompatible with the object oriented view of the world in which whole objects are primary and data is encapsulated and thus hidden.
The object relational impedance mismatch does not refer to a language problem, it refers to a difference in paradigm between OO and relational. It exists in all language regardless of the languages abilities and it's not a problem that can be solved, only mitigated, if you want to use both paradigms. You can solve the problem by avoiding using two paradigms, by either bringing the relational model into the application and not using OO or by using an object database.
Well, a vehicle that can squeeze down an alley won’t have the cargo room of a giant highway truck. The inter-city drivers will hate its poor capacity. Likewise, one that has a big 20-speed transmission for for hauling heavy loads is going to drive the city drivers nuts. They’re going to end up with one single interface to all possible roadways that everyone can come together and agree to hate.
If the database API is so free-form that you can store anything in it, you won’t get the advantages of PostgreSQL’s strict typing and lightning fast joins. If you make it so regimented that your data model ends up looking like a set of tables with foreign keys, then it won’t be able to make full use of MongoDB’s... whatever it does well.
They’re different animals. Choosing one highly affects the rest of your system design, from how you arrange your data to how you add new data to how you search for it. PostgreSQL and MongoDB have fundamentally different strengths and weaknesses, and if you make something that works equally well with both, it’s inherently going to suck equally on either.
1. Well, they probably own the car though which is a better parable. You don't need to rent your car, you can simply purchase it just as you can purchase servers. People build their own garages.
4. So you should only used already written software? Sometimes it is just faster and better to write it yourself. You get less dependencies, you know how the whole thing works etc.
5. Why not? That is exactly what one should do.
Adding an abstraction on top of DB query won't help much if you are moving from one type of DB to another, say SQL to document, or document to graph.
Do people use ElasticSearch as a primary data store? In my limited experience implementations don’t treat it as the source of truth.
We unfortunately inherited a cluster from someone who thought it would be appropriate as a SoT for forensic data, which causes me no end of grief.
https://www.elastic.co/guide/en/elasticsearch/resiliency/cur...
I'm a little skeptical of the idea that you can do custom software development with a custom database schema and realistically expect to outsource DB management. But sure, you can try. And in any case, you'd hope a largely read-only and document oriented dataset like a newspaper's has a relatively simple schema; without too many crazy schema quirks.
I'm sure you can get advice or buy know-how; but they're too coupled to think you're not also going to need to spend some time too. (At least: assuming your workload is large enough and complicated enough that naive brute force isn't an attractive option).
Anecdotes like this are just that, anecdotes, there are no numbers in this article to show a trend away from Mongo, Mongo is actually continuing to gain adoption. See https://db-engines.com/en/ranking
> -8.82%% As of 11:56AM EST. Market open.
1 month chart still looks green so this might just be a correction.
It's under `/info`, which un-suffixed redirects to `/about`, which I _can_ find from the home page (it's linked in the _footer_) but I can't find this engineering blog from `/about`, or anywhere else I can get to from `/`.
I don't think it's so much a positive choice to use the same system as it is just using (a possibly separate instance of) what they already had, and still squirreling it away in a corner.
This requirement made some sense in a world where a rogue employee might yank your database server out of a rack and walk off with it, but I don't understand why this is still considered relevant in an AWS context. First of all your data is never really "at rest", a huge point of Dynamo is that the data is always available (at least 99.999% of "always" anyway). Second "at rest" where? Literally on a physical disk? Even if there was a single physical disk that held it I'm willing to bet that even if you were in the right AWS data center and had support of a willing AWS employee you couldn't find that disk and even if you could there'd be no way to remove it. I'll guess its literally millions of times more likely for you to fail to deprecate some AWS API key that's sitting in clear text on a laptop you sold as surplus than anybody gets their hands on your at rest Dynamo data.
The regulations and the attack risk aren't always in alignment unfortunately.
Interesting. I didn't know you could make indexes for things /within/ the JSON.
The only limit that I've run into in comparison with NoSQL databases is that you cannot index fields inside arrays within the column (json -> arrayIndex -> field), you need to normalize the array to a different table.
https://dev.mysql.com/doc/refman/5.7/en/create-table-seconda... contains some JSON-specific examples.
The proper solution is to onboard the proper talent and tools or outsource it all. MongoDB does have fully-managed offerings that will automatically migrate and run your database, across cloud-providers if you need, and would've cost less than all the time and effort spent on this. This is just another example of poor technical competency at media publishers.
Great blogpost, though; well written, which is always a surprise.
I'm guessing their original choice for MongoDB was more for the schemaless development flexibility rather than horizontal scaling capabilities. This seems to have worked well enough for them. DynamoDB seems like a natural fit coming from a document store and they did very well not to choose it.
I think AWS is due for a managed NewSQL store comparable to Google Cloud Spanner or MS Cosmos DB.
It's not just media companies rewriting something significant every few years; everyone in the tech industry does it. There's a very good reason for it: it provides interesting and high-paying work for tech employees.
Who wants to go to a company, set up a great system that doesn't need any more work except maintenance, and then just settle down into maintenance mode? No one who wants a rising career. "Maintained system using $oldTech" doesn't look good on a developer's resume, whereas "migrated critical system to $newTech" does.
And it's not just the developers: managers who oversee big teams of developers working on a big new project have more prestige and are paid more than managers who just oversee some boring already-done project that's in maintenance mode.
Now, in this case, there were probably very good genuine reasons for making a big change (I'm no fan of MongoDB), but many times if you look closely with a critical eye, you'll probably find that people padding their resume and making up a reason for their job to exist is the real reason something is being done.
Busy-work is never interesting, and if it's "high-paying" someone is getting fleeced (and will do what they can so as not to get fleeced in the future). At some point, the constant rewriting will have to stop. In contrast, maintenance work on a well-engineered critical system can be quite interesting in its own right (and it's not like the "maintenance" doesn't involve writing some new stuff on occasion. Nothing is ever truly "done").
Just to expand on this: if I were to take part in a project to move a critical service from Postgres to MongoDB (which I believe to be a worse solution), this would still be a better learning experience for me than just maintaining an existing installation. I'd have to learn about both DBs, and write a lot of new code to work with the new DB. That wouldn't happen with a Postgres system that's already set up and working fine. And my resume wouldn't look like someone that can come in and get to work on a big project and see it through by working on a maintenance project.
>At some point, the constant rewriting will have to stop.
At that point, the developers and managers who pushed for the big new project will have moved on to another company. Remember, moving companies every few years will generally yield a much higher salary than staying at the same company for your entire career.
>In contrast, maintenance work on a well-engineered critical system can be quite interesting in its own right
Sure, but you're not going to get a great salary doing that.
MongoDB, on the other hand, recently changed their license to an abomination that many people think is no longer Free: https://news.ycombinator.com/item?id=18301116
Mongo made a lot of money building a projects that runs on top of many other projects that were released as Free Software. Now they're upset that other people are building on top of Mongo in the same way.
Explain to me how a startup of 10-20 people can compete against AWS once they grab what you're working on to make an AWS service?
Changes Mongo, Redis made to their licence were made to protected against those practices.
2. HAproxy is still GPLv2.
3. Redis Labs's CCL and MongoDB's SSPL are not open source licenses, but their purveyors sure do like to give off the impression that they are open source. If you want to keep your code proprietary, keep it proprietary. Don't pretend to be open source. If you are not okay with others using your work, even making money off of it, as long they adhere to the rules of the open source licence you used to license your work to them, then don't license your work to them under open source licenses, or don't cry foul when they use it under the terms of the license.
If you care about using FOSS - as many companies do - MongoDB is no longer open for consideration.
> approximately 2.3m content items.
I had a previous project where we did a similar thing (except with HSTORE instead of JSONB) and it exploded rather dramatically (very simple queries took multiple minutes or timed out entirely) after around 30m rows. I hope the Grauniad doesn't run into a similar issue, or at least anticipates it better than we did.
Having issues accessing Xmillion rows seems like something that would be caught by someone focused on performance of the database full-time.
At my previous startup I was ingesting Hearthstone games at a rate of 1-2M / day. Before being handed off to permanent storage (s3, redshift etc) a bunch of data would get stored in a JSONB, with 14-day retention. This all ran on a 200GB t2.large instance on RDS, was our smallest instance and never really caused an issue.
1) Why use Scala to write (a relatively simple) internal CMS?
2) Why use a clustered database for 2 million records?
3) Why write your own proxy? (in Akka, none the less)
4) Why would you migrate articles from Mongo to Postgres using a script that runs overnight in screen?
The Guardian is, prima facie, a Wordpress blog. A simpler architecture would be:
1) Any CRUD web framework to build the CMS for reporters to draft their articles (Django, Rails, etc). Any basic RDBMS with read replication will do. Or, ditch the webapp entirely and just make a simple Markdown editor that commits to a git repo, a la Prose or Netlify.
2) When a reporter "publishes" an article, generate HTML for it and push to the CDN network. (I can't easily tell by looking at their HTTP headers, but I assume they're doing this already)
Okay, I'm being a little tongue-in-cheek. It's probably not that simple. But, one has to wonder, when you're serving up 100 million static HTML pages a day, if it really has to be this complicated.
Django itself was literally developed to suit the use cases of a newspaper.
A bloated web framework makes your code simpler because there are many things that you don't need to reimplement by yourself.
The problem is when developers use a framework without understanding it well and:
1) Reimplement in the code features that are already present in the framework. 2) Fight against the framework because their business needs conflict with the conventions choosen by the framework. 3) Fight against the framework because they don't agree in the way the framework solved a particular problem and want to solve it their way.
In both cases, the origin of the issues is not the framework itself.
> The most trivial way to implement a website for high traffic is simple static content generator
Agree. But in any case, those kinds of sites are not hard to catch neither.
I spent 4 years at a large financial news company, where we benefited from many ex graunaids who decided to migrate over the river.
They helped us create the new front end to the website, to much acclaim. However, it was hard for a number of reasons(this is from the financial news company, not the gruan.):
1) The journalists hated change, especially as they couldn't see any benefit. They just want to keep their same interface exactly as it is, bugs and all. They also had an active union.
2) There is 20 years of "micro services" moving data from the CMS, through various things to allow stuff like translations, syndications (very important source of money) data extraction, meta data processing, physical page layout, and many many more. Most of which is done by a legacy ETL framework pushing to and from a solaris FTP server that is old enough to join the army.
3) there is more than one way to enter data into the CMS.
4) The type of article, and the data in said article changed depending on where it came from, and what services nobbled it.
5) looking after the journalist's interface, curating the data, sorting the articles and adding meta data, looking after paying subscribers, and finally the front end, were all different departments that refused to talk to each other.
This meant that unlike a rational place, there was no source of truth for the CMS. It wasn't like you could call up article 342923 and display it. There was no guarantee that it would have all of the metadata (like were we allows to publish it) required. Add to that the inter department rivalry, which meant that for some reason the membership department were allowed to spend 4 years re-writing the same bit of functionality over and over again. (user management and payment gateways is a solved issue, but alas it took the best part of 25 million quid to find that out.)
To answer your questions:
1) because it scales maaaaan, looks good on my CV, I don't want to spend time doing boring work, I want to learn a new tool
2) see 1
3) see 1
4) Because I suspect that they've never seen a working ETL system
To answer your bonus questions:
1) Journalists have unions, changing the editor requires a _boat_ load of training, and is almost never worth it. Buy over build every time. But yes, its just text. However its the metadata that makes it. Whos in the article, whats the subject. Is it a lifestyle piece, does it have photos, who owns the copyright for the photos, is the article syndicatable, can we syndicate this article, who edited it. Etc, etc, etc. The text entry is the easy bit, its the parts that make it a real news paper that are hard.
2) Nope, almost certainly never done like that. The article will be given a UUID, and dumped into the CMS DB. The front page generator system will then dynamically pull out the articles based on parameters given first by the editors, (front page image, leading headline etc) then the related articles might be curated by hand, or by keyword/metadata or user's preference.
Then the advertising and tracking bits have to be injected, which account for 50-70% of the effort.
CDNs now allow a lot of logic to be pushed to the edge. (see https://labs.ft.com/2014/10/caching-user-agent-specific-resp...) which means that its not overly taxing to host a very large website.
Their proxy is one example where Rust will give the same capabilities as Scala and will be more robust, in the future, if it's not already the case today.
I haven't tried it yet, and there are some limitations at the moment, but I find it extremely promising.
Most of the times a boring database is a better fit.
Not everybody can be an early adopter, especially on the software which stores what is arguably your most precious asset.
Elaborating on my comment though, it would make sense to consider a drop-in replacement for a technology for which you've already invested yourself, but got bitten enough time to think about moving away to a different model (operational burden, data loss, etc).
I'm not sure if "baroque" is a good qualifier for FDB.
It's recent in the open source community, but has been running in production for many companies (including Apple) for some years. It's based on sound architecture, compare to many others.
This is one of the reasons why building on top of the old API wasn’t an option. There was very little separation of concern in the original API and MongoDB specifics could be found even at the controller level. As a result the task of adding another database type in the existing API was too risky.
This seems like the problem was more related to the API layer code quality rather than where the data was stored.
We moved that specific collection to MySQL, no problem and there was virtually no change in the data structure. Both were of course indexed.
SQL is not RDBMS is not ACID is not NoSQL.
When comparing NoSQL to something you should really compare it to RDBMS which is what most people conflate with "SQL". This conceptual understanding gap is IMO what led to the huge growth and misuse of "NoSQL" for a decade.
This bit sounds like a cautionary tale about the importance of a layer and well structured API. There’s no reason why details about your data store should leak into an API Controller.
This happened to us numerous times before and after we migrated to their Atlas platform. We are a major streaming company with more than 120M registered user. Disclosure, we are moving away from MongoDB as well.
- They'd been running the two systems in parallel and comparing their outputs for months. By the time the pulled the plug, they were confident that PostgreSQL was working fine.
- They had written systems to copy data from the MonogoDB-backed service to the one running on PostgreSQL. Once they started treating PostgreSQL as the official store of record, they'd be faced with either writing another system to mirror PostgreSQL back to MongoDB or committing to lots of double entry. Either of those sound painful.
They definitely want more than a key value store. They at least want date and content indexing, editor, author, etc. I imagine they probably have a few internal requirements as well wrt analytics, tooling and more. S3 doesn't work for all this and another user's suggestion that this is all treatable like static content sounds way off to me. The best you can do is pregenerate some of the HTML.
postgres all the way.
THe FT use(d) cassandra to store their membership details (6 million rows of largely static, read-only well structured data) The cluster was massive (12+nodes in at least two regions, from memory) slow and was impossible to upgrade reliably.
The support from datastax is shite. Backups are not reliable. imports even less so, and you are beta testing the whole system every time you do a point release.
CMSs have highly structured data. swallow your pride, map the data and build a proper schema. Its really not that hard.
Yes, cassandra has a graph layer, no its really not worth it. Yes it has gremlin, no you shouldn't need it if you've modelled your data correctly.
I've seen cassandra shine when it comes for write optimised loads. Pipe a bucket load of data into gremlin and magic happens.
But thats a specific workload, which is pretty rare, and certainly not suited 999:1 read to write ratio. Thats not cassandra's fault, thats the fault of the idiot that chose it, and the boatload of idiots who carried on and added loads of systems that makes it much harder to migrate away.
Then we come to support. Datastax is the defacto support provider. They make a lifecycle manager, backup/restore tool, and push a load of patches into the main codebase. But it is shit
o Backups fail silently
o The only way to make alerts work (ie do an action, rather than create a popup when you log into the ops center) requires work to navigate the API
o Its full of CVEs, which are script kiddy-able
o migrating data from backups to new clusters was impossible to do without a boat load of manual work, failed 50% of the time
o restoring from automated backups was impossible until august.
o it couldn't use instance profiles on AWS until august
o upgrading to a point release silently breaks backups, _always_
Basically I spent the first half of this year QA very expensive software. There are some very very good support people, but there were some terrible ones as well.For backups, we found a tool on github to snapshot to S3. Worked fine as far as I know. It's the guys in the office next to me that were handling this, not me, never heard of any major issue.
Datastax builds a graph layer on top, or you can use something like JanusGraph, but it's never as good as using a real graph database with natively designed storage system.
In Amazon it's harder to compare directly as the support contract is paid across all of our accounts, but we're spending around $13,000 for a highly available db.r4.xlarge postgres instance.
Performance wise, querying without an index is SLOW, basically not worth it - as we end up doing a scan of the entire database. Fortunately we don't need to do this as we can usually rely on the guardian content API for proper searching stuff. The average API (a Scalatra app) response time reaches 150ms at 'peak time'. This isn't a high-performance use case - around 1000 requests/minute at 'peak' time.
My company does bespoke software development for large enterprises. If they're on AWS having all your services within a VPC you control is something of table stakes for a lot of large company engineering and security teams.
My SAs seem cautiously willing to try it, but I won't put them through the work of learning how to manage it unless I can find a clear advantage over Postgres+JSONB.
"Goodbye {Postgres|Mongo|MySQL}, Hello {MySQL|Mongo|Postgres}"
Would you say the same thing about programming languages?
I was actually listening to the Full Stack Radio podcast, and the latest episode talked about the power of moving more into the database. What struck me as a strong idea was that we can't treat databases as equivalent - different database servers have different strengths (for example, some have better JSON support or time evaluation functions).
I think it's particularly telling that you seem to be thinking of the database as a "tool", but if you're like most, you fight religious battles over your language of choice. In my experience, the database IS the application (and in many cases, the business) and is a far more important choice than the implementation language, the cloud provider, etc.
https://en.wikipedia.org/wiki/Narcissism_of_small_difference...
Moving into databases is not something I understand. I worked the last 10 years as data engineer but I like to move out from databases a lot more than moving into them. :)
OMG yes. Having worked in the Midwest, I promise you that farmers absolutely do trash talk each other about John Deere vs New Holland vs International Harvester. Tell a Ford driver that you like his Chevy pickup and prepare to hear about it for the next hour.
> Do you thinks that builders would fight over which hammer brand is better.
You seriously haven't spent much time around blue collar workers, have you.
But lemme ask this question, what are the things in which mongo is better than postgres?
/dev/null is more efficient than just about anything at ingesting billions of records, but not very useful, is it?
You are talking about replica sets, which is different concept.
You have to add capacity (shards) 3 nodes at a time, two third of which sit unused. It's not scalable at all.
replica set is for redundancy and availability, but you can use replicas for reads, so scale your reads traffic. pgsql works absolutely the same way, you have one master which accepts writes and read-only slaves/replicas.
OpsManager is great for team that don't have dedicated devops I think and have great dashboard/visualization.
However, run your own MongoDB is very easy. Not like Postgres(unless you used RDS). However, when using RDS, you still have downtime when upgrading db, it still have a small amount of time the standby in MultiAZ is promoted to master, DNS is updated, and during that yourcurrent primary is not writeable or even not available. MongoDB is way easier to operator. You add/delete/node on fly and client auto discover network topology. Plus https://docs.mongodb.com/manual/administration/production-no... this links give great info to tune it: thing like run on XFS, mount with notatime options...puts opslog on high iops volume etc...
Peformance isn't a factor to pick your database much nowsaday. They looks great on the benchmark. However, try for yourself before pick one on your workload. Postgres does has its own ward.
Pick a database based on how well your team confident with it, how does the database help you move fast enough or deliver business value. Don't follow the hype or silly benchmark.