Right, so the solution is more complexity? Of course it is. Sigh
Right, so the solution is more complexity? Of course it is. Sigh
Our Twitter-scale Mastodon implementation is a direct demonstration of this. It's literally 100x less code than Twitter had to write to build the equivalent feature-set at scale, and it's more than 40% less code than Mastodon's official implementation. This isn't because of being able to design things better with the same tooling the second time around – it's because it's built using fundamentally better abstractions.
This is also true of most of the main SQL database engines, no?
I fundamentally disagree with the premise of your blog post but not the premise of your company, so why don't I make an ask for you to write a much more practically useful application. Design a basic shopping cart using your system and compare and contrast it with a well designed relational equivalent. A system that allows products to be purchased and fulfilled is a far closer match to what the majority of companies are using software for compared to writing a Twitter clone.
Here's my take -- at a certain point in scale and volume using a database, it actually does make sense to rewrite all of the following from scratch:
- query planning
- indexing (btree vs GIN, etc) and primary/foreign/unique key structure
- persistence layer
- locking
- enforcing ACID constraints
- schema migrations/DDL
- security, accounts and permissioning
- encryption primitives
But more crucially, my belief and experience is that most companies making lots of money from software products will NEVER remotely reach that scale -- and prematurely optimizing for that is not just the wrong decision but borderline professional malpractice. You can get very far with Postgres and JSONB if you really need it, and you'll spend more time focusing on your business logic than reinventing the wheel.
I'd like to be proven wrong. But I get sinking feeling that I'm not wrong, and while your product is potentially valuable for a very specific use case, you're doing your own company a disservice by distorting reality so strongly to both yourselves and your prospective customers.
I'll round out this comment by linking another comment in this thread that goes very well into the perils of event sourcing when the juice isn't worth the squeeze:
We started with a Twitter demonstration because: a) its implementation at scale is extremely difficult, b) I used to work there and am intimately familiar with what they went through on the technical end, and c) the product is composed of tons of use cases which work completely differently from each other – social graph, timelines, personalized follow suggestions, trends, search, etc. A single platform able to implement such diverse use cases with optimal performance, at scale, and in a comparatively tiny amount of code is simply unprecedented.
It's a very impressive demo. You should be proud, keep your chin up, and keep us updated as Rama's value prop. continues to grow.
You're not really addressing the substance of my comment so I am left to assume that your omission is because you cannot address it. Be that as it may, you'll hopefully at least take my final comment at its face value that you do your own product a disservice by overhyping what its use case is towards areas that it is objectively a poor fit. People have tried event sourcing many times and it's just not a good fit for many if not most workloads. You can't be everything to everyone. There's nothing wrong with that.
My advice to you is this: call out the elephant in the room and admit that, and focus on workloads where it is a good fit. That extra honesty will go a long way in helping you build a business with sustainable differentiation and product market fit.
This "Twitter-scale mastadon implementation" is when my red flags went up. It's meant to demonstrate a simpler and more performant architecture, but it actually demonstrates "things you should never do" #1: rewrite the code from scratch.
https://www.joelonsoftware.com/2000/04/06/things-you-should-...
"The idea that new code is better than old is patently absurd. Old code has been used. It has been tested. Lots of bugs have been found, and they’ve been fixed."
The "1M lines of code" and "~200 person-years" of Twitter being trashed on in this article is the outcome of Twitter doing the most important thing that software should do: deliver value to people. Millions of people (real people, not 100M bots) suffered thru YEARS of the fail-whale because Twitter's software gave them value.
This software has only delivered some artificial numbers in a completely made-up false comparison. Okay it's built on "fundamentally better abstractions", but until it's running for people in the real world, that's all it is: abstract.
Please don't tout this as a demonstration of how to re-create all of Twitter with simpler and more performant back-end architecture.
- We stress-tested the hell out of our implementation well beyond Twitter-scale, including while inducing chaos (e.g. random worker kills, network partitions). - We ran it for real when we launched with 100M bots posting 3,500 times / second at 403 average fanout. It worked flawlessly with a very snappy UX.
The second-system effect is a real thing, but there's a difference when you're building on radically better tooling. All the complexity that Twitter had to deal with that led to so much code (e.g. needing to make multiple specialized datastores from scratch) just didn't exist in our implementation.
That is exactly correct. All testing is make believe except for real case studies by real customers with an intent to pay, and barring that, real pilots with real utilization. Otherwise, you run the risk of building a product for a version of yourself that you are pretending is other people.
At some point in time the same argument was made for relational databases despite there being stable systems built without them based on ISAM. The newer relational systems took a lot less work to implement but that didn't imply that it made sense to rewrite the ISAM based systems.
Joel was talking about commercial desktop software in the extremely competitive landscape of the 90s, he wasn't talking about world-scale internet service infrastructure. The architecture that delivered a set of X features to the initial N users isn't always going to be enough for X+Y features to the eventual 1000*N users that you promised your investors.
Companies like Google are quite public about how much they rewrite internal software, and that's just what the public hears about. A particular service might have been written with all of the care and optimization that a world-class principal engineer can manage, serve perfectly for a particular workload for several years, and yet still need to be entirely replaced to keep up with a new workload for the following several years.
You wouldn't tell another engineer that they shouldn't rewrite a single function because software should never be rewritten, so it also doesn't make sense to tell them not to rewrite an entire project either. It's their call based on the requirements and constraints that they know much more about.
Nobody should be rushing out to rewrite Linux or LLVM from scratch, and yet we wouldn't even have Linux or LLVM if their developers didn't find reasons to create them even while other projects existed. In hindsight it's clear those projects needed to be created, but at the time people would have said you should never rewrite a kernel or compiler suite.
If you can convince more developers to apply what is effectively canonical store -> workers -> partitioned denormalized materialized views as a pattern where it makes sense, then great.
But you can do that with just the tools people already have available. Heck, you can do that with just multiple postgres servers (as the depot, and for the "p-stores" and for the indexing functions), and then you don't need to ditch the languages people are familiar with for both specifying the materialization and the queries.
Part of the reason we use a "hodgepodge of narrow tooling" however, tends to be that it allows us to pick and choose languages, and APIs, depending on what developers are familiar with and what suits us, and it allows people to pick and mix. Convincing people to give up that in favour of a fairly arcane-looking API restricted to the JVM is going to be a tough sell a lot of places.
Does your implementation match the features of Mastodon?
Yes, we implemented the entirety of Mastodon from scratch.
But, at a certain level of scale everything is a data engineering problem, and sometimes this is the (relatively) simple solution when viewed in the context of the entire system.
'Just use mySQL/SQLite/Postgres' is great advice until it isn't.
I think it would a mistake to approach a domain and assume nothing can be made simpler or more straightforward.
More complexity? The writer makes it very simple. Just use their product Rama.
I genuinely believe we would all be better off if Martin Fowler's original article about Event Sourcing had never been written. IMO, it's a bad idea in 99% of cases.
One thing there is: These were massive engineering departments in very large companies. Think of several dozen data processing teams, each very much their own small company, all working on different domains with event and stream consumers and utilizing the possibility of restarting their stream consumers some time ago. Impressive Kafka-Clusters ingesting thousands and thousands of events per second and keeping quite the backlog around to enable this refeeding, outages and catching back up and such. At their scale and complexity, I can see benefits.
However, at that scale, Kafka is your database. And you end up with a different color of the same maintenance work you'd have with your central database. Data ends up not being written, data ends up being written incorrectly, incorrectly written data now ends up with transitive errors. At times, those guys ended up having to drop a significant timespan of data, filter our the true input events, pull in the true events and start refeeding from there - except then there was a thundering herd effect, and then they had to start slowly bringing up load... great fun for the management teams of the persistence.
Note that I'm not necessarily saying e.g. Postgres is the solution to everything. However, a decently tuned and sized, competently run Postgres cluster allows you to delay larger architectural decisions for a long time with run-of-the-mill libraries and setup. Forever in a nonzero amount of projects.
We use SQLite and Postgress for pretty much everything. And about 90 percent of the server code is auto generated using an inhouse code generator.