Shopify Ruby on Rails distributed monolith runs 19M queries per second on MySQL
twitter.com
twitter.com
Few months in, both services reached 12 hours to run. We had to change the start time just so they'd happen back to back. I rolled up my sleeves and got to work on the db.
A week later, the job's runtime was down to 2 hours. I was nearly promoted out of the job. I kept making improvements to the database, reorganizing the application, and optimizing. One day I manually ran the job, then went to get coffee. When i came back, prompt was awaiting my next command. I thought it had silently failed. I ran it several times throughout the day and checked the response. It ran in ~17 minutes. The backup was also reduced to less than 20 minutes. This was all on mysql 5.7
We literally gained 23 hours of availability. We had no idea what to do with it. I was fired shortly after.
I've made similar improvements (40x speedup, 8x reduction in RAM footprint) and had been shown the door. This was for software that was a bit over 30 years old too so it was a bit involved.
The team winning does not help a narcissist feel better. You'd be better off getting nothing done while stroking their ego daily.
Source: worked on Azure's managed Citus pg
At this point I'm pretty sure anything IT relate became a bullshit job.
Unfortunately the case when you do your job too well, always leave a good 15-25% on the table to keep yourself 'required'.
Such a simple summary of how companies will chew you up and spit you out.
I guess they can't help themselves and need to make their position known. Personally, I love ruby, and rails is pretty good too.
Sure it only had highly optimized queries (something rails makes kinda easy too), but that's what you are supposed to do anyway.
There wasn't ever a need to recode to run it on a $10 instance. Nether do I think 99% of all database driven projects, including mine, are ever going to reach as many concurrent users.
IMO: talking about performance for ANY major coding stack is just premature optimization
The gist of it is that there's one codebase with multiple separate "modules". This codebase is packed and linked as a library and then we build different super slim hosts that load different parts of the monolith in production containers. Usually just different environment variables or config.
But locally, we can run the whole thing in one process. We're.using .NET so `dotnet run` brings up the whole app. Whereas we might run parts of the app in different console hosts in prod, locally they are hosted background services in-process.
From a debug perspective, this is super awesome since you can just launch and trace one codebase. If we broke it out into 3-4 separate services, we'd have to run 3 processes and 3 debuggers. 3 sets of configuration, 3 sets of CI/CD, 3 sets of testing. Terrible for productivity.
We have parts of the system connected to SQS for processing events and if we need more throughout, we simply start more instances of the container all running the same monolith.
I think GCP is probably one of the best platforms for building modular monoliths because of its tight orientation around HTTP push.
1. As part of the same entity, so you could scale your operation.
2. As part of an ecosystem, so you could for example create an entire open banking network or just a regular network with bank transfers and card payments using proper protocols such as ACH, ISO8583 and ISO2022 for example.
And there is a separate concept of a configurable application which can completely reshape its component graph according to some high level configuration flags (like database=prod|dummy), we call it "multi-modal applications".
And we created the perfect tool for wiring them: https://izumi.7mind.io/distage/index.html
For us, it just becomes a matter of configuring the correct construction of the dependency injection container at host startup using some flag (usually environment variable) to pick the right bits and pieces to load into the container and which services to run from the monolith.
Then each of the host "partitions" gets its own Dockerfile.
Monolith on the other hand is a single application with all of the logic in one place.
Distibuted monolith is a set of applications\services like in microservices pattern but they can share common data storage and depend on each other.
Stateful vs Stateless.
Your monolith is a binary that gets distributed to hosts to perform some function. The binary has multiple entry points that can be envoked. Most calls are via internal library call.
Microserverices (also stateless) have a different artifacts for each component, services call other services via a private API (often grpc/httprc).
Then to handle any load we need to build autoscaling and spin up the toaster to medium potato and 20 instances (this costs 60k a month, but no worries we only pay for what we use so we will only run this for 27 minutes during our big sale).
Oh what wonderful world we live in and the pain we inflict on ourselves.
GJ Shopify for running a sane (tm) tech stack :)
(btw of course Rails scales it's shared nothing setup, spin up infinite app servers as long as the db can handle it. It's pretty expensive for compute though)
If you have a thriving business already using Rails, it's difficult to justify moving off of Rails...now that painful upgrades seem to be mostly in the past. However, I do find isomorphic JS components & state management to be a pleasure to develop & maintain compared developing an app in two languages.
Upgraded a .NET 4.7 project to .NET 6 in 3 days and haven’t touched it since. And the .NET project was larger.
I haven’t done rails since that upgrade and I hope i never have to touch it again.
If you squint, yes. Erlang has had NIFs and ports since forever precisely because it isn't great at CPU intensive tasks.
In most cases however, you're still looking at an order of magnitude less servers than the equivalent Rails/Django application.
checking again:
CODE: 198235 LOC total
SPEC: 127708 LOC total
We do the same at Shopify, and running off the edge allow to catch bug much sooner and identify them much easier.
It also very significantly cuts down on maintenance cost because we no longer have to work-around bugs, we can fix them upstream and so a small update.
As for you pointing at the Rails 5 migration taking a very long time, it's true that certain major migrations were a pain this one in particular, but it's because a major API had been removed (attr_protected). We (both Shopify and GitHub) work with the edge also in part to make sure the community won't have to suffer this kind of harsh migrations ever again.
All this to say you are blowing things out of proportion. The Ruby and Rails teams at Shopify and GitHub are not an overhead, they pay for themselves, and any tech company as large as Shopify or GitHub you will find major contributors to the stack they use. e.g. There used to be a JVM team at Twitter (half smaller than Shopify even at peak).
I don't see the JVM team as the same thing as a team dedicated to working off a project's core. The JVM team looks closer to the YJIT project. Making a JIT also calls into question Shopify's scalability: Rails was slow enough that they allowed a JIT to be built internally? That's quite the trade-off to make.
That's the same team (Rails + Ruby)...
> Making a JIT also calls into question Shopify's scalability: Rails was slow enough that they allowed a JIT to be built internally?
You are conflating scaling ability with language speed, there is of course some relation between the two, but it's mostly orthogonal.
At the scale of Shopify (several thousands developers) having a few people focus on improving Rails and Ruby is a drop in the bucket and payoff immensely.
Rails and Ruby are both Open Source projects not backed by for-profit organization. It's perfectly normal for an user of such project to contribute patches... That's how Open Source is supposed to work...
There is plenty of organizations contributing patches to the Linux Kernel (Google, IBM, etc). Using your logic that means Linux is slow enough that Google allowed a new scheduler to be built internally?
That's quite the silly line of reasoning...
But we continue to see cost / performance improving. Ignoring the Rails framework and Ruby VM both together has likely gotten 2x speed improvement. We will get 96 Core Graviton v3 or 128 Core Zen 4 EPYC. Cost / Core performance is coming down and will continue to do so at least until Zen 6.
Depending on your App, somewhere along the line it will surely lean towards Ruby Rails's flavour. Assuming you do value what Rails have to offer.
I am just waiting for fibre/ async ( or something similar ) to be built into Rails and Active Record, along with even faster RubyJIT.
b) They run multiple app servers so each component can in fact be scaled separately.
c) They built their own PaaS which runs on Kubernetes.
d) You would have to be utterly incompetent to not care about optimising costs.
But I am glad we both agree that Shopify is doing a great job running a sane stack.
The database would be the bottleneck and that's got nothing to do with rails - it's all mysql.
I would argue that the computing costs of any computing endeavor (short of AI, at this point) are dwarfed by the human costs, and my personal experience is that Ruby/Rails is at least a 10x headcount savings over the programmer sprawl required for Java/JS.
Java and JS sure are good for job security. I'll give them that. Throw in an unequivocal demand for Oracle, and Cisco networking gear, and you have the whole Fortune 500 world that got dumped on us from the people who couldn't manage the mainframes, either.
I’m not a ruby hater, but the average joe can’t accomplish this. If you consider each store an individual instance with an isolated/shard of a db it makes sense. But the underlying foundation is immense.
Partitions are usually the key to scale.
I’d imagine parts of their infrastructure would be better served by different runtimes. They’d save a lot of money.
But if your entire team is hyper focused on ruby there is something to be said for a huge monolith.
Each instance is probably just a k8s pod
/s
The average joe does not need this.
If you have decent margins on the pages you're serving, Rails is fine. Where you might want to investigate other things is if you're, say, an ad driven business with really thin margins and you want to minimize costs. Or if you've got things dialed and you're just not changing things much any more and you want to eke out some savings.
And in any event, Rails is a good choice to figure out the problem space you're working in. Even Twitter started with it, and objectively, Twitter is very much not the sweet spot for Rails.
If you’re doing a tremendous amount of parallel processing it will fall over without throwing lots of compute and scaling horizontally. Rails doesn’t scale vertically. You need to give it compute, and every other resource it uses will also need more compute. Average joes are fine with a 2 core VPS. Lots of businesses are not.
I need you to send 5-10 million API calls per day to 20 different API’s … are you using Rails? Every API is rate limited differently, with different batch sizes (each having unique parameters) and it needs to literally be done ASAFP to make certain deadlines. If you want to throw money at the problem, sure. If you want to do the same thing with 1/10th of the resources you’ll use a better runtime like Erlang/Elixir or Clojure/Golang something with CSP.
HN is so quick to dismiss things like kubernetes and wax poetic about simpleton life but there are very legitimate reasons to choose alternative tools for your problem space.
It seems most really big companies make the boring but safe choice to go with the JVM.
JVM is indeed a powerhouse and that’s where I’d go Clojure without hesitation. Scala might be an easier sell but I’m a lisp fan and have used Clojure in prod (was briefly CTO at FarmLogs a YC company where most of our core infra was Clojure) and it’s really an amazing language but it requires very senior engineers to do correctly.
For an online store (compared to, say, a live action game) so much of the content that you serve will be cached that regardless of your apps runtime a lot of the user experience will be defined by how effective your caching strategy is.
I think that Rails still offers enough compelling advantages for developer productivity to offset the (possibly) higher hosting costs.
Although they probably spend most of their time on the application layer, I suspect the most important and trickiest work was done to scale the database.
Yes.
At it's core, what they provide is (largely) read only set of products, and a (largely) append only set of purchases.
The only write operation that needs to lock the database is when you adjust the quantity of available inventory. You need to pay attention and think things through to do that with good performance, but it's not that complex. and they wouldn't be doing millions of sales per second.
Edit: To be clear, I agree that this is an example of distributed, high-performance which is why the comment made little sense to me.
Nothing will run at that scale on a single VPS. All companies will have a wide range of languages used.
If this is not Rails supporting high traffic then what do we need more?
What you seem to be getting at, isn't distributed systems, but the totally self-inflicted pain of a service oriented architecture
Having spent the last 28 years building distributed network-connected systems, this comes across as wildly obtuse.
The point is that there are orders of magnitude differences in complexity when scaling a system with few communications paths and little distribution of state across process or network boundaries as there is when scaling one with many paths and state distributed in many locations. We don't tend to start talking about distributed systems when you have a tiered stack of a horizontally scalable component sandwiched between a load balancer and a database even though in a very strict technical sense already that is "distributed".
Once you start adding message queues etc., then it certainly becomes more and more reasonable to talk about a distributed system, but there is there as well a distinct grey area if dealing with e.g. queues just triggering jobs in the same code base against the same database with respect to the intent clearly expressed by the original comment.
Put another way, ignore the word "distributed", re-read the original comment, and consider that irrespective of which label you're comfortable with, what the comment is doing is drawing a distinction between two classes of systems with wildly different complexity in the distribution of responsibility and state. Where precisely you draw the line is entirely irrelevant.
> What you seem to be getting at, isn't distributed systems, but the totally self-inflicted pain of a service oriented architecture
No, it really was not. This separation between basic 2/3 tier apps and systems with a more complex data flow pre-dates the SOA buzzword literally by decades.
https://shopify.engineering/software-release-culture-shopify
"We split off 2 things in to small services because it made sense" is rather different. I mean, I don't really care if you call this 'microservices" I guess, but it is different, right?
Or create organizational problems.
- "Who owns the flip-flop service?"
- "That's Bob's team, we fired them last summer!".
"That's Bob's team..."
So unless you're using managed services e.g. RDS you're going to be exposed to the same complexities as distributed computing.
Especially with the cloud where instances can die at any point.
You could turn that one Rails app into a complex microservices architecture and do a conference talk about it, and get a promotion. Then you can undo the microservices architecture, write a blog post about returning to the majestic monolith, and do another conference talk about it, maybe get another promotion. Abstract, de-abstract, bundle, unbundle, rinse and repeat.
It feels like a tragic situation that's nobody's fault, just the reality of human psychology being wired to reward the wrong things.
Otherwise known as RDD (Resume Driven Development) - https://rdd.io
We value:
- Specific technologies over working solutions
- Hiring buzzwords over proven track records
- Creative job titles over technical experience
- Reacting to trends over more pragmatic options
The only thing boring are the people on HN who are still having this played out monolith versus micro-service argument.
Whereas actual developers in the real world have moved on and realised that both are useful in different situations.
The cost argument about this monolith is just a straw to clutch at. Microservices are not cheaper than a monolith. Operationally or infrastructure wise. Logs, monitoring, tracing and what not for each Microservice.
Shopify probably looks more like what you ridiculed than not. we can guess that it's not one big team, and it's not hundreds of identical copies of this big monolith (but configured during deployment to run in different roles).
So what you are saying is that microservices are not a solution to a technical problem but a solution to organizational problems?
It made sense for Netflix, because they had a big Cassandra cluster, too much money, and a very picky organizational/hiring culture, and so on.
https://www.infoq.com/presentations/microservices-netflix-in... (https://www.youtube.com/watch?v=TOM6UhCetQ0)
I’ve interviewed twice and there seems to be a resistance from engineers both the times I argued that MySQL wasn’t the right approach. Sure, MySQL could possibly run any use case possible in the world if you throw enough engineering resources at it and highly optimize for that use case, but why do that if there is another engine specifically designed for your use case. The argument was “Shopify runs on MySQl and we can handle millions of queries… blah… blah…”.
It's fine to argue with interviewers, but the threshold where it indicates to both sides that you're probably wrong for that job is pretty low.
A whole lot of technical decisions we - me included - have very strong opinions about don't really have that much of a measurable effect on technical outcomes.
I once, many years ago, had a conflict with someone reporting to me because I refused to entertain rewriting our entire frontend in Rails. This was just after Rails was released, and we had a lot of PHP code. We had PHP code because PHP frontend devs were "cheap" and plentiful, not because any of us liked PHP.
I agreed with him about the preference for Ruby, and we used Ruby for other things (ironically all our Ruby use was on the backend), but he kept pushing on the basis of no insight into the relative market conditions for hiring at that time and on the basis of making assumptions about how "trivial" a rewrite was because he had no insight into why our then-current frontend had all of the capabilities it had which he'd chosen to ignore because he didn't know the roadmap.
He went to my boss - the CEO and co-founder - and tried to get me fired. The CEO went to me and asked if I wanted to fire him instead. I didn't, but I did have a rather serious chat with said developer about our respective roles, and how it was about more than technical preferences, and how he might get a lot further if he actually tried to work more constructively with me instead of thinking he had the clout to get me fired.
Said CEO hired me to run a development department again in my previous job, while he to my knowledge has never again worked with that developer - looking like a troublemaker to the wrong person can have long term effects.
And what engine is that ?
MySQL has been proven by Meta and numerous others to scale to ridiculous levels.
Looking at it holistically, MySQL is ALWAYS going to be the correct approach, if you already have an internal team of MySQL DBAs. Either your problems are small enough that it really doesn't matter if you use CSV files, MongoDB or MySQL or they are large enough and important enough that you want to stick with technology you know and understand, even if it requires 25% extra hardware.
We did a project where we where looking into OpenStack, which was objectively the technological correct choice. Factoring in training, ramp up cost and hardware, it made more sense to just pay VMware.
Without knowing you, I suspect the issue is in how you answer such questions, not whether or not that you're technically correct. I'd go with the route or presenting two options, the one you find to be "the right approach" and the one that fits into the company's current infrastructure. Coming mostly from the operation sides of things, I find that developers can be pretty clueless about the cost and complexity of operations. Often to the point where you wouldn't trust them to design anything unsupervised, because doing so would end up in an operations nightmare.
If not, how do you know it was wrong for them?
FYI Uber migrated from Postgres to MySQL https://news.ycombinator.com/item?id=26283348
This without context is meaningless. What is the cost in $$ and engineering time to scale to that level? Would a native image be able to scale to the same level at half the total cost?
Even if I would be a consultant for them I will first try to understand the current situation and then imply a native image will have half the total cost is a good solution. What if reducing the hosting costs will actually damage their speed of pivoting and adapting to changes?