How Pinterest scaled
read.engineerscodex.com
read.engineerscodex.com
The thing to strive for 12 years later should be how to scale to 100M+ users with 6 engineers. NAND and CPU core has an order of magnitude reduction in price / performance. 1-2TB DRAM is obtainable. MySQL or Postgres has a decade of improvements made, along with simpler, battle tested solution for scaling. Both Python and Ruby, Rails or Django has all the lessons learned from real world.
At least that is what should have happened.
Vertical scaling is more expensive without the benefit of failover. With modern clouds, Horizontal scaling is cheaper with commodity hardware and gives the added benefit of failover redundancy.
So maybe one of the core reasons for Silicon Valley's love for horizontal scaling is due to plentiful AWS credits.
I can both horizontally and vertically scale in AWS. I can both horizontally and vertically scale in my home office.
Horizontally scaling is just more cost efficient if you can dynamically adjust to load. Clouds enable that because they have massive reserve compute.
In the startup world, there infinite compute, but finite money. It’s cost savings.
If you're using cloud stuff like EC2 on AWS for cost savings, you're in for a bad time. It makes sense for stuff that you need to be able to dynamically scale on a whim, but once you know what kind of setup you need, it's almost always cheaper to go dedicated.
At AWS's data center AWS also buys big servers. But their advantage is that they only have big servers, and then rent it out in pieces. They don't care whether they rent you 2 48 core 96GB instances or one 96 core 192GB instances. Both options run as VMs on much larger servers and take the same resources, so both are priced the same (a machine of the same category with twice the resources costs exactly twice as much). Thus there's a much smaller benefit to vertical scaling.
If you compare actual prices on large-ish instance types between AWS and regular dedicated servers (rented in a DC, not doing your own thing) then AWS is very expensive. And the hopes of paying less due to better utilization somehow never pan out. AWS has lots of advantages, but price isn't one of them; and their price structure will influence the architecture you choose.
- Data migrations, schema changes, backfills, backups & restores, etc., take so long that they can either cause or risk outages or just waste a ton of engineer time waiting around for operations to complete. If you have serious service level objectives regarding time to restore from backup, that alone could be a forcing function for horizontal sharding (doing a point-in-time backup of a 40TB database while dropping some unwanted bad DELETE transaction or something like that from the transaction log is going to be very slow and cause a long outage).
- The lack of fault isolation means that any rogue user or process making expensive queries impacts performance and availability for all users, vs being able to limit unavailability to a single shard
- When people don't have horizontal scalability, I've seen them normalize things like not using transactions and not using consistent reads even when both would substantially improve developer and end-user experience, with the explanation being a need to protect the primary/write database. It's kind of like being in an abusive relationship: you internalize the fear of overloading the primary/write server and start to think it's normal not to be able to consistently read back the data you just wrote at scale or not to be able to use transactions that span multiple tables or seconds as appropriate.
IME vertically scaled replicas/hot-stand-bys are a lot more stable to operate if your requirements allow you to get away with it. OtoH you better already be prepared if/when you hit scaling limits.
Those big ass SQL boxes cost serious money.
That being said I'm partial of having fleets as small as possible to simplify operations. While trying different instance sizes, I found our bottleneck ended up being network. The largest instance types saturate the network first, and hence we can't use servers as efficiently.
Like millions of other devs, I have made a (decent) living out of this career, but you do wonder what might have been.
I think there is a new stack - a new Dev Manual - of what to do and build to keep it simple today (Hadoop is involved :-)
Of course this doesn't detract merit from the post. The devil is on the details, and there are many complexities other than raw scale.
Or are those 3 years mostly maintenance and the bulk of work was done in a few weeks at the start?
Serving ecommerce sites with a couple of Java or .NET servers, backed by Oracle or SQL Server instances, surviving heavy load days like black Friday.
See Stackoverflow architecture, for example.
I would hazard a guess the Pinterest usage model has a somewhat higher write-load, and is more sensitive to slow writes/non-propagating writes.
Do I think their arch is still overkill? Yeah, do I think equating them with SO and friends is reasonable? No.
The scaling issues we had with Tcl in the dot-com wave with our AOLServer clone, is one of the reasons why I wouldn't use something like Python.
There was a post on HN a couple of months back about Pinterest saving 2m/month on infra costs by swapping from Python to Elixir/Erlang. It copped flack in the comments because basically “ppphhh $2m that’s not even worth saving, how dumb, Python is good enough”.
When is anything worth doing then? Do people really think there's magic everywhere?
Historically I would agree. There are options now that are both easy to develop in and are blazingly fast (C#, Go, Rust, Kotlin to name a few).
The elephant in the room is that both Ruby and Python fell behind in relative ease of use and computational performance with respect to other ecosystems and instead of addressing pressing matters there is a lot of mental gymnastics and cognitive dissonance at play such as presenting false dichotomies.
I never experienced this mythical speed of development of Django or Rails, and I've been using both on/off for the last 15 years. Back then I was using stuff that was quite fast for development already.
But as soon as you have something moderately sized the slowness of languages like Python or Ruby becomes a problem for running tests, migrations, debugging, increased server costs, downtime, overhead in things like queues. I worked in some quite large projects though, so maybe for small stuff they work, but I absolutely hate them for writing big apps.
Like you said, the speed of development is IME fantastic with languages like Golang, C# or Node.js, and they don't suffer from the performance issues. Lots of companies local to me moving to Kotlin as well and my friends have only good things to say.
It would probably be very interesting to compare then and now.
The average tenure of a software engineer/developer is just 2 years.
People come and go and the next thing you know you're using 6 different database with services written in 5 different languages.
Instead if you kept the same 6 engineers around for decades you could probably scale to 100m with those same engineers.
It one of the reasons why it seems like Google can never release a product. They can't keep people around long enough to see it through.
The thing you seem to miss are the other common denominator.
Huge amount of money and unreasonably far into the future expectations of returns.
Means there is no short to medium term pressure to optimise for efficiency or returns, which means one of the fundamental element of good engineering environment is missing.
These companies build in a vacuum of limitations in term of cost and a vacuum in term of goals.
https://stackoverflow.blog/2022/04/19/whats-the-average-tenu...
The floor is 0. You can't have a negative tenure at a company.
So it is possible to avoid this, regardless of employees turnover.
Engineers get promoted for creating complex architecture. I explicitly remember my manager asking me in my yearly review for promotion "what is so hard about that" .
Managers get promoted for hiring more people to support the said complex systems.
They knew eachother so well and of course a side effect from that was that they produced absolutely astonishing things.
I was fended by these guys and the level of tooling set up for me not to fuck up, and their reviews, test strategies etc. was lean, direct and nutured by every single individual. They had built these skills as a team through years of hardship. It was quiet a Ln experience to be honest.
So yes, I totally agree with you on that people are what matters most and being in an organisation which hands down are backing it.
Scaling to 100m users is more about your users/product than your engineering IMHO.
If your product doesn't appeal to 100m users, you can't scale that high. It's not necessary to be ready for usage that's unrealistically high. And your team won't develop experience scaling if it's not happening.
Mega scaling gets more approachable all the time. Ebay had to shard their databases pretty early, because what they could fit on the biggest Sun machine to run Oracle they could buy isn't much. Now you can run a dual socket Epyc and get 256 cores with terrabytes of memory and petabytes of storage in one system. Might not be the best way to run your database, but if your queries aren't super awful, you can do a ton of queries per second with that much ram.
I think your theory is less theory and more fact.
Another problem is that engineers will use a new tech project at work as an opportunity to use a new, hot technology that you can then put on the resume for your next job. That's when you see over-complication.
How do I know this? Bc I have done it - I am part of the problem. I do not want to, but the job market is competitive and I know my employer will never counter an offer I get from a new company.
A lot of this would be solved if companies would try hard to keep their engineers. I would not have made some of the decisions I have made if I knew that I would have to maintain my own garbage. I do try to write good design docs and code comments at least.
I doubt it. Effort does not scale linearly with user scale, in my experience, primarily because features don't. This means that horizontal scalability is insufficient to go from 10M to 100M. And that means that more complex systems are required to support it--I can easily see a team of 6 very experienced engineers with deep domain knowledge of the application becoming overwhelmed as that scale change occurs.
There are people whose lives rotate around having different boards on Pinterest. They collect ideas for recipes, interior design, clothing, vacations and _everything_ on Pinterest. It's not a site where people click three links and come back in a week for their next three clicks.
This is so annoying people have created browser extensions to remove Pinterest links from search results
As a database person, I dislike reading this. I understand their need and priority to keep things running. But I would have loved it if they emphasized that doing something like this has so many disadvantages.
Normalization, referential integrity and a powerful query language bring so many benefits. I see young people oftentimes seeing relational vs. NoSQL as two totally valid opinions, like Coke vs. Pepsi. And not as one thing runs the world for decades and is perfect for >90% of use cases, and the other one is for niche cases and fast hyper-scalers.
I’m the guy on most teams pontificating about SQL and data-oriented code, but even I’ll concede that large scale live systems probably shouldn’t serve from an SQL database.
If that is the case, is a relational db still the right place to keep such data? It seems that if the goal is to store this kind of data in one or a few huge sharded tables, a different tool could be optimizing to do the job better.
It seems that I’m missing something…
how storage became efficient https://medium.com/pinterest-engineering/evolving-mysql-comp... https://medium.com/pinterest-engineering/evolving-mysql-comp...
Why? key/value data stores have consistent hashing which scales to planet scale built in. No need for sharding or caching or other bs that brings down slack and GitHub every other month.
I have had much better luck sticking to a plain old relational DB stack for those phases, because it’s fast to iterate, and then consider moving to single-table NOSQL when the system/DDD has “gelled” and traffic is starting to hockey-stick
Really, this is pinterest. They store 'boards' and 'pins', you may as well have a bit of fun with it. What use cases do they have for the features you list? They probably only used a database for write transactions?
I deeply value database technology and use it daily, but I don't think something simple like this needs to worry too hard about it.
The hard part is finding 11 million users.
I'd use .NET Razor pages, Dapper, Postgres, Kestrel on Linux?
What's amazing is that 14 years later and the major browsers are still hoarding the Pinterest masonry layout behind a flag!
https://developer.mozilla.org/en-US/docs/Web/CSS/CSS_grid_la...
Curious to watch https://www.infoq.com/presentations/Pinterest/ at some point.
[1] https://web.archive.org/web/20160203120655/https://engineeri... [2] https://github.com/lexical-lsp/lexical
66 MySQL DBs + 66 secondaries
59 Redis Instances
51 Memcache Instances
That’s interesting.For every 1 database, they also needed 1 redis cache + 1 memcache
This article sparked in my mind a bit-twiddling hack I implemented to get a ton more scale out of our single giant MySQL box that survives in their code base to this day, much to every one of their dev’s chagrin (crazy to think that 15 years later the weekend project I launched still has tens of thousands of monthly subscribers).
At Hatchlings, we had several hundred collections of Easter Eggs (usually 7-12 unique eggs in each collection). Each user would “open” and then hunt for a few collections at a time until they “finished” them. New collections would come out every week (to ensure there was new content daily).
As we scaled both in users and number of collections I realized that the “normal” SQL layout you’d use to represent this (a users table, a collections table keyed on user_id, collection_id, and an eggs table keyed on user_id, collection_id, egg_id) was growing rapidly (each user was adding new rows every day) and needed multiple joins in our “hot” routes (when a user searches for an egg we need to know which collections they have access to, and when a user clicks an egg we need to know if it was new for them and if it newly completed that collection) which were being hit tens of thousands of times per second during peak times. Additionally, every egg collected increased the user’s score.
So I implemented an optimization: store the collection unlocked and completed info compactly in the users table in bit fields. These were several 64 bit integers where eg collections_unlocked_1 being equal to 4 (binary 100) meant that the collection with ID 2 was unlocked. (collections_unlocked_2 equaling 1 would mean collection 64 was unlocked since there were 64 bits in 1). Your “active” collections were the bits that were 1 in an unlocked bit field but 0 in the parallel completed bit field.
This did a few things:
* Reduced the rows per user from total_collectionstotal_eggs to ~collections_in_progress (just in-flight collections) — about a 100x reduction in our total DB size
* Reduced joins to 0
* Since all the “common” egg collections were already completed by all the most active users we didn’t have to check the eggs table for the vast majority of finds because if a collection was marked finished we knew you already had all the eggs in it
* Made everything needed for the hot route compactly storable in a single memcached entry (which allowed us to greatly reduce writes because we would read scores and collections from memcached in the hot route and only write changes to the DB once/minute)
It was a great speed and scaling optimization… But it was really tricky to deal with and reason about… and a binary arithmetic error could completely nuke months of game progress for people. And we also had to remember to add more columns to the users table every year or two or we’d run out of space for collections in our bit fields (we forgot about this a couple of times and it lead to downtime which tended to happen during the most peak times where we had hyped that we were going to release a ton of new content all at once).
Because that's exactly how I use Pinterest.
And articles like this, about a presentation older than 10 years.
Our stack was similar than pinterest, django but postgres and redis. Team of 2.
Indeed, keep it simple. Cache things. Use queue, etc.
You can scale vertically a lot. Db perfs and disk space were our points of focus.
I assume since they're a viral / addictive platform the numbers will be higher. Plus they're not a blog or so, but rather a personalized social media platform, which makes stuff more complicated (you can't just cache their news feeds, they're all personalized).
It's easy to say meh, 11m is nothing, there are other platforms with more users (esp. since you don't bring up examples how you managed otherwise). But I think it's a big technical feat to do this with 6 people.
Basic servers are capable of handling hundred of the housands of those if you optimize your db and use a low level language.
We used python and an orm and we handled 1000 r/s on peak hours daily.
And we are talking user accounts, comments, votes, upload, transcoding, video streaming...
Todays computers are mind blowingly fast.
Afaik a transcoding session holds the entire CPU core, for example malformed gif avatar typically can lie down a single-server web-forum for a few seconds.
If they were once a good idea, those days have long passed.
Google Image search is almost useless if you don't block out crap sites like this.
https://www.quora.com/Why-do-some-movies-show-On-Netflix-in-...
(The irony of linking to Quora as evidence doesn’t escape me.)
They do.
> Google Image search is almost useless
So your issue is with Google and not Pinterest then.
It is just repackaged content from the original Pinterest presentation back in 2012:
This was an article meant to resurface learnings from something that happened a decade ago, with added images and a clear distillation so that somebody doesn’t have to watch the 45 minute presentation to understand what happened.
I’m sorry you didn’t like it.
But I hope you can see the value I’m trying to provide here, as someone myself who doesn’t really have the time to sit through hour-long presentations to learn something.
Parent comment was also valuable to me, for seeing how important it is to listen to others before jumping to conclusions about the worthiness of something.
I like the 11 million unique users a month early in the company -- simple architecture, popular programming tools and languages, small team, and likely enough ad revenue to pay the bills and get some earnings!
I wrote my code using Microsoft's Visual Basic .NET, ASP.NET, SQL Server, and one use of platform invoke. The code appears to run as intended.
I like the mention of DB (relational data base): I wrote the code using a free version of Microsoft's SQL Server. Right, it's only for development work, has some severe limits on DB size, gets expensive for a production version, also a pain since have to count processor cores, was glad to see Postgres and MySQL since have been planning to use one of those. I liked seeing the mention, and remark on power, of key-value stores (e.g., Redis) since I wrote my own key-value store, with all the data just in main memory, using two instances of the .NET collection class.
Now with TB (trillion byte) main memories am even toying with the outlandish idea of keeping nearly all the DB data in main memory with SQLite -- outlandish!
And, liked seeing the steps up in capacity. That there are likely better architectural ideas now doesn't disappoint me!
The outlines of the architectures of even huge Web server farms was really good to see. You mean Google, Facebook, Amazon, Microsoft, etc. have surprisingly simple architectures??? WOW.
You can’t please everyone.
Yes! Let's censor stuff we don't like! /s
But on the other hand, when I do want to use it, it's a great scrapbook and recommendation tool -- easy to capture images I find while browsing, and good at figuring out what else I'd like to see related to them on the same theme. A visual notebook with suggestions.
Things like:
- keep track of the bags I've been thinking about buying so I can see them all at once and find others I missed
- pin a bunch of home office photos I like, and then it suggests even more ideas (or garden ideas, or kitchen ideas, or workbench ideas...)
- pin all the recipes I've enjoyed making and that I want to make, and get even more suggestions that have the same dietary restrictions
- pin some things with a particular vibe -- let's say "motorcycle photos I think are cool" -- and get a better idea of the whole aesthetic or subculture or whatever they're from
For better or worse, it is really good at what it does. No argument that it pollutes search results, no argument that there are too many ads in the form of promoted pins, but underlying that, the recommendation algorithm is really useful.