Backend of Meta Threads is built with Python 3.10
twitter.com
twitter.com
So they created their own language Hack so allow a slow transfer from PHP to Hack. Which is basically the idea behind Kotlin.
Sometimes that actually meant rewriting the thing from scratch, and other times that meant adding something else as a middleware to prevent touching the original thing.
For the most public success story, React was a complete rewrite of the thing before it, for example.
It's mostly hyperbole but there's definitely a kernel of truth to it. I generally dislike rewrites and I wonder if FBs approach influenced me more than I realized at the time.
Why would a massively successful company toss a stick in their own spokes by tearing the product to ground?
I'm also very skeptical about big bang rewrites, but there are points where you need to migrate off certain tech. Ideally you can make that transition piece by piece though, but that also introduces its own set of problems (now you have two systems and the new one has to inherit some of the baggage of the old one in order to be compatible).
Those systems are well understood and well tested, so it's not cost effective at the moment to embark on a complete rewrite. Current AI coding systems are also very unreliable, so it's just a matter of time before those two meet on the graph and voila, another boom in moving the legacy world into safe languages.
Facebook was already at a scale larger than 99.99999 percent of sites before they had HHVM
Ruby is the exception to the rule.
Friends don't let friends start new Ruby projects in 2023.
Elixir? Okay, it's a nice language that is also dynamically typed. It has some advantages but good luck finding anyone with knowledge of how to program in it or has even used a functional language before.
I don't personally use Ruby these days but to say it's super behind is just silly. Golang is as similar to Ruby as C++ is.
That's what I mostly meant by saying Ruby is behind Elixir and Golang. And mind you, Elixir is strongly but dynamically typed and I still find it much better than Ruby.
However, Ruby also has a lot of creature comforts that Elixir and GO don't have. Maybe there wouldn't be a product at all with GO or Elixir. It's hard to say.
In Ruby you get nothing like that, all your function arguments are just variable names. This increases testing friction a lot. I've been all over the spectrum: from PHP and Ruby through Elixir (combining best of static and dynamic types IMO, though still flawed) to Golang and finally to Rust which is super strict. To me Elixir and Golang are close to perfect. Rust takes it too far and development velocity can suffer a lot until you become an expert (which can take quite a while).
Plug and play is nice but my opinion remains that it's oversold. Quickly whipping out prototypes is not the only virtue of a programming language (though technically that's a feature of Rails, Ruby's killer app, and not of Ruby itself).
Stuff like Elixir's LiveView and its imitators (like PHPx and I think C#'s Blazor?) are where things get better but since I am not interested in frontend, I leave that work to other people.
Still, at one point it's OK to admit the original tech is no longer as useful and migrate to something else IMO.
We can disagree about usefulness I guess. Ultimately Ruby has worked just fine. They didn't report massive issues, and the DB is going to be the bottle neck way before Ruby, so I just don't see an issue.
If I knew I was building a site for a massive load would I use Ruby? No. But I don't think it's a terrible choice either.
Not only that, the community felt the pressure to write several JIT compilers already.
So....
Is there any large company at all that hasn't invested in making their tool chain better?
You talk about it like they are a failure. Lol
Besides, you can even use optimized C and get none of the benefits (but still all the drawbacks) because of the algorithms you’re using or your database or something else entirely such as microsservices etc. making differences in language speed negligible.
I guess the point is the language speed on backends don’t matter as much as many other things (network, database, architecture setup etc.) and usually the gains you get with your language being faster aren’t high unless you’re in a tight loop or you’re like Amazon where even a few milliseconds gain is more money.
Exactly. And I am saying the following as a fan of dynamic languages but let's be realistic: this is practically their ONLY true benefit -- quicker time to market. They lose out on pretty much any other metric, with the exception of also slightly quicker iteration in the day-to-day work.
- Ruby on Rails prototype: 2.5 days.
- Golang's Gofiber: 4.5 days.
- Elixir's Phoenix: 7.
- Rust: don't ask (too long lol).
- Python's Django (made by a friend): 3.5 days.
Not to split hairs here but if you are in a situation where 5 days more for a prototype matter then I am not very sure you would have succeeded even short-term after. Putting your foot in the door can be extremely important, absolutely, but after whipping out a quick MVP you'd likely immediately stumble upon the next obstacle which is not guaranteed to be overcome.
As mentioned in my previous comment, the only argument in favor of the fastest-to-prototype languages could be that they allow slightly higher velocity day-to-day as well. But I've worked with PHP and Ruby for a long time, and I've worked with Elixir, Golang and Rust for quite a while now as well and again, that particular advantage of PHP / Ruby / Python (a) does exist, yes, but (b) is not as big as people make it out to be.
By your own admission here GO takes nearly twice as long (which is about what I've seen). So if you aren't building a trivial app but instead building something that takes 6 months in Ruby you'd be spending nearly 12 months in go. Sounds like a real significant amount of time if you are rushing to get to market.
They have a incredibly successful company. What did ruby do to hurt you personally.
This whole line is co oketely irrational. You want thier webapp written in assembly?
Nuance matters. I wouldn't pick Ruby for anything except scripts nowadays. You get absolutely zero guarantees and being able to whip out an MVP in 2 days is not as important as people make it out to be. Extremely often being able to have an MVP in 10-20 days is just as viable but then you will have to write less tests and have more guarantees from your compiler.
You mentioned Elixir. You literally suggested a language that was created 6 years after Shopify began.
You wouldn't pick Ruby, thats your choice. People are still making billions of dollars with it.
So, since you took up the mantle, what language should they have used 17 years ago when they started the company?
Not arguing with their results here. I am saying that they made it work despite how terrible it is. You have no guarantees about anything and you have to have much more tests than you would have in other languages just so you have basic stuff ensured. Not cool at all and I am saying this as a guy who made good money with Rails for 6.5 years. Not impressed to this day and I am happy I left it behind. It's good for scripting I'll admit, though nowadays I just learned bash/zsh better and use Golang for the same task occasionally.
> what language should they have used 17 years ago when they started the company?
Whoops, you are 100% correct that I ignored the historical timeline (Elixir didn't exist back then, yes), sorry about that, I derped big time.
That being said, all of PHP, Ruby, Python, Java and C# were indeed valid choices back then. But I've been part of successful rewriting efforts (biggest one was about 270k coding lines which I'll admit is much smaller than what Facebook and other big corps are dealing with) and to me the downsides of rewriting are overplayed:
- You need extensive tests? Well, you need them even if you never rewrite.
- You will now have two systems? So what, you already have a load balancer, you will just put a few more rules in it.
- Hard to find engineers for $new_language? That's true if you do it in its first 5 years of life but I've been part of growing ecosystems twice (Elixir and Rust) and it only gets easier with time. Solution: don't be a super early adopter, that's obviously too risky. E.g. both Elixir and Rust are beyond 10 years old at this point and are now a safer choice.
- Have to duplicate the old system's bugs? Tough nut to crack and I partially agree with this one but what my previous teams did was write down these bugs in the docs and made sure to start fixing them after the rewrite was completed. And in all cases fixing the bugs in the original language was planned anyway but was eternally kicked down the road.
--
My main point here is that it's OK to admit at one point that the tech you used for a while has outgrown its usefulness. I understand and recognize some of the downsides but to me they are overblown.
My entire complaint is just that you want to say it's outgrown it's usefulness. That is such an absolutist view. And the fact you are even bringing rust into this is crazy. I know C, C++, ASM, GO, and a few other low level languages. Rust is harder to learn than all of them, even ASM. And Rust in some ways can be a dumpster fire when you have to unwrap everything. A very intelligent person can take months to be comfortable in Rust or be similarly comfortable in GO in a couple days.
It's such a dogmatic view to say Ruby has outgrown it's usefulness, that's fine though you are entitled to it. Ruby isn't going away, it'll still be there and I'm sure people will continue to make money with it.
> My entire complaint is just that you want to say it's outgrown it's usefulness. That is such an absolutist view.
I don't speak for all programmers and IT companies -- and neither do you. We are both right at the same time due to the fact that the world is big.
I've been in super small teams (think 3 people) all the way to 50+ and stricter typing is more and more sought after the bigger a team becomes. It's very hard to code confidently the weaker the typing of the language is after a project gets beyond a certain amount of coding lines. Every dev must spend extra time to self-onboard in the context of others' work before being able to meaningfully contribute.
Of course this curve tapers off eventually and some teams become super well-oiled machines -- that's a fact. But we don't live in an ideal reality; people leave, some get sick and are gone for 2-3 weeks, others get reallocated to other projects. Things happen and the ideal stable-ish state of a team is rarely achieved.
Having stronger stricter typing gives peace of mind when coding. It does reduce prototyping (and sometimes the day-to-day) speed but from one scale and up it is worth it and that cost is quickly surpassed by confident and quick delivery of features or bugfixes in the mid-term.
> And the fact you are even bringing rust into this is crazy.
Oh I agree. Rust can be infuriatingly hard and slow to progress with. I've made the mistake to try and prototype one thing with Rust and I lost one small business opportunity because of that. Won't ever do that again. Believe me, I've experienced this craziness first-hand and I agree with you there.
> It's such a dogmatic view to say Ruby has outgrown it's usefulness. Ruby isn't going away, it'll still be there and I'm sure people will continue to make money with it.
Maybe it's dogmatic to you because you imagine that if I was right Ruby would be no longer widely used and since you are not seeing that, you thus think me wrong? Well if so, (1) the world is big and even if Ruby is being gradually shown the door it can take decades until you and I notice, (2) your assertion that Ruby is not going anywhere is 100% true but does not contradict anything I've said because there's a place for everyone in the IT sphere and that still does not mean people are not seeing cracks in Ruby's perfect image (people from Twitter and Shopify have expressed displeasure with it a good amount of times in the past... and now I regret I never kept the archive links because I always get asked for source and can no longer provide it -- sigh).
BTW many people came to Elixir from Ruby and said they're never going back unless they have to build an MVP in a weekend -- they would go back only for that.
> Maybe there wouldn't be a product at all with GO or Elixir. It's hard to say.
I am not claiming either way. Don't imagine me such an extremist, please -- I am not. I am only saying "OK you did it with PHP, Ruby or Python, the tech did the job okay-ish but it's beginning to suck for you -- why are you so averse to admitting that it's time for a change?". I got a very cynical answer and at the risk of you thinking me even more extremistic I'll name it: a lot of programmers love cozy jobs and they might hate what they do daily but they still love the wages and are not gonna rock the boat. Obviously I can't claim any percentages but I've been around and I've met plenty of such guys and girls. So IMO you should factor that into your analysis.
Inertia and network effects do exist and sadly they often have nothing to do with the quality of the thing they are carrying (think of all the sucky software we all must use if we want to communicate with others -- that's a good example of inertia / network effects not being correlated with the quality of software).
> We can disagree about usefulness I guess. Ultimately Ruby has worked just fine.
Due to tenacious teams. Not due to Ruby's technical virtues which are not that many. I admire people who can make anything work but I've also drank beer with some of them and they said that some days they think of committing suicide (I hope they were joking but they genuinely looked and sounded unhappy). I am talking PHP, Python, Ruby, Javascript.
> They didn't report massive issues, and the DB is going to be the bottle neck way before Ruby, so I just don't see an issue.
I remember the times in several apps I consulted for, long ago. We had good metrics systems in all of them and we had DB requests go from 5ms to 50ms and ActiveRecord itself (when subtracting the DB times) has routinely eaten 150ms to 250ms, sometimes more. I am aware that these issues have been fixed a while ago but it did leave a sour taste in my mouth and I can't just non-critially accept a claim that the DB dominates every latency in Rails world. Maybe it's true nowadays but during the dark times of Rails 2.x and 3.x it definitely was not.
Now Elixir... the last companies I contracted with we were talking 10-100 microseconds of Elixir code and 5+ ms of DB requests. Pretty neat.
...And then I consulted for two Golang projects where we had something like 100 to 400 nanoseconds of app code and 5+ ms of DB requests. Insane.
So I don't disagree with you on the concept but I do disagree somewhat that Ruby / Rails is not a resource hog. At one point it definitely was and it was humanly noticeable even.
But I can concede that nowadays this is very likely no longer true. I got no beef with any tech, I am simply a guy who is always looking for something better and thus I don't get attached to any tech. Elixir has served me very well in the last 7 years (together with Golang and Rust sprinkled in) but if e.g. Rust gains Erlang's / Elixir's transparent concurrency / parallelism abilities and fault tolerance and speed of development then I'd have zero qualms abandoning Elixir.
> So, wait, all apps are trivial in nature and can be built in less than a week?
Come on now, you are starting to sound like you want to misrepresent what I am saying. :) I was talking about quick MVPs / prototypes. Of course beyond those everything else is very different.
> So if you aren't building a trivial app but instead building something that takes 6 months in Ruby you'd be spending nearly 12 months in go.
You assume the curve of Elixir or Golang is linear; it's not. It's an asymptote.
To give you a contrived example:
- Month 1: Rails app at 20%, Elixir app at 10%
- Month 2: Rails app at 35%, Elixir app at 25%
- Month 3: Rails app at 45%, Elixir app at 40%
- Months 4 to 6: Rails app at 55%, Elixir app at 60%
- Months 7 to 9: Rails app at 80%, Elixir app at 90% and starting to work on final touches.
And so it went the few times I had the privilege of witnessing parallel Rails and Elixir development (including rewriting a few Rails apps to Elixir's Phoenix). The Rails guys knew their stuff pretty well and I liked working with them -- but even they admitted that they are regularly blocked on too many checks in tests or in the controllers themselves due to much weaker typing. (Though there were other reasons as well but hey, this comment became an essay already.)
---
I understand that you are skeptical. A lot of us cope with the fear of missing out by simply denying there's something to miss out on. Plus we can't be everywhere all at once so we eventually find our own corner and become experts there. All of that is completely fine and I am not judging; we can't all chase some theoretical perfection, plus when we hit 32-35 y/o the reality of "this is still just a job even if I like programming" sets in and we are all very excused for not researching and knowing every single alternative of doing things.
If I am reading you correctly, you disagree that the older tech (Ruby included) has outgrown its usefulness. Ultimately I don't think we disagree with each other because we are both right at the same time: maybe where you work and the people you communicate with Ruby / Rails are still deemed super good at what they do, they get the job done, the product brings enough revenue to drown out the drag that Rails can be and everyone is happy. Cool, more power to you. I however changed sub-careers in programming several times in a row and I am solving for tangibly different problems than what most Rails companies I've met back in the day did. And I've had a lot of financial success with Elixir, Golang and Rust, and my customers were super appreciative of the work done. Even now I am tutoring people who have 30+ years of programming experience and they are very happy with Elixir in particular.
We being in different bubbles is 100% fine. The world is big and rich and interesting. I am not disparaging your choice. Hopefully I am offering you an alternative take from another vantage point instead.
I'm just not going to respond to this because it's completely disingenuous when you make statements like that.
You are free to go on with your irrational views and I'll go on ignoring them.
The part you quoted was the only piece that could be interpreted as seeking conflict -- though it had a more charitable interpretation that you decided to ignore -- and you latched onto that.
Oh well, I tried. Let the future readers judge for themselves.
Being able to handle spikes and iterate quickly is probably more important.
In other words, Facebook felt compelled to write a compiler when they had hundreds of millions daily active users.
They're using a derived language called Hack.
> Does the compiled code use a garbage collector? Or reference counting?
Not sure but it's open source so I'm sure you can dig up the answers one way or another: https://github.com/facebook/hhvm
It's not the end of the world because they have deep pockets and a problem domain where preserving the interface (I.e., keep everything still running on PHP) and optimizing the infrastructure that provides it is more cost effective and least disruptive than switching to a tech stack that is more performant.
Yea actually. Bash/CGI could handle plenty. If thats your most comfortable language then go for it.
About 80% of the code I now write is in bash, because usually the problem ones trying to solve has a trivial solution if you just use existing tooling in creative ways.
At one point I wake up two hours later and realize I could have finished this in 15 minutes with Golang or even Ruby.
Absolutely shocking.
Some worthwhile references:
[1] https://engineering.fb.com/2016/08/31/core-data/myrocks-a-sp...
[2] https://research.facebook.com/file/529018501538081/hhvm-jit-...
[3] https://research.facebook.com/file/700800348487709/HHVM_ICPE...
Also, Facebook used to do ahead of time compilation (HPHPc) but eventually HHVM managed to outperform it.
But… how? Considering all the transactions in flight, and everything? And did you ever Test disaster recovery with that setup?
I’ve worked on relatively big projects, but FAANG engineering is like some entirely different field of software engineering. Fascinating.
> Considering all the transactions in flight, and everything?
If I remember, we used --flush-logs --master-data=2 --single-transaction, giving it a consistent point-in-time dump of the schemas, with a recorded starting point for binlog replays, enabling point-in-time and up-to-the-minute restores. Nowadays you have GTIDs so these flags are obsolete (except --single-transaction).
--single-transaction does put extra load on the database—I think it was undo logs? it's been a minute—which caused some hair-pulling, and I believe they eventually moved to xtrabackup, before RocksDB. But binary backups were too space-hungry when I was there, so we made it work.
Another unexpected advantage of mysqldump vs. xtrabackup, besides size, was when a disk error caused silent data corruption on the underlying file system. Mysqldump often read from InnoDB's buffer cache, which still had the good data in it. Or if the bad block was paged back in, it wouldn't have a valid structure and mysqld would panic, so we knew we had to restore.
> And did you ever Test disaster recovery with that setup?
Yes! I wrote the first version of ORC. This blog post is from long after I left, but it's a good summary of how it worked: https://engineering.fb.com/2016/10/28/data-infrastructure/co...
It wasn't the best code (sorry Divij)—the main thing I'm proud of was the recursive acronym and the silly Warcraft theme. But it did the job.
Two things I remember about developing ORC:
1) The first version was an utter disaster. I was just learning Python, and I hit the global interpreter lock super hard, type-error crashes everywhere, etc. I ended up abandoning the project and restarting it a few months later, which became ORC. In the interim I did a few other Python projects and got somewhat better.
2) Another blocker the first version had was that the clients updated their status by doing SELECT...FOR UPDATE to a central table, en masse, which turns out to be really bad practice with MySQL. The database got lock-jammed, and I remember Domas Mituzas walking over to my desk demanding to know what I was doing. Hey, I never said I was a DBA! Anyway, that's why ORC ended up having the Warchief/Peon model—the Peons would push their status to the Warchief (or be polled, I forgot), so there was only a single writer to that table, requiring no hare-brained locking.
It was a lifetime ago I ever did DB administration (postgres in my case), but the write-ahead-logs being replicated out independently was extremely important for point-in-time recovery, such that you could always take the latest backup, zip forward through the WAL, and recover to any arbitrary point in time you want, so long as the WALs were available. I wonder how much something like this would have been done at FB scale.
I'd think the same or bigger performance benefits can be reaped if you just start with stuff like OCaml / Haskell / Golang / Rust from the get go.
Arguably picking almost any other tech would have worked as well because they would have doubled down on it as well and made sure that all the pieces that they need are in place and are working.
Thanks for elaborating, though I definitely don't share your conclusion.
Have they really done this for OLTP workloads running in mysql? I know they do it for OLAP though.
Ultimately these these choices are all about trade-offs. Maybe python is fine for them, maybe they've built themselves into a corner. Time will tell.
Plus, instruction and cycle counting is low hanging fruit compared to memory latency. You can cache-optimize a program in any language so long as the memory representation of some data is relatively transparent.
That's a very valid question but in the case of Facebook they already have them so why not use them for that?
I mean yeah, they made their choice -- use HHVM and it likely served them very well. I am just pointing out that in their case sourcing extra (or even any at all) C++ devs is a non-issue because they already have plenty.
Fully agreed with your memory latency remark.
And I have definitely seen projects fail due - in part - to language choice. Of course projects can succeed in almost any language but that doesn't mean the language choice is irrelevant.
Just Google "python vs c++ speed" and you will find hundreds of examples (or "python vs rust speed" - Rust is essentially the same speed as C++).
Here's the first result - they got a 25x speedup:
https://towardsdatascience.com/how-fast-is-c-compared-to-pyt...
Technically choosing C++ over Python will save you several orders of magnitude more than nanoseconds.
Though C++ might be a bad example. I'd replace Python web app with Elixir or Golang.
Asking because I’m curious about these early stage mega-tech companies that are working at a scale I can’t fathom.
Facebook was founded in 2004. Whatever fancy tech you're thinking that would fit the round hole, it didn't exist then. They also never thought it would serve a couple billion of users one day, so they built it in whatever they had and knew at the time.
Same for Instagram and Python (they started with Django).
No idea if this was a credible leak of actual 2007-ish FB php, but it doesn't look like what you're describing.
"It's running on Instagram's #Cinder fork that includes a JIT, lazy-loaded modules, pre-compiled static modules, and a bunch of other interesting tweaks against vanilla Python 3.10"
From the Tweet.
Personally, I don't care what language people use. We all know that if you have money, you can just throw in more servers. Until your CFO and CTO decided that the next big thing is to do "Cost Saving" (in the eye of recession).
I'm surprised in 2023 people are still debating mainstream programming language "PROD" readiness.
It's like a "cool kid" competition.
Definitely not true at Facebook scale. You need to be smart, fast, _and_ throw more servers at it.
We both know the issue has always been how you store the data. Be it your main data storage, or the need of multiple storages for different use cases, etc.
The app server itself can be implemented in any languages. It's just a data orchestration + transformation logic anyway...
This is a huge loss for the world, and makes me a bit angry.
Edit: I had falsely assumed that because GP missed out on the JITing etc. described in the linked tweet, that Twitter was still inaccessible to the open web. However, trying again now, Twitter is again viewable without an account. So Twitter's communication about needing an account to view it was lacking, but so was my diligence.
Yesterday, I read an article about Tour de France stage and it linked to about 4 tweets with videos which are not playable without an account.
I just hope this gets noticed by mainstream media and they change their practices instead of assuming that everyone has a Twitter account.
That’s off by multiple orders of magnitude, if you believe their reports of 30 million sign ups in under a day.
https://www.theverge.com/2023/7/6/23786108/threads-internal-...
https://www.techempower.com/benchmarks/#section=data-r21&l=z...
the claim is usually more about how the maintenance costs are just way higher versus a statically typed language
But it’s not “just Python”.
The OS state that needed to be loaded is not Python.
The network gear firmware is not just Python.
The databases are not Python.
Where it becomes just Metas Python is after a whole lot of other necessary work is done by other smarter people (the ones who master the physics of building the machines), not just some dweebs who got pulled the internet onto the host.
These articles are just primate rage bait
Why would they use Django? I did some small projects in it but I assumed it wasn't very fast and wouldn't be suitable for a big app like this. I would like to know the pros and cons. Why didn't they build up something in C++ or Rust? Won't python limit the speed of responses despite being build the hard stuff in compiled languages? Sorry for being so naive, I am an amateur.
Personally evren I prefer php over python.
I get that Facebook is still a huge success, but I do find it telling that they opted to put Threads under Instagram, rather than Facebook.
It's Facebook by Meta. Instagram by Meta. Why isn't it Threads by Meta?
Switching from hack to php and viceversa should be extremely easy.
That being said, they chose python because Instagram is written in python, they are probably using a lot of existing libraries.
What percentage of python webapps do you think are hitting this as their latency and throughput limit?
(Assuming effective DB use of course, i.e. not doing dozens of DB roundtrips to server a single result or getting megabytes of data and filtering on the client etc.)
(Heavily modded, run on a custom Python JIT, and using an extremely custom FB-developed database, also used for IG and FB.)
It's Django because IG was originally written in Django back in the day. FB's general approach to scaling is keep the interface generally similar, and slowly replace things behind the scenes as necessary — rather than doing big rewrites. It's seemed to work pretty well.
Ultimately the language used for the web server isn't a huge deal compared to the performance of the database, for these kinds of massive social apps. Plus, well, they do have the custom JIT — but that's very new, and when I first joined IG in 2019, we were running vanilla Python in production.
Thanks also to every other answer I received for my question, I appreciated all of them.
Why Django? Because that's what it originally was. Same with YouTube frontend btw. The apps just grew and grew.
Some will use `python manage.py runserver` in production, and they are using the defacto wrong set up. Don't ever do that.
It’s millions of lines of code, you can’t just change it to be Java one day.
Is this just a Flask problem, or does Django have the same issues?
If you want to purely make APIs though, give FastAPI a try.
The ORM and django admin are killer features out of the box that the other frameworks don't have. I will say though that FastAPI is really nice, especially if you need async support. However, I have found that using django ninja [1] adds a lot of nice to haves that FastAPI has to django that makes it much more fun to use again.
There are some footguns around N+1 queries, but they are general to all database interfaces and pretty easy to avoid IMO.
I know django doesn't have that "shiny factor" to it these days - but it's very reliable.
> mixed messaging on best practices for scalable apps
The WSGI stuff can be kinda confusing and is used across a lot of python frameworks including django and I think flask?
My advice for "simple scaling" is to start with a separate Postgres instance, and then use gunicorn. Use celery immediately if you have any "long lived" tasks such as email. If you containerize your web layer, you'll be able to easily scale with that.
Finally - use redis caching, and most importantly - put Nginx in front! DO NOT serve static content with django!
> the ecosystem feels overbloated with vapor ware extensions.
This still exists to some degree for some more niche stuff, largely because of it's age. Although impressively they'll generally still work or work with minimal modifications. It's popular enough and old enough that most normal things you'd want to do have decent extensions or built in support already.
The current state of the art is apparently Prisma, but it covers only a small part of the full picture.
Using Django is probably the reason why I can stand using SQLAlchemy, it's way to complicated for everything I do and it's just not a nice an experience.
Better than Entity Framework? Consider me impressed.
Entity is way more flexible, in the sense that you can use in any .Net project really. The Django ORM have no value outside Django, I have yet to anyone use just the ORM, without the rest of the framework.
Such a strong assertion!
I've used multiple ORMs (not Django's), including some of Haskell's type-safe ORMs (e.g. Persistent and Beam). I could not imagine going back. What makes Django's ORM so great?
Anyway, you can use the Django ORM without Django :)
Django ORM is tolerable on small projects, and quickly gets in the way on anything a bit more complicated.
Hey, quick question from a relative newbie who is currently trying to solve this exact problem.
Besides Celery, what are good options for handling long-running requests with Django?
I see 3 options:
- Use Celery or django Q to offload processing to worker nodes (how do you deliver results from the worker node back to the FE client?)
- Use a library called django channels that I think supports all sorts of non-trivial use cases (jobs, websockets, long polling).
- Convert sync Django to use ASGI and async views and run it using uvicorn. This option is super convoluted based on this talk [0], because you have to ensure all middleware supports ASGI, and because the ORM is sync-only, so seems like very easy to shoot yourself in the foot.
The added complication, like I mentioned, is that my long-running requests need to return data back to the client in the browser. Not sure how to make it happen yet -- using a websocket connection, or long polling?
Sorry I am ambushing you randomly in the comments like this, but it sounds like you know Django well so maybe you have some insights.
---
[0] Async Django by Ivaylo Donchev https://www.youtube.com/watch?v=UJzjdJGS1BM
Celery is mature, but has bitten me more than anything else.
For scheduling, there are many libraries, but it's good to keep this separate from Celery IMO.
For background tasks, I think rolling your own solution (using a communication channel and method tailored to your needs) is the way to go. I really do at this point.
It definitive is not using async, I think that will bite you and not be worth the effort.
Huey is worth a look.
I usually use `asyncio.create_task` from async views for small, non-critical background tasks. Because they run in a thread you will lose them if the service crashes (or Kubernetes decides to restart the pod), but that's fine for some use cases. If you need persistency use Celery or something similar.
Django combined with an async-ready REST framework such as Django Ninja is very powerful these days.
If that's only 99% true, you might want to investigate other options.
I've mostly used Node, Java, and Rust, with about 12yrs experience now. Only about one year with Python and Django recently. I am so much more productive than anything else I've tried. Using anything else feels like a mistake for web apps or CRUD apis. I also use type hints.
I use Django + django-ninja + django-unicorn for dynamic UIs.
Building sidewaysdata.com with django right now! Govscent is in Django too, the ORM is nice to combine with AI stuff already in python.
I'm also following the Ash framework which looks promising.
It's a different strokes for different folks sort of deal.
Django is very opinionated, in a way that is "eventually" correct (i.e. they might have had some bad opinions many years ago, but they have generally drifted in a better direction). If the opinions don't align with what you're building, the escape hatches are not generally well-documented or without penalty.
Flask is minimalist and flexible. Once you find a "groove" to building with it that fits your sensibilities, it's quite, quite nice. That being said, the most recent versions of flask have excellent documentation, imo, and the tutorial is a bit more opinionated in a highly productive way. The "patterns" section of the docs is also super useful for productionizing your app.
Personally, I prefer the combo of Flask+SQLAlchemy, and eventually Alembic once you decide that's good for your app's maturity level. I respect Django a lot, I just enjoy the explicitly "less magical" aspects of a Flask stack, which is an opinion-based trade-off imo.
If so, do you have any recommendations or suggestions for someone getting their feet wet on this?
Some of my team mates work on driving internal adoption of Cinder, rolling out new versions of Python everywhere, improvements to our build and packaging systems, supporting other Python tooling teams, etc. There's a lot of cross-functional project work, and our primary goal is to improve the Python developer experience, both internally and externally wherever we can.
It's probably hurried along python code and fairly simple horizontal scaling. Some population of individual instances are probably pegged, and random requests are frequently landing on their backlogged request queues.
This seems to have been rushed out the door to capture a unique market opportunity, and they'll clean up the engineering once they get engagement.
So not entirely just python 3.1.
Great work, much appreciated.
Python doesn't go by semver rules. Semver isn't a definition for all software project versioning systems everywhere.
"To clarify terminology, Python uses a major.minor.micro nomenclature for production-ready releases. So for Python 3.1.2 final, that is a major version of 3, a minor version of 1, and a micro version of 2.
* new major versions are exceptional; they only come when strongly incompatible changes are deemed necessary, and are planned very long in advance;
* new minor versions are feature releases; they get released annually, from the current in-development branch;" — https://devguide.python.org/developer-workflow/development-c...
European Russia is 40% of the total European landmass. 15% of its total population. That's 110 million Russians in only the part that is inside Europe! Germany is at a low 80 million. Total population of EU is 445 million. Besides, numbers doesn't change because you are from the EU and disagree. I'm also in the EU.
However, I have a much larger following on Twitter, so it was the tweet that got picked up by someone here on HN.
We'll see how this looks like in twelve months.
Edit: Someone kindly pointed out that you can see the content in the browser without an account from following a direct link. (You just can't have the app installed or it will open that and drop you on the login page.) This makes me a lot happier with the product!
I was seriously considering trying it out, but if there’s no search engine visibility, there’s no point. Most of the searches for my real name pull up tweets.
Hopefully it’s just a missing feature that will be added later. Even TikToks are linkable, and those have very little google visibility.
Edit: Someone kindly pointed out that you can see the content in the browser without an account from following a direct link. (You just can't have the app installed or it will open that and drop you on the login page.) This makes me a lot happier with the product!
I hadn't seen one of these deep links before I installed the app. After I installed the app they just opened the app but did nothing because I don't have an account. But I just uninstalled the app and now I see. This makes me a lot more positive on this product!
They just extracted Instagram comments and threw them in a sea of whitespace. No thought is given to anchoring content by importance or context, nor is there any esthetic apeal to it.
So until it either solves its privacy issues of becomes part of the fediverse, audience is crippled.
Edit: Someone kindly pointed out that you can see the content in the browser without an account from following a direct link. (You just can't have the app installed or it will open that and drop you on the login page.) This makes me a lot happier with the product!
In fact a Thread thread has already been on the front of HN.
Still, mostly true.. read-only is just dumb. Nice to be back in 2008, when mobile-first turned into mobile-only.
The entire point of this thread was that this is a Python/Django web app, but it's not actually a web app at all, so what's the point of pointing out any of this?
But there's no broader browser yet it seems.
Instagram can be used in a desktop browser just fine.
Twitter removed that behavior under Elon, then developed something even more restrictive. Great.
Neither are good, and unfortunate that tumblr has recently ish copied this (for default css at least). Real shithead behavior.
Unfortunately, Facebook looks the best out of all of them. If a restaurant posts their specials on Facebook, I can see them. Basically not true for other social sites.
Based on that, I think it's safe to assume that a full featured web version of Threads should be up within a year. Threads already has some web presence (you can view threads and profiles).
Get it? Threads, threading...? OK, I'll see myself out.
Python calling a concurrent c thing is super common. For more complex things there is multiprocessing, but that essentially is simplification of IPC.
It’s still perfectly acceptable for Twitter users to discus other platforms on Twitter.
We need a social network that removes the noise from the conversation
Its not open to the web yet. Supposedly, it will be
Wordpress wouldn't be this successful if this approach didn't work. After all, without caching it will happily take over 20 seconds to render the main page.
Would you recommend using Cinder for stack made of - Django - Cython - Numpy
Thanks!
For example: you start a main process, warm it up with a few requests, run the JIT compiler and then fork off worker processes to handle the main chunk of traffic.
As of now, it requires hand-tuning to get the best possible performance.
In terms of use cases, Cinder does the best when faced with "business logic" code (lots of inheritance, attribute lookups, method calls, etc). It can speed up numerical computations too, but you're probably better off using a library if that's the majority of the workload.
Is Cinder something that could help optimize real-time streaming? We had a UDP stream and then through multiple gstreamer and nvidia deepstream magic (which I believe the senior dev implemented in Python) we perform some ML inference on the stream in real-time.
However, latency is a major issue here and to get to our MVP we didn't really prioritize optimization, as is tradition.
So now I'm wondering if Cinder as something that can be used to optimize real-time data streaming is a thing or whether me asking this just shows I don't understand its use case.
Either way, thank you advance for your insight.
(Also we used Django which I am now wondering if I should have switched out for FastAPI, but that's a separate question)
Because it's adding well over a million new users each hour which is pretty staggering.
But still there probably is all sorts of juicy details that would be neat to hear about
But I think on net it was still an easier problem than scaling a startup product from zero, because the tooling and processes to support huge usage is already so mature at these big companies.
A few million new people is less than 0.1% of the total. It's entirely within expected variance.
My naive assumption is that they have some neat simulated user/load system to help identify any scaling issues that might arise when launching a service, and the extreme % growth. That 3.8 billion was spread over 15 years.
Surely it makes massive difference to hardware utilization?
I like Python and javascript but there’s perfectly fine tools around now for building systems that allow one server to do a lot more.
They say you should not prematurely optimize. Equally you should not be prematurely unoptimised.
It's better, but not the limiting factor.
I’d use the analogy of a Japanese bullet train as the compiled language and a diesel train as Python.
Python is my primary language along with JS but if I was building a new service for Meta I wouldn’t use either at the back end.
Python is likely faster to push out an app like this than say rust or even go.
Also why spend time optimising something that will be dead in a few months?
Instagram stories crushed Snap’s DAU, Reels is gobbling up TikTok user-minutes…bleeding other services is what Meta excels at
It loads a javascript-driven animation of what appears to be a galaxy, any not much else. There is a QR Code at the bottom right corner that appears to be invalid.
Not sure what I'm missing. 10MM+ signups through what mechanism? There is no signup form as far as I can tell. Do they restrict full content access for linux-hosted browsers?
edit: after saving and reversing saturation, I've resolved the QR Code to https://www.threads.net/download/redirect, which only offers binary download options, and only for mobile platforms.
Where is the web interface?
But that's really odd because most mobile MVPs are either just PWAs or some React Native shit compiled to binaries.
Why the hell wouldn't they just release it on the web?
I'm a data engineer trying to teach myself React so I can build a full stack app using some APIs and it's certainly a good exercise, but I could also certainly build that in Python. Your app looks great and I'd like to explore that more!
P.S - fellow runner and cyclist and out and I'm checking out the app more now.
Its an interesting dynamic: popularity draws attention, attention helps adress pain points, this makes the platform more attractive and the popularity increases...
No point in supporting that user-hostile tire fire.
This could have been on Mastodon.
Of course, who knows what it'll be doing tomorrow. I'd definitely think it's probably time for HN to prefer Mastodon (or Threads, I suppose) links, where a post is available on multiple microblogging things; Twitter has had any one of about five behaviours for non-logged-in linked tweets (work properly, work but without replies, spin forever, redirect to login, generic error page) over the last week, depending on when you try.
(Mastodon link works.)
Of course he can. And he already does I assume, considering his other remarks in the comment. So including him in the 'we' does not really make sense ... if you consider it a question. Which it is not.
> "Can you ban..." who is "you" in that case?
That would be us dude, but not him as there is no reason to ask him, his position is clear.
So he is telling us what to do, dressed up as a question.
It is not a question.
Geez, it is you who is the authoritarian for coming so hard at the person who asked a pertinent question.
Geez.
Don't you find it ironic that you are proposing we do on HN exactly what Twitter did, which is probably why you suggest not linking to Twitter?
I would say that the amount of tweaks that prop up Python is no cause for celebration.
Too weak for someones hobby project, strong enough for global scale.
Team best practices, enforcement of coding style and technique, good project management, infrastructure, and team cohesion etc.
And weeding out bad engineers who think that "switching to a new [language|framework|religion]" will solve all problems.
10+ years ago Facebook itself did a remarkable job of scaling up, with hardly any outages, a PHP LAMP stack thingy to one of the hugest traffic websites in the world. Meanwhile back then Twitter mucked around with every novel technology they could find or invent, and had constant outages.
Also Google didn't "switch to Golang." There's plenty of C++, Java, Python there. I worked there for 10 years and encountered only a handful of Go projects. Lots of cloud services in C++.
Would I personally choose Python for a project? No, I don't like it, and I work in Rust full time because I prefer it. But if I worked at Meta / Instagram and had a team of Python engineers and existing libraries & infrastructure, this would be the right approach.
I think this is just an Instagram skin with some feature flags.
My personal suspicion is that the secret sauce is simply to have a sufficient budget to hire an army of code quality engineers. I could be wrong, though.
The first one is a big productivity booster as it shows you bugs before you commit them.
Here's a post about this, not exactly new but still describing the general principles very well: https://instagram-engineering.com/static-analysis-at-scale-a...
Nah, asyncio on uvloop is plenty fast.
That being said, my personal experience suggests that both are really hard in Python, and not independent.
Sadly, all the attempts I've had with static analysis in Python screamed that while the language and tools make a valiant effort at supporting a reasonable set of annotations, 8-9 years after PEP 484, the libraries are simply not yet ready for it (not even, in many cases, the standard library).
Unit tests, staging environment, effective log aggregation are all important tools, of course.
When people are hyper critical of Python I feel it’s typically they haven’t seen real professional Python before and just throw scripts around… if you work with some hardened Python pros you pick up the tricks really quickly and it’s very enlightening.
I’ve always struggled to find all of those tricks in one online resource personally.
My questions remain, though.
You have years of experience but couldn't point to anything in particular?
Heck, I like python and I can complain about dynamic typing issues in for loops, or that they are adding features like generators/decorators that make code more difficult to understand, which goes against the zen of python.
(But I still think python is great)
Those features are 20ish years old which made the word "adding" seem a little weird. Interestingly the Zen of Python itself is not much older than generators.
The biggest pain with it is of course refactoring. It's tedious, but I think still comes out ahead in terms of productivity for high level stuff.
The same is true here. They spent so much money getting Instagram to scale. They created Cinder to try and fold their hacks/tweaks back into Python. Spinning up some Rust/Java/Nim/Zig/Pony/Whatever stack sounds like it'd be fun... but spinning up a stack with which you're already VERY familiar, sounds like a money-maker.
I'm totally agreeing btw, I realize this came off like it might be a counterpoint.
Something, something privacy regulations.
It's not only available in the U.S. It's available in 100 countries. There's more to "global" than the E.U.