Why do we need Flask, Celery, and Redis? (2019)
ljvmiranda921.github.io
ljvmiranda921.github.io
Fyi - python ASGI frameworks like fastapi/Starlette are the same developer experience as go. They also compete on techempower benchmarks. Also used in production by Uber, Microsoft,etc.
A queue based system is used for a very different tradeoff of persistence vs concurrency. It's similar to saying that the usecase for Kafka doesn't exist because go can do concurrency.
Running "python -m asyncio" launches a natively async REPL. https://www.integralist.co.uk/posts/python-asyncio/#running-...
Go play with it ;)
Forget email, say you have a app that scans links in comments for maliciousness. You rely on an internal api for checking against a known blacklist, which follows shortened links first, and an external api from a third party. You want the comment to appear to submit instantly to the poster but are comfortable waiting for it to appear for everyone else. What are your options?
You could certainly use message queues and workers. If you’re cloud native maybe you leverage lambdas. Maybe you spin up an independent service that does the processing and inserting into the database in the background, and all you need to do is send a simple HTTP request on an internal network.
Your solution depends on your throughout requirements, the size of your team and their engineering capabilities, what existing solutions you have in place. Everything has its pros and cons. Pretending that celery/redis is useless and would be solved if everyone just used Java ignores the fact that celery and redis are widely popular and drive many successful applications and use cases.
More off-topic (or, rather, on-topic), I find lambdas great for things like a static website that needs a few functions. I especially like how Netlify uses them, they seem to fit that purpose exactly.
Me too! It makes me irrationally angry when people regurgitate linguistic clichés. I was already mad with:
"python does not have a go-like concurrency story"
when it would be enough (and 1000x less cringe) to say:
"python does not have go-like concurrency"
I think these mindless clichés make language really ugly and dysfunctional, and even worse they are thought-stoppers, because they make the reader/listener feel like something smart is being said, because they recognize the "in-group" lingo. In my experience, people get really offended when you point this out. It's kind of an HN taboo to discuss this. Which is also interesting in itself.
Going forward we should pay more attention to our communication use cases. Btw: I wonder if we can stack several of these clichés. For example: "leverage" + "use case" = "leverage case".
I've found it's a taboo to discuss anything even slightly personal. People are averse to feeling bad, so criticism needs to be extremely subtle in order to not offend.
> Btw: I wonder if we can stack several of these clichés. For example: "leverage" + "use case" = "leverage case".
I hate you for even thinking of this.
The personal association you made between "discussing anything even slightly personal" and "criticism needs to be extremely subtled" makes it sound that your problem isn't language or Orwellian discourse but the way you subconsciously link discussing personal matters with harshly criticising those you speak with for no good reason.
If your personal conversations boil down to appease your own personal need to criticise others then I'm sorry to break it to you but your problem isn't language.
You also went to "Orwellian discourse", which has a specific meaning, from a text by Orwell I mentioned. It seems to me like you got personally offended, interpreted my comment in the most uncharitable way, and chose to lash out at me instead, and I'm not sure why. I wasn't even talking about anyone specifically.
I repeat, your problem is not language. Your problem is that you manifest a need to criticize others. That problem is all on you.
This shows from the fact that you're saying "criticize" like it's a bad thing.
https://news.ycombinator.com/item?id=22911497
PS: I work with like 65-70% of that stack daily
Right now, we do hypercorn multiproc -> per-proc quart/asyncio/aiohttp IP-pinned event loop -> Apache arrow in-app cache -> on-gpu rapids.ai cache.
But not happy w event loop due to pandas/rapids blocking if heavy datasets by concurrent users. (Taking us back to celery, redis, etc, which we don't want due to extra data movement..) Maybe we can get immutable arrow buffs shared across python proc threads..
Ideas welcome!
If you don't need all that, then it's not a problem in Python either. You don't even need asyncio, you can just use a ThreadPool or a ProcessPool, dump stuff with pickle/shelve/sqlite3 and be on your way.
It's not in the standard library, and there are probably other options too now (not a heavy Python user any more), but Python has had easy parallelism for at least a decade (probably longer).
Performance is pretty close for both. I was disappointed to not see Python get an official green thread implementation. The counter-argument I commonly see cited is https://glyph.twistedmatrix.com/2014/02/unyielding.html. I personally don't find it to be a very convincing argument.
The coroutine and queue model is the same right ?
Cool thing in 3.8:
Running python -m asyncio launches a natively async REPL.
The vast majority of the time, gevent's monkeypatching works without any issues. With asyncio, you basically have to rewrite everything from the ground up to always use the new async APIs, and you can't interact with libraries that do sync I/O.
with Pool(5) as p:
print(p.map(f, [1, 2, 3]))
This runs f(1), f(2), and f(3) in parallel, using a pool of five processes.Python standard library since 2.6. Pretty much the definition of "in python".
Besides, Go has its own set of problems with parallelism. None of those are best in class.
As for Go's "set of problems with parallelism", they're pretty much just that sharing memory is hard to do correctly without giving up performance. No languages do this well; Rust and Haskell make it appear easier by making single-threaded code more difficult to write--requiring you to adhere to invariants like functional purity or borrowing. If you're writing Python, you very likely have values that are incompatible with these invariants (you want to onboard new developers quickly and you want your developers to write code quickly and you're willing to trade off on correctness to do so).
Go is absolutely best-in-class if you have typical Python values.
"Best in class" is not a relative term. Neither Go nor Python are appropriate choices for highly parallel intercomunicating code. Yes, Python is more limited than Go here, but hardly makes a difference when you avoid it.
Haskell and Rust do make it easier, by forcing developers to organize their code in a completely different way. Erlang does the same, with a different kind of organization. None of those languages are more difficult to program in, but yes, they are hard to learn.
thats an extraordinary claim which needs evidence.
Name any single Go feature aimed at helping parallel computing.
On its own, yes. For webapps, you can easily combine it with a multi-process WSGI server (like gunicorn or similar).
And it's showing great promise ...while it could be delayed.
https://www.mail-archive.com/python-dev@python.org/msg108063...
2021 is probable going to be the "Gone Gil" moment !
You can't say that when Python shoves Pip/Poetry/VirtualEnv/Black/Flake and what not into people. In contrast, Go has built-in package management and gofmt. Python is essentially gate-keeping new devs.
There was also a bad decision of using Python code for installation (setup.py) instead of a declarative language.
Most of that issues are actually fixed in setuptools if you put all settings in setup.cfg and just call empty setup() in setup.py.
Like here: https://github.com/takeda/example_python_project/blob/master...
If nothing else, Go lets you distribute a static binary with everything built in, including the runtime. Python's closest analog is PEX files, but these don't include the runtime and often require you to have the right `.so` files installed on your system, and they also don't work with libraries that assume they are unpacked to the system packages directory or similar. In general, it also takes much longer to build a PEX file than to compile a Go project. Unfortunately, PEX files aren't even very common in the Python ecosystem.
> Pip/Poetry/VirtualEnv
"packaging" refers to the way the language manages dependencies during the build and import process, not how you distribute programs you have built.
Python has a deservedly poor reputation here, having churned through dozen major overlapping different-but-not-really tools in my decade and a half using it. And even the most recent one is only about a year into wide adoption, so I wouldn't count on this being over.
Go tried to ignore modules entirely, using the incredibly idiosyncratic GOPATH approach, got (I think) four major competing implementations within half as long, finally started converging, then Google blew a huge amount of political capital countermanding the community's decision. My experience with Go modules has been mostly positive, but there's no really major new idea in it that needed a decade to stew nor the amount of emotional energy. (MVS is nice but an incremental improvement over lockfiles, especially as go.sum ends up morally a lockfile anyway.)
But if Python finally "picks" poetry, sticks with it for a few years and incrementally fixes problems rather than rolling out yet another new tool, that will also be better.
You can only identify the end of the churn for either retroactively. Python just looks worse right now because it's been around longer.
go: tends to wait and implement something once the problem is understood. took 2 years after go maintainers decided to solve the dependency issues. and as of the latest release its finally been labelled production ready.
and honestly the proxy issues were not real. go modules was still optional. you could just turn it off.
python is how old now? couple decades? and it has only gotten worse over time.
> go: tends to wait and implement something once the problem is understood.
Go's modules provide no additional "understanding" over any of the other Bundler-derived solutions in the world. MVS was the primary innovation, but wanting checksum validation means I have to track all the same data anyway.
> took 2 years after go maintainers decided to solve the dependency issues
This is revisionist history. There were other official "solutions" before ("you don't need it", "vgo is good enough", and "we'll follow community decisions"). If this one sticks, it's fine. But you can't say it's good now just because it's the one we have now - it's good now if it's the one we still manage to have in five years.
Go's track record is not "good" (in that regard I think only Cargo qualifies). At best it's "mercifully short."
> and honestly the proxy issues were not real.
Documentation was poor, the needed flags changed shortly before release, the design risks information leaks, and the entire system should not have been on by default for at least one more minor version.
> python is how old now? couple decades? and it has only gotten worse over time.
Yeah, that's exactly why I said "Python just looks worse right now because it's been around longer." It hasn't gotten worse though, it just also hasn't stopped churning. And if Go doesn't stop churning, in 10 years it will look the same.
The age argument works both ways - multiple major versions of Python predate Bundler. Go has no excuse for taking so long to reinvent "Bundler with incidentals", just like every other language.
With elixir, you run "mix release" and the release pipeline is set up to automatically gzip the release (it's one line of code to include that). Shoot the gzip over (actually I upload to s3 and redownload), unzip, and the entire environment, the dependencies, the vm, literally everything comes over. The only thing I have to do is sudo setcap cap_net_bind=+ep on the vm binary inside the distribution because linux is weird and, as they say, "it just works".
Kubernetes - which is one of the biggest projects built in go - has been struggling with dependency and package management.
Here's the CTO of Rancher commenting on his struggles
https://twitter.com/ibuildthecloud/status/118752909888666419...
https://twitter.com/ibuildthecloud/status/118753821015230873...
This is not trivial stuff..and it shouldn't be trivialised into a go vs python flamewar. Because it can't be.
To work with Python packages, you have to pick the right subset of these technologies to work with, and you'll probably have to change course several times because all of them have major hidden pitfalls:
* wheels
* eggs
* pex
* shiv
* setuptools
* sdist
* bdist
* virtualenv
* pipenv
* pyenv
* sys.path
* pyproject.toml
* pip
* pipfile
* poetry
* twine
* anaconda
To work with Go:
* Publish package (including documentation): git tag $VERSION && git push $VERSION
* Add a dependency: add the import declaration in your file and `go build`
* Distribution: Build a static binary for every target platform and send it to whomever. No need to have a special runtime (or version thereof) installed nor any kind of virtual environment nor any particular set of dependencies.
Python packaging is a mess, but Go doesn't even bother. "Just download from some VCS we'll pretend is 100% reliable and compile from source" is not a packaging solution.
You can certainly debate the difference in uptime between specific services; I don't know either way, but if you told my that PyPi had higher uptime than GitHub, I'd believe you... but that's kinda missing the point. If you depend on an online service to host your release artifacts, if and when that service goes down, it's gonna hurt.
Meanwhile, Python's packaging wars continue to rage on. Go's is simple: a release is a tag in a VCS repository. I'm sure there are issues with that as well, but that should come as no surprise, considering there are issues with literally every packaging solution. At any rate, there's little moral difference between downloading a tarball (or a wheel, or... whatever), vs. pulling a tag from a git repo. It requires equal levels of trust to believe that no one has tampered with prior releases in both cases.
I'd like to also point out that I don't have a dog in this race. I've done a little Go here and there, but frankly I don't like the ergonomics of the language too much, so I stay away from it. I've done (and continue to do) a decent amount of Python. I like the language, but tend to prefer strongly-typed, functional languages, and languages with performant runtimes, so I tend to only use it for smaller projects.
go has support for a proxy system, tooling is still immature though.
that's what vendoring is for and the proxy cache. this problem hasn't existed since like go 1.8 and is completely resolved in go1.14.
Can you provide some context for this statement? I've used Python asyncio extensively in Fargate (no ASGI frontend), and the developer experience is far from Go; however, I don't see how an ASGI framework can fix this. It seems like it offers the same general course-grained parallelism that you get from a containerized environment like Fargate except that Fargate abstracts over a cluster while ASGI frameworks presumably just abstract over a handful of CPUs.
For example, we have a large data structure that we have to load and process for each request. We want to parallelize the processing of that structure, but the costs to pickle it for a multiprocessing approach are much too large. We've considered the memory mapped file approach, but it has its own issues. We're also looking at stateful-server solutions, like Dask, but now we're talking about running infrastructure. In Go, we could just fork a few goroutines and be on our way.
i dont claim to have expertise in your business domain, but you should get the answer here.
Of course if you are stuck with Python it's better than nothing.
(1) it's undersupported (e.g. you can in theory download s3 files using async botocore, but in practice it is hard to use because of strict botocore version dependencies)
(2) it isn't natural - once you're in the event loop it makes sense, but using an event loop alongside normal python code is confusing at best.
(3) it got introduced too late. The best async primitives are only available in pretty recent versions of python3.
The difference with go is that it go has these primitives built in from the beginning and using them doesn't introduce interoperability problems.
If you have CPU intensive workload, an optimizing compiler can help. Well, for that you have C or other languages with more mature and aggressive compilers.
https://www.businessinsider.com/mcdonalds-spending-millions-...
This is also why McDonald's introduced table service, which is only in restaurants that have a layout where it's impossible to hide how many people are waiting. It costs manpower to deliver food to tables, but the additional orders are worth it.
McD's don't mind if you have to wait, they mind if you leave before you order. "Busy-looking queue" is a much more frequent problem than "totally-packed-restaurant".
Source: I like hamburgers + I geek out over stuff like this. I.e., Just Guessing.
Chipotle is about as fast as it gets
I get that people like it and I'm certainly not going to discourage anyone from doing what they enjoy, but I also feel like social media has turned certain fast food chains into memes where you can't merely just be satisfied with something, you either need to looooovveeeee iiiitttt or demand it be "canceled". There's no middle ground.
If I had the money I'd spent it to open a good old McDonald's I'm sure more people would come.
My reasoning is that they are just trying as hard as they can to decouple order-taking from order-preparing and distributing, taking their clue from Starbucks. Volume of this or that is not really an issue: with the exception of chips, these days they hardly prepare anything at all before it’s been ordered, so it doesn’t really matter whether they make a burger or a muffin.
The kiosk system is human-less thinking, we're not machines operating with pure logistics in mind. We like to have responsibilities in a way. When I'm talking to a cashier, he/she feels a duty to do something from A to B. With decoupled ordering.. nobody knows who I am, nobody really cares (McDonalds doesn't pay nor train for welcoming and service mindset). I'm just a thing carrying a ticket.
I've seen things so absurd, people walking around not knowing who to talk to, where to go, what to assume; both customers or employees. It was a surrealist comical situation.
Also my average time to serve for 1 hamburger with no other customer around is 7min. (5-15). Talk about fast food :)
Another point, I don't mean to make people work more, but I even prefer busy waiting lines with hectic kitchen action. That's what McDonalds were, high throughput grill. It felt something.. now it's all dull and clinic.
Oh and lastly .. the kiosk are fugly. They break the room space, break the flow of people, are way too big for their purpose (100$ a stupid tiny 80s monochrome terminal, would do better :p)
ps: This decoupling idea was tried at a company I worked for. Exact same principle (which I also gave some credit at first), split everything in small chunks so people can go faster.. it all went worse because nobody took responsibility for anything since a single task was now a dozen tiny bits done by a dozen people not really knowing what their bit was for. They just passed the products from hand to hand, not being able to track who or what was wrong until the last guy received all the shit because he's the one to show the result to the managers :)
I like minimalism, but sometimes batteries are included for a reason.
If someone finds that Redis and Celery are more complication than they need for a given task, then I think they're probably not using an orchestration framework.
If you're worrying about tweaking Celery for performance, then I suspect your uses may be a bit more complex than uwsgi's mules are designed for though.
https://uwsgi-docs.readthedocs.io/en/latest/Caching.html https://uwsgi-docs.readthedocs.io/en/latest/Mules.html
The cache is a simple key-value store. Works well. I later swapped it out for Redis, because I needed to shared my cache between multiple machines. Switching to Redis was a very quick and easy replacement, so you don't need to worry about lock in.
I used mules in a couple of ways. I had some background task mules which mostly just ran in a loop with a `sleep()` call. An example is deleting old records from a database once a day. Another example is listening to an Amazon SQS queue for files being uploaded to an S3 bucket.
I also had some mules which were triggered by web requests. These are normally what you would use "farms" for. The web request sends a farm message, and a mule picks it up and acts on it. For example, I used this for sending webhook callbacks in response to certain web requests.
You could probably also combine this with uwsgi's async mode. That would be useful if you needed a web request to wait for a long running task to finish before sending a response back. I handled that kind of situation with the aforementioned webhook callbacks instead.
A sibling comment has mentioned Legion. I've never used that, so can't comment on whether or not the caching and messaging works together with that.
Agree, Mcdonalds has definitely upped their ordering game recently. Thank you and I appreciate all the helpful comments!
1. Quite often you don't (I've built dozens of websites without needing Celery)
2. Even if you think you do there's often a much simpler solution that is enough for most needs (Use cron, spawn a process etc)
Celery is a big, heavy lump of code to add to most websites and it increased the deployment complexity.
I've usually had to build a small python cron runner using croniter in previous systems - which I think is a pretty clean solution - it just deferred tasks to rq workers. But having direct support in the lib might be nice.
Questions like: should I send this email in-line in the web request? Get a very easy answer: no, just stick do it later. Sure, sending an email is probably fine to do in-line for now, but months in you may realise that things are slow, that you’re sending emails and rolling back transactions later, or committing the transaction but losing the email that needed to be sent, or all manner of other annoying edge cases. Queues don’t solve everything, but they can be an ok answer to a lot of stuff for a long time.
For basic sites, yeah maybe not necessary, but a reliable background processing system has always been a significant accelerator in my projects.
Several developers like to overengineer and "go for celery" (also applies to other technologies with other uses) even for small things.
You don't need Celery to run a batch job every day for example. Or to even do some parallel processing.
"Oh but python multithreading sucks" do you know when it does not suck? When your thread is waiting on something else. Also there's the multiprocessing module with a lot of "batteries included" for basic use cases.
Not to mention (my biggest pet-peeve of Celery) is that it "forces" you to work with a task queue model. Not a data queue model. Sure, it helps a lot when you need that, but sometimes you just need a queue.
For example, sending a mail can take a while because mailservers have queues and whatnot. If you send a mail after signing up a user for them to verify their email, you don't have to wait until the mail is "really" sent before letting them know that their signup has been processed. You can tell a background worker to send the mail and return to the user much faster. For another example, in a previous job I worked for a big file sharing service. If user wanted their files deleted that caused all sorts of calls to AWS to actually delete the files, which could take a while. However, from the user perspective it was pretty fast because all we had to do was set the file in the database to "delete in progress" state and tell a background worker to delete the files. Then we could show the user that their files were being deleted within a couple dozen milliseconds instead of having to wait for all the AWS calls to complete.
Although this delay chain might be considered worth it you don't want to scale the frontend to multiple workers for some reason e.g. single-threaded evented runtime, or GIL runtime (python, ocaml), or if you want to avoid CPU-hard tasks being executed on your frontends.
In that case, it might be valuable to transform CPU waits into IO waits by moving the CPU work to a jobs queue, possibly running its workers on a different set of machines entirely.
Really though, I think a lot of people use celery for offloading things like email sending and API calls which, IMHO, isn’t really worth the complexity (especially as SMTP is basically a queue anyway). Of course, YMMV depending on your use case.
However, I find it is often more worthwhile for:
1. Tasks which take a long time to run
2. Tasks which need to happen on a schedule, rather than in response to a user request.
There is an option 3 too, which is for inter-software communication. Eg events or RPCs, but I found Celery to be very much a square-peg-round-hole for this, which is why I developed Lightbus (http://lightbus.org). Lightbus also supports background tasks and scheduled tasks. /plug
Honestly, I think that mostly anything that doesn't depend directly in the current state of your infrastructure should be done asynchronously. I've had a lot of issues with systems that start up doing everything synchronously: you'll probably need to refactor it to be asynchronous in emergency mode during a crisis.
But anyway, how is your application supposed to respond after any of those failures? Is it just supposed to ignore the failure and thread on like if nothing happened, leaving your users on the dark? Is it supposed to reliably log every task so that it can retry anything that fails and in the worst case feed failures into some monitoring system/process? Or is it supposed to inform the user of any success or failure before the user can move on?
Queue software is only a good match for the first. For the second you will need to roll your own interface with the monitoring system anyway, so it's much easier to roll your own queues and get control of everything. The third one is best done synchronous, it doesn't matter the nature of the process or how long it takes. But funny thing is, I have never seen the first situation on the wild.
But as we decided to implement an MVP with a single HTTP request from the client, this whole separation doesn't make any sense, exactly as you noted.
The correct pattern should be a client submits a request to get something done to a thin layer, receives a ticket that allows it to claim the result and goes away to either check for the result via polling for the ticket or receives a call back.
It depends. If you needed a coordination layer or you needed to isolate certain types of traffic then it makes more sense. I assume your alternative here is "why not just have a new client tier doing work" which is a reasonable architecture too
> One aspect of this set up I’ve never been able to understand is how the application then gets the result from the worker?
Often they don't in these architectures. I found it a little strange that a job queue was being used to serve (what seems like) synchronous traffic. Usually I see job queues in the wild used for async/send-and-forget workloads
Personally I would rather chain multiple synchronous service calls if I needed a synchronous workflow. It's just simpler to me to stick web workers behind a load balancer and scale that. This is less elegant when things take a very long time though or are prone to retries. Either the client or server needs to be responsible for queuing/message persistence/retries - with services the client does it, with job queues the server does it
The above post walks through sending emails out with and without using Celery, making third party API calls, executing long running tasks and firing off periodic tasks on a schedule to replace cron jobs.
There's links to code examples too in an example Flask app which happens to use Docker as well.
Also consider if the machine running that process just disappears and that process dies. Putting work into a task queue allows you to do it durably until it can be processed so that it's not lost in some typical "that machine/instance died" scenario.
Every problem is not the same. Work queues introduce a magnitude or more of complexity to an http application. Sometimes that is very needed. Sometimes its overengineering. Sometimes it is a gray area.
What happens when you want to retry jobs with exponential back off, or rate limit a task, or track completed / failed jobs?
You can wire all of this stuff up yourself but it's a hugely complicated problem and a massive time sink, but Celery gives you this stuff out of the box. With a decorator or 2 you can do all of those things on any tasks you want.
I use it in pretty much every Flask project, even on single box deploys.
In Python too:
# or ThreadPoolExecutor depending of the type of work
from concurrent.futures import ProcessPoolExecutor, as_completed
with ProcessPoolExecutor(max_workers=2) as executor:
a = executor.submit(any_function)
b = executor.submit(any_function)
for future in as_completed((a, b)):
print(future.result())
Any task you would put in celery would be a good candidate for being first passed to a process pool executor.No, the real reason we use celery is that it solves many problems at once:
- it features configurable task queues with priorities
- it comes as an independent service accessible from multiple processes
- its activity can be inspected and monitored with stuff like flower
- it has the ability to create persistent queues that survive restart
- error handling is backed in, they are logged and won't break your process
- it offers an optional result backend for persisting results and errors
- many distribution and error handling strategies can be configured
- tasks are composable, can depend on each others or be grouped
- celery also does recurring tasks, and better than cron
- you can start tasks from other languages than Python
Now some People use celery when they should use a ProcessPool, because they don't know it exists. But that's hardly because of the language: you didn't seem to know about it either.It is true that Python concurrency story is not comparable to something like Go, Erlang or Rust, but very common use cases are solved.
In the same way, we can perfectly use shelves instead of redis.
We use redis beacuse it solves many problems at once:
- it's very performant on a single machine, and can be load balanced if needed
- it has expiration baked in
- it offers many powerful data structures
- it embeds numerous strategies to deal with persistence and resilience
- it's accessible from multiple processes, and multiple languages
- its ecosystem is great, and it's hugely versatile
It's almost never a bad choice to add redis for your website. It's easy to setup, cheap to run, and and it shines to manage sessions and caching for any service with up to a few million users a day.Once you are at it, why not later use it for other stuff, like storing the result of background tasks, log stream, geographical pin pointing, hyperlolog stats, etc. ? There are so many things it does better than your DBMS, faster, easier or for less resources.
It's such fantastic software really.
But no, nothing prevents you in Python to create a queue manually and serialize it manually. It's just more work for less features.
I also understand the benefits of task queues. However, there are many cases in which you do not need any of those. Specifically, everything you wrote applies to the web backend/distributed systems usecase. Doing things in the background in a simple application, not so much. My problem is exactly with introducing distributed systems machinery for a local process on a single machine that doesn't need any of that.
You have an operating system that you can use. You don't really _need_ concurrency when you have a machine that can timeshare amongst processes. That's the world Python was designed for
> In most other languages you can get away with just running tasks in the background for a really long time before you need spin up a distributed task queue
This is partly true, but not a lot of people do this because you need to persist those tasks unless you want them dropped during a reboot. Same logic goes for entire machines disappearing. You _can_ get by without a distributed system, but you will need to tolerate loss in those scenarios. Those losses are non-trivial for most apps/companies, so it doesn't seem all that practical to me to consider a world without a distributed system (and yes, persisting things in postgres/mysql before they are being worked on is still using a distributed system)
I wonder how many other people have celery just for email.
It can be hard to find the right analogy. If the subject of the analogy is enough of a hot topic, it will get attention itself. As you can see, you've got comments in here even talking about how McDonald's isn't faster with their new system, or how they could have better optimized for customers, links to articles about McDonald's business model etc. Some of the earlier negative comments about McDonald's were deleted - probably due to downvotes. Since my advice to other writers was sincere and I believe useful, I'm keeping my comment.
For sure I'm not expecting you to change your article. Just hoping that my tip might help you with future technical writing. If not, no worries.
Had you got e.g. a Java app running with a multithread application server model, you could serve and process everything within a single process. No Celery, no Redis, no MQ.
This doesn't mean that the above stack has no use. But, whenever picking a tech, you should understand the use case. The "simpler" Python/Flask solution has an increased complexity when the task at hand is not simple anymore.
They do, it's a WSGI application. Flask has been multithreaded for many years. Python also has multi-process queues that don't need an extra process.
The application server model makes it so there's a running application, with a state, and which exposes an HTTP endpoint.
> Flask has been multithreaded for many years
So? You run a blocking thread to performing a long-running task in Flask? With Python? Try, then report what you find.
Flask/Django are mostly designed to work with a stateless approach. Nothing wrong with that, but it's got drawbacks.
> multi-process queues that don't need an extra process.
More processes usually imply more complexity. And still, since you don't have a central application with a state, you NEED an extra piece to manage the result from the queue.
I did it for years. It works just fine. The GIL is essentially like running an app on a single core, which works just fine for many use cases. CPU cores are quite powerful.
> More processes usually imply more complexity.
Right, but I'm only saying you can have more processes without requiring Redis.
> And still, since you don't have a central application with a state, you NEED an extra piece to manage the result from the queue.
A regular thread can do that.