Breaking up with Flask and FastAPI: Why they don’t scale for ML model serving
modelserving.com
modelserving.com
Why? You are doing the inference in the same request, which is synchronous from the perspective of the caller. The request can be memory-intensive or cpu-intensive. And the issue is you can't efficiently consider all the workloads for a single machine without being bottlenecked by Python.
I would say that the problem is in your approach trying to use the webapp hammer for all the different flavors of nails in your system, using a language that isn't suited for concurrency. What I would do is decoupling the validation/interface logic from your models via a queue. This way you can scale your capacity according to workload and make sure the workload runs on hardware most relevant to the job.
I have a feeling trying to throw a webapp at the problem might not solve your root issue, only delay it in time.
In a system that requires a response to the customer within a few milliseconds through web api, how do you ensure the performance in Python? I'm genuinely interested as that sounds outside of what stock CPython can do except trivial logic.
Another aspect that's completely ignored are the requirement - you might need the response time in milliseconds, couple of times per day. Others might need to serve hundreds of requests. You also need to consider the budget available.
One approach is to translate the synchronous call into an async call plus polling on the webapp side. You push onto a queue, with the callback queue in the message body. But that gives you problems when you want deploy a new version of your webapp - existing connection will be disrupted and the state lost.
Since you need to deal with retries, anyway, you can move the logic into the client itself. It will get the request id on the initial response and then ask the service for results.
You see, this solution can vary wildly depending on your scalability, durability and resiliency requirements. And on your budget. Its not wild to expect the response to be big, so you might want to upload it to s3. You might use websockets, too. Technology gives you a lot of options here, of different levels of complexity and scalability ;)
In that case, you trade off safety whilst the setup is still simple. But if querying the model takes less that 100ms I would argue whether thats cpu/memory intensive at all. Remember that you still need to allocate and populate the memory, and Python doesn't free the memory back to OS.
I will say, though: the runner webapp MVP exists to do basically what you're describing (and keeps an internal queue). Yatai is architected so that the runner instances run on a separate Kubernetes cluster to the pre- and post-processing code that that's run in the main webapp itself, and can be scaled separately.
All models at scale eventually need to be executed by an async queue processor which is fundamentally different from a request response REST API. For simplicity managing this outside of the process making the web request will help you debug issues when people start asking why they are getting 502 responses. If you are forced to use python for this, I would always suggest of going to celery/huey/dramatiq as an immediate next step after the REST API MVP. I hear Celery is getting better but I have ran into issues over the year so it pains me to recommend it.
Prime example of the difference is whether you accept disruption of inference when you deploy a new version of your webapp.
Very high chances you don't, thus you will start implementing a queueing mechanism without realizing it.
Yup. There's a huge amount of work that you need to do to do the whole ML lifecycle, and FastAPI doesn't support that out of the box like a full fledged ML Platform.
But you probably don't actually want a full ML Platform because they're all opinionated and if you try and fight them it's often worse than just serving it as an API via FastAPI...
It would've been delightful to see "instantiate a runner in your existing Starlette application". I don't want to instantiate a Bento service. Perhaps I can mount the bento service on the Starlette application?
Apologies if I am still grossly misunderstanding. I tried to look through some of the _internal codebase to see how the Runner is implemented, the constructor signatures are very complex and the indirection to RunnerMethod had me cross-eyed.
Instantiating and using a runner can be done anywhere with `init_local`, but it's really the runner ASGI app that does the work of queuing and batching. We've thought about allowing users to spin that app up separately but it's not a focus right now; instead we're trying to ensure that the system is as easy to use for data scientists as possible and have that workflow fully ironed out before we support the more advanced use-cases.
The whole runner situation is quite complex because we wanted to support user-created runners in the nicest way possible, and also leave the space open for non-python runners (and service app) in the future.
> We've thought about allowing users to spin that app up separately but it's not a focus right now; instead we're trying to ensure that the system is as easy to use for data scientists as possible and have that workflow fully ironed out before we support the more advanced use-cases.
I like this, and I understand. I might like to take a stab at implementing this more modular interface.
This confuses me. How is that FastAPI’s fault? Can’t you just asynchronously delegate them to a concurrent.futures.ThreadPoolExecutor or concurrent.futures.ProcessPoolExecutor? What does Starlette provide here that FastAPI doesn’t? If the FastAPI limitations are due to ASGI, shouldn’t Starlette have the same limitations?
Definitely not FastAPI's fault and yes Starlette has the same limitations. BentoML builds additional ML features/abstractions on top of Starlette. We introduced a "runner" concept which automatically creates separate processes for models to run in.
We've been using Flask for years as the foundation for our 0.13 version. Our choice to move away from Flask and FastAPI as a core part of our library is based on our experience with hundreds of users and use cases
For model serving, we were thinking Triton (native vs python server) as it is a tightly scoped problem and optimized: any perf comparison there?
Will a project be abandoned because users are pointing out ways to improve it? Does work stop at some point because "we can't get to every issue"? Sure, the maintainer could get burned out, but that is not a given.
Projects with similar stars, momentJS or Jquery don't see much active development have very few issues and pull requests.