I wouldn't put the model under fastapi or any similar framework, i would serve it from a different process to also allow me to serve multiple versions of the model as well (similar to tf serving). But eventually we have an API call to some web framework calling this different process and requiring a response with the model recommendations to be returned in a few milliseconds, how is a message queue appropriate for such a real-time use case, could you elaborate?