Would you describe getting response from the model within few milliseconds cpu or memory intensive? What I assumed as to the characteristics of inference is a process that takes multiple seconds to minutes.
In a system that requires a response to the customer within a few milliseconds through web api, how do you ensure the performance in Python? I'm genuinely interested as that sounds outside of what stock CPython can do except trivial logic.
Another aspect that's completely ignored are the requirement - you might need the response time in milliseconds, couple of times per day. Others might need to serve hundreds of requests. You also need to consider the budget available.