> The key advantage of doing batched inference at a huge scale is that once you maximize parallelism and sharding, your model parameters and the memory bandwidth associated with them are essentially free (since at any given moment they're being shared among a huge amount of requests!)
Kind of unrelated, but this comment made me wonder when we will start seeing side channel attacks that force queries to leak into each other.