We've been experimenting with nice compromises for using pydata on-the-fly compute with caching. Basic tension is indiv requests get blocked by their blocking APIs, while caches (eg, IP-pinned) are best when same-python-app-thread.
Right now, we do hypercorn multiproc -> per-proc quart/asyncio/aiohttp IP-pinned event loop -> Apache arrow in-app cache -> on-gpu rapids.ai cache.
But not happy w event loop due to pandas/rapids blocking if heavy datasets by concurrent users. (Taking us back to celery, redis, etc, which we don't want due to extra data movement..) Maybe we can get immutable arrow buffs shared across python proc threads..
Ideas welcome!