Kew – Task Queue That Runs Inside Your FastAPI Process
github.com
github.com
Kew (https://github.com/justrach/kew) solves this by running directly in your FastAPI process. I'm using it in production for my AI apps where it handles thousands of LLM API calls daily with proper concurrency control.
Technical details: - Runs directly in your FastAPI process using asyncio - Uses Redis for persistence - True concurrency control with semaphores (crucial for managing expensive AI API calls) - Circuit breakers for handling API timeouts and failures - Millisecond-precision task scheduling
Simple example:
# Create a queue with strict concurrency control
await manager.create_queue(QueueConfig(
name="llm_tasks",
max_workers=4 # Strictly enforced for API rate limits
))
# Use it in your FastAPI endpoint
@app.post("/generate")
async def generate_text(prompt: str):
await manager.submit_task(
queue_name="llm_tasks",
task_func=call_llm_api,
prompt=prompt
)
Why: Current solutions (Celery/RQ/Huey) require separate worker processes which add complexity to FastAPI deployments. They weren't designed for async frameworks and require sync/async context switching. This becomes especially painful when dealing with AI API calls where you need precise control over concurrency.This is running in production on several of my AI applications. It's particularly good at handling concurrent LLM API calls where you need to respect rate limits while maintaining high throughput.
The code is MIT licensed and available at https://github.com/justrach/kew.
Would appreciate feedback especially on the concurrency control implementation. Also curious to hear from others building AI apps about their task queue patterns.