286 karma · joined November 18, 2022
step 1: sparse index retrieval (FTS/BM25) - say with k = 10 step 2: re-rank the 10 records using embeddings
the difference in this approach is during step 2, you convert text to embeddings on the fly - when you're running the retrieval pipeline, meaning you don't need to have all of your corpus pre-embedded in a vector db
and for the record, my last employer was still using gpt-4o and gpt-4o mini last year. and they are an F500 (not that it means anything, just for context).
While looking at the traffic I noticed Copilot’s own LLM requests used X-initiator: agent and didn’t consume premium quota. That made me curious what would happen if I changed my own requests from user to agent.
Turns out the model still responded, but those requests didn’t decrement my premium quota or show up in the billing analytics.
A few interesting things I found along the way:
- watched model/capability discovery and routing happen in real time - looked at what gets injected into context and sent with ghost completions - found that recent edits can pull in context from files other than the one you're currently editing (including infamous .env) - found the SQLite session store behind Chronicle, including previous prompts/responses - watched the model query that history through tool calls
I then went through the VS Code source to reconcile some of what I was seeing on the wire with the actual implementation.
Overall some interesting lessons around how their harness is implemented.
So I built openai-agents-redis – a drop-in replacement that uses Redis instead. Same API, but now your agent sessions can be shared across processes and survive container restarts.
It's lightweight (just a thin adapter) and handles connection pooling, serialization, and cleanup automatically.
GitHub: https://github.com/rafaelpierre/openai-agents-redis PyPi: https://pypi.org/project/openai-agents-redis/
Maybe others running into the same SQLite limitations will find it useful. Curious for your thoughts!
In addition to those, PyJaws also provides some nice features such as cycle detection out of the box.
Folks who have used Python-based orchestration tools such as Apache Airflow, Luigi and Mage will be familiar with the concepts and the API of PyJaws.
If you work in a DS team where you're the only DS, then it probably suits your needs. Otherwise I can't imagine how you could achieve anything production grade