HNHacker News
TopNewBestAskShowJobs

j0selit0

286 karma · joined November 18, 2022

submissionscomments
j0selit0··on I don't need Codex agents
I need changes I can understand and review
j0selit0··on RAG Is Simpler Than You Think
that makes me even more humbled :)
j0selit0··on RAG Is Simpler Than You Think
author here - apologies for it not being clear. the idea here is:

step 1: sparse index retrieval (FTS/BM25) - say with k = 10 step 2: re-rank the 10 records using embeddings

the difference in this approach is during step 2, you convert text to embeddings on the fly - when you're running the retrieval pipeline, meaning you don't need to have all of your corpus pre-embedded in a vector db

j0selit0··on RAG Is Simpler Than You Think
you would also need to maintain multiple indexes in multiple languages. I never had to do that - but I assume it's a pain
j0selit0··on RAG Is Simpler Than You Think
good one. yes, that's pretty much my experience with query rewriting as well
j0selit0··on RAG Is Simpler Than You Think
to be fully transparent, I used LLM to review my draft yes. and interestingly enough, this article had tens of thousands of more views and estimulated significantly more discussions than my other, non-LLM written articles. so there's that.
j0selit0··on RAG Is Simpler Than You Think
and for the record, my last employer was still using gpt-4o and gpt-4o mini last year. and they are an F500 (not that it means anything, just for context).
j0selit0··on RAG Is Simpler Than You Think
sorry if it came across as name dropping - the point I wanted to make is actually that F500s are chasing RAG (as in pure semantic search without any care for chunking strategies, evals etc) instead of starting with the basics.

and for the record, my last employer was still using gpt-4o and gpt-4o mini last year. and they are an F500 (not that it means anything, just for context).

j0selit0··on RAG Is Simpler Than You Think
this is quite nuanced. in financial markets you have a combination of natural language questions that involve technical keywords / slang / acronym. and for these specific terms, an off-the-shelf embeddings model fails miserably.
j0selit0··on RAG Is Simpler Than You Think
my personal experience - I have been involved with such projects for the last 2 years. interestingly enough, a lot of times such initiatives didn't take off because people/stakeholders were overcomplicating things and wanting to use semantic search for everything - without having a minimum knowledge of chunking strategies, pros/cons etc
j0selit0··on RAG Is Simpler Than You Think
I work on a really similar scenario. We ended up having to create a so-called "semantic layer" containing metadata (table and schema descriptions) and glossary terms. Still, there is a lot of work involved maintaining glossary terms since some of them are ambiguous and people have different understanding/interpretations for some of them.
j0selit0··on RAG Is Simpler Than You Think
I wish everyone thought like you, in my experience unfortunately it's not the case
j0selit0··on RAG Is Simpler Than You Think
author here - that's an amazing idea. would be an insanely large article though - maybe will write up a series
j0selit0··on RAG Is Simpler Than You Think
author here - thank you!
j0selit0··on RAG Is Simpler Than You Think
author here - thanks, I'm honored you would use my agents :)
j0selit0··on RAG Is Simpler Than You Think
I'm sorry is this ironic or not? doesn't sounds simple at all
j0selit0··on RAG Is Simpler Than You Think
author here. at most companies I've worked for recently (F500) RAG is still quite trendy. this was what frustrated me a bit and motivated to write this article - along with other experiences that definitely relate with some of the folks in the comments above
j0selit0··on GitHub Copilot trusted a client header for premium-request billing
Follow-up to my previous Copilot/mitmproxy post.

While looking at the traffic I noticed Copilot’s own LLM requests used X-initiator: agent and didn’t consume premium quota. That made me curious what would happen if I changed my own requests from user to agent.

Turns out the model still responded, but those requests didn’t decrement my premium quota or show up in the billing analytics.

j0selit0··on What I learned by putting GitHub Copilot behind a MitM proxy
thanks! corrected in the article
j0selit0··on What I learned by putting GitHub Copilot behind a MitM proxy
I was curious to understand how Copilot implements its harness, and also how I was exhausting my quota so quickly. End up going down a rabbit hole of intercepting its network traffic with mitmproxy.

A few interesting things I found along the way:

- watched model/capability discovery and routing happen in real time - looked at what gets injected into context and sent with ghost completions - found that recent edits can pull in context from files other than the one you're currently editing (including infamous .env) - found the SQLite session store behind Chronicle, including previous prompts/responses - watched the model query that history through tool calls

I then went through the VS Code source to reconcile some of what I was seeing on the wire with the actual implementation.

Overall some interesting lessons around how their harness is implemented.

j0selit0··on OpenAI-agents-Redis: Native OpenAI Agents SDK session management using Redis
I was building a production app with OpenAI's new Agents SDK and hit a wall with session management. The built-in SQLite support works great for prototyping, but doesn't cut it when you need to scale across multiple instances.

So I built openai-agents-redis – a drop-in replacement that uses Redis instead. Same API, but now your agent sessions can be shared across processes and survive container restarts.

It's lightweight (just a thin adapter) and handles connection pooling, serialization, and cleanup automatically.

GitHub: https://github.com/rafaelpierre/openai-agents-redis PyPi: https://pypi.org/project/openai-agents-redis/

Maybe others running into the same SQLite limitations will find it useful. Curious for your thoughts!

j0selit0··on PyJaws: A Pythonic Way to Define Databricks Jobs and Workflows
PyJaws enables declaring Databricks Jobs and Workflows as Python code, allowing for: - Code Linting - Formatting - Parameter Validation - Modularity and reusability

In addition to those, PyJaws also provides some nice features such as cycle detection out of the box.

Folks who have used Python-based orchestration tools such as Apache Airflow, Luigi and Mage will be familiar with the concepts and the API of PyJaws.

j0selit0··on Who needs MLflow when you have SQLite?
I mean, come on, SQLite doesn't even support concurrency. Are people seriously considering using it in a production scenario?

If you work in a DS team where you're the only DS, then it probably suits your needs. Otherwise I can't imagine how you could achieve anything production grade