How we index images for RAG
kapa.ai
kapa.ai
For example you might identify a car in an image but the context is the car running a red light. A new model might pick that up while an old one doesn't. These context adjustments might sometimes require you to rerun your LLM processing or potentially have a one to many relationship for multiple runs so you can take the best of or combine results.
Actual usage will also reveal most commonly used assets and you can target the ones that are most trafficked and save a ton on processing that way.
This is what I've been doing in my Obsidian infodump for a while. If I know that an image is important, I generate a text description (Mermaid if possible, English if not) and paste it after the image in a block. This lets agents see the image if they don't really see it. Though my process is manual, the improvements in outcomes for agents that rely on text search/retrieval is very real and is worth it.
Descriptions of images that are charts or diagrams to start with?
Retrieving based on text and then giving the generation model the image instead is much smarter than retrieving based on image. Image-based retrieval is slow and expensive.
Same with giving the model an image vs a structured representation of it.
How was the accuracy compared to pre-parsing the image and doing search in the text?
But the experience was that it was able to find small details in PDFs, in technical diagrams, and this was really not captured well at all with OCR.
In general, OCR I think should be used more as an add-on to retrieve data, not given to the generation model itself. Similar to retrieving based off a text description and then giving the generation model the image.
https://github.com/Qbix/AI/blob/6753f6e453908682401f49760002...
https://github.com/Qbix/AI/blob/main/config/observations.jso...
wrote it up here a few months ago: https://community.safebots.ai/t/building-cultural-infrastruc...
- Marketing material? check
- Bloated to the extreme? check
- "Get a free trial" at the end? check
- Entirely LLM generated? check
Man I hate that AI writing tic. I appreciate the instincts for sharing the workflow. It's still very difficult to get AI to put an info dense description together though, we tend to get long and vague.
Multimodal retrieval does not suit this domain. CLIP-style embeddings wash out exactly the fine detail that matters in charts, tables, and annotated screenshots, and short technical queries ("how do I configure X") give too little signal to match against image vectorsSo you include colour, shapes, etc?