If Gemini can do semantic chunking at the same time as extraction, all for so cheap and with nearly perfect accuracy, and without brittle prompting incantation magic, this is huge.
If Gemini can do semantic chunking at the same time as extraction, all for so cheap and with nearly perfect accuracy, and without brittle prompting incantation magic, this is huge.
The pre-step, chunking and semantic understanding is all that counts.
So I can ask Gemini to return chunks of variable size, where each chunk is a one complete idea or concept, without arbitrarily chopping a logical semantic segment into multiple chunks.
- BM25 to eliminate the 0 results in source data problem
- Longer term, a peek at Gwern's recent hierarchical embedding article. Got decent early returns even with fixed size chunks
For others interested in BM25 for the use case above, I found this thread informative.
We use it in combination with semantic but sometimes turn off the semantic part to see what happens and are surprised with the robustness of the results.
This would work less well for cross-language or less technical content, however. It's great for acronyms, company or industry specific terms, project names, people, technical phrases, and so on.