Hmm good insight there. I've done some experimenting formerly by chunk length and it's been pretty troublesome due to missing context.
Hmm good insight there. I've done some experimenting formerly by chunk length and it's been pretty troublesome due to missing context.
Define a custom recursive text splitter in langchain, and do chunking heuristically. It works a lot better.
That being said, it is useful to maintain some global and local context. But, I wouldn't use overlapping windows.
When working with more extensive documents, the process gets a bit more intricate. In this case, your embedding database might need to hold more information per entry. Ideally, for each document, the database should store identifiers like the document ID, the starting token number, and the ending token number. This way, even if a document appears more than once among the top results from a query, it's possible to piece together the full relevant excerpt accurately.
If you use local models then it's a fantastic idea.
https://unstructured-io.github.io/unstructured/bricks.html#p...