What I want is a dynamic chunking - I want to search a document for a word - and then I want to get the largest chunk that fits into my limits and contain the found word. Has anyone worked on such thing?
grep -C $n word document
will get you $n lines of context on either side of the matching lines.Using the models underlying a library like this, there's maybe room for fine-tuning as well if you have a set of documents with specific semantic boundaries that current approaches don't capture. (And you spend an hour drawing bounding boxes to make that happen).
[0]: https://en.m.wikipedia.org/wiki/Longest_common_substring