I am working on something similar. I have all the 10-k and 8-k docs.
I’ve pulled out the structured data separately, and now looking at breaking up the text into paragraphs to get embeddings.
Why are you using inverted index style text search instead of embeddings? I can see doing both perhaps…