Strongly recommend using Apache Tika[1] for this. It's industry standard for ubiquitous document text extraction.
You can take the text output from Tika, chunk it with something like Chonkie[2], and embed it for your search index.
You can take the text output from Tika, chunk it with something like Chonkie[2], and embed it for your search index.