Needless to say, working with PDFs makes me want to pull my hair out.
I also ended up writing the SpacySentencizer Executor instead of using a "vanilla" sentencizer. That led to consistent sentence splitting (so "J.R.R. Tolkien turned to pg. 3" would be one sentence, not 5)
For testing, Jina allows you to swap out encoders with just a couple of lines of code, so trying different methods out should work just fine.