Very cool! I've had great results with Mistral OCR as well. However, in terms of embedding models I've had a hard time getting consistent results with really large pdfs (over 40pgs). Ideally would like to be able to build indexes with just my two 4090s.
Would be super interested in that writeup!