nice - did you write a custom parser for PDF/DOCX? we wrote one for XLSX after running into event loop issues with sheet JS
[1] https://github.com/J-F-Liu/lopdf
testing was mostly manual with a test corpus we generated. its not perfect but its pretty close for most files we've seen
We wrote (should say are writing) our own xlsx parser in Rust on IronCalc: