Reading PDFs as html would
Be nice, but as a ML engineer having high quality conversion would be very useful for large scale analysis and information extraction of pdf documents!
In case you are looking for an API to extract structure rich content like tables from PDFs or images, look into this https://extracttable.com (p.s. I contributed to it)