The model is technically in everyone's Windows installs, but we don't have the C++ projected WinRT headers to use the Microsoft.Windows.Vision library.
The model is technically in everyone's Windows installs, but we don't have the C++ projected WinRT headers to use the Microsoft.Windows.Vision library.
Image-to-text models, filtered by "Microsoft": https://huggingface.co/models?pipeline_tag=image-to-text&sor...
It looks like it doesn't recognize more than one line at once, but combined with another model or algorithm to detect text bounding boxes it'd be handy.
The one somewhat unique offering in Azure is the Document Layout model which gives you back the OCR with titles, headers, paragraphs, and tables all labeled.
This is really, really good for RAG since it's often useful to stuff the nearest header into the chunk of text when generating an embedding (much, much better results this way).