I’d imagine their capabilities mirror that of Mistral OCR [1]. Mistral outputs markdown, the image would have to be convertible to a reasonably useful markdown structure (charts, tables etc).
Pretty much anything with a different colored background gets returned as (image)[image_001].
Example: https://omni-demo-data.s3.us-east-1.amazonaws.com/test/17398...