I’ve been trying to figure out how to process hundreds of my own scanned photos to determine any context about them. This was convincing enough for me to consider google’s vision API. No way I’d ever trust OpenAI’s apis for this.
Edit: can anybody recommend how to get similar text results (prompt or processing pipeline to prompt)?