How Does GPT-4o Encode Images?
oranlooney.com
oranlooney.com
In my experience it works remarkably well for features like scanning documents in Notes and in copying or translating text embedded in images in Safari.
It is not open source, but free to use locally. Someone has written a Python wrapper (apple-ocr) around it if you want to use it in other workflows. The model files might be in /System/Library/PrivateFrameworks/TextRecognition.framework if you wanted to port them to other platforms.
Text extraction is included (including the ability to specify custom words not found in the dictionary) but there are also utilities for face detection, classification, etc.
Beats everything else, truly international and multi-lingual, including Chinese (as it is made in China)
Their PaddleLayout models are also miles ahead compared to LayoutParser or TableTransformers in both inference speed and output quality
The reason I use it is to test whether it's just analyzing letter-by-letter (even if they claim it does more) or if it's actually scanning the letter/word in its context. If it's letter-by-letter, I get hilariously awful results.
Sure, it got things wrong. But it also figured out some things even I couldn't decipher.
But the whole "point" of LLM (forget it, it's not AGI) is you don't need to make many specialized models and cursed pipelines anymore, to solve a definitely-in-reach-without-LLM problem your farmer neighbor wants to pay $500 for.
Before LLM it's not going to be done as it takes more than $500 engineer hours. Now we just brute force. Sure, more compute, but we get it done!
I guess your OCR dream is covered by this.
Could you please list some? I am developing a tool that relies on OCR and everything I've found refers to tesseract as being the best choice
Improving OCR would require innovation within CV - separate from transformer architectures and frankly I don’t expect much new work to happen here
[0] - https://github.com/microsoft/unilm/tree/master/kosmos-2.5
It's _extremely_ slow, about 30 seconds/page on an A10G. Maybe there's room to improve that, I don't know, but for now that's a problem.
The actual character recognition is superb, probably comparable to the big cloud offerings. On a sample of pages of moderately challenging typed text it was literally flawless aside from non-ascii characters.
It can do _neat_ handwriting with reasonable accuracy. This surprised me since there doesn't seem to be any handwriting in the training data (for Pix2Struct either). However, it will sometimes just skip handwriting.
The structured (markdown) output is sometimes impressive, occasionally a mess. I noticed two weaknesses in particular: it often starts a table as an HTML table and then switches to markdown, and it struggles to distinguish multi-column layouts unless they're pure book-like paragraphs or a table with clear gridlines. This is probably a result of sane/straightforward layouts from READMEs and scientific papers being most represented in the training data (the industry I'm in produces lots of documents with layouts that are wild even to a human).
One other thing: as a generative model it can and will go off the rails. One document I gave it with a lot of messy handwriting produced the typed header and then just 1500 lines of greater-than symbols. To be fair I couldn't read it either. While I didn't see it produce any valid-looking but "hallucinated" output, that's a possibility too.
Clip embeddings can absolutely “read” text if the text is large enough. Tiling enables the model to read small text.
I have done some other LLava stuff in it
The pyramid strategy loosely tracks with renormalization group theory, which has been formally studied for years as a method of interpreting machine learning models:
https://arxiv.org/abs/1410.3831
I love the convergence we're seeing in the use of models from different fields to understand machine learning, fundamental physics, and human consciousness. What a time to be alive.
Sounds like an arbitrage opportunity for all those gpt wrappers. Price your cost per token the same, send over the prompt via image, pocket the difference?
I get that there's competition from other providers now so they have an instinct to keep implementation details secret, but as someone building on their APIs this lack of documentation really holds me back.
To make good judgements about how to use this stuff I need to know how it works!
I had a hilarious bug a few weeks ago where I loaded in a single image representing multiple pages of a PDF and GPT-4 vision effectively hallucinated the contents of the document when asked to OCR it, presumably because the image was too big and was first resized to a point where the text was illegible: https://simonwillison.net/2024/Apr/17/ai-for-data-journalism...
If OpenAI had clear documentation about how their image handling works I could avoid those kinds of problems much more effectively.
I've also seen some people recommend Paddle OCR, but I find their documentation to be lacking and I haven't got that one working yet to evaluate.
If you make your system and it "works", then how will you see the one time out of X where it confidently provides you false information that you happily use because it usually work ?
You don’t. You treat it like you would a human worker: set your process to detect or tolerate wrong output. If you can't, don't apply this tool to your work.
If you need to be able to reason about multiple objects in the image and their relative positions, then don't you need to use a tiled approach?
You cannot have this level of control by prompting Dalle, also GPT-4o isn't using Whisper (older GPT-4s yes).
Another possibility is that OpenAI is mapping each image to ~170 vectors in an embedding space that is shared with token IDs. If that's the case, the architecture of the image-to-fixed-number-of-tokens model has not been disclosed. It could be a standard CNN, a ViT-like model, an autoencoder, a model that routes a variable number of vectors with RGB data to a fixed number of vectors, or something else that has not yet been ublished. The whole thing is likely trained end-to-end.
Yes, I agree.
Further down the road, I imagine we will end up finding interesting connections to the symbolic approaches of GOFAI, given that the embedding of a token, object, concept, or other entity in some vector space is basically a kind of symbol that represents that token, object, concept, or entity in that vector space.
Interestingly, old terms like "representation" and "capsule," which didn't become as widely adopted as "embedding," tried more explicitly to convey this idea of using vectors/matrices of feature activations to stand in for objects, concepts, and other entities.
For example, see Figure 1 in this paper from 2009-2012: http://www.cs.princeton.edu/courses/archive/spring13/cos598C... -- it's basically what we're talking about!
A pyramid of overlapped tiling resolutions is of course possible too.
What I found is that you can sometimes get more text out of 85-token image than you can out of 85 tokens of text! That said, I think there will be plenty of edge cases where it actually loses some information, and maybe you could argue that if you remove every other word in the text, it could still restore the text.
I never went deeper on this, but I believe there's something clever to be done in the context window with the fact that images are relatively cheap tokens-wise.
https://museumofbadart.org/zoo/
I found it while Googling for a test image for the malicious prompt test, where it was used as the lead photo for this blog post:
https://www.artsy.net/article/artsy-editorial-bad-art-good
There's definitely something eye-catching about it that really makes it stand out from the crowd.
The author is speculating about an embedding model but in reality they're speculating about the image-tokenizer.
If I'm not wrong the text tokenizer Tiktoken has a dictionary size of 50k. The image tokenizer could have a very large dictionary size or a very small dictionary size. The 170 tokens this image tokenizer generates might actually have repeating tokens!
EDIT: PS. What I meant to say was that input embeddings do not come from another trained model. Tokens come from other trained models. The input embedding matrix undergoes back propagation (learning). This is very important. This allows the model to move the embeddings of the tokens together or apart as it sees fit. If you use embeddings from another model as input embeddings, you're basically adding noise.
But why only choose 13x13 + 1? :(
I'm willing to bet that the author's conclusion of embeddings coming from CNNs is wrong. However, I cannot get the 13x13 + 1 observation out my head though. He's definitely hit on something there. I'm with them that there is very likely a CNN involved. And I'm going to put my bet on the final filters and kernel are the visual vocabulary.
And how do you go from 50k convolutional kernels (think tokens) to always 170 chosen tokens for any image? I don't know...
If the external-model also undergoes training along with the model then I think that might work.
It's easy enough to get it to hallucinate those things. It doesn't actually tell them to you.
Models are improving in truly hiding or ignoring information these days though. As the author of the article states, you'll have a hard time tricking GPT-4o to read text in images as instructions, most likely thanks to this research: https://openai.com/index/the-instruction-hierarchy/
I do feel pretty confident that when the model is happily spitting out its system prompt, and all metadata around the image, but not its pixel dimensions, that probably those dimensions were not provided in any system/assistant/tool message. So maybe part of the image embeddings also encode the pixel dimensions somehow (it would also help the model not think of the image as a squished square for non-1:1 images that have been resized to 512x512).
Single token?
> GPT-4o must be using a different, more advanced strategy internally
Why
There is no way that this is Tesseract.
-> Tesseract accuracy is very low, it can barely do OCR on printed documents.
For example, GPT4 with some vision capability would be able to fill in the incorrect OCR with the additional word co-occurrence understanding.
I've tested this approach with purely text LLM to correct OCR mistakes and it works quite well.
Also note that in some newer OCR pipelines that don't involve LLMs, there is a vision component and then a text correcting model that is in some ways similar to some forms of spell check, which can further improve results.
multimodal llm would of course blow it all out the water, so some llama3-like model is probably SOTA in terms of what you can run yourself. something like https://huggingface.co/blog/idefics2
I doubt they are running an OCR model, but if they actually were it would likely be an in-house one trained with more modern techniques.