Visual Reasoning Is Coming Soon
arcturus-labs.com
arcturus-labs.com
For visual reasoning practice, we can do supervised fine-tuning on sequences similar to the marble example above. For instance, to understand more about the physical world, we can show the model sequential pictures of Slinkys going down stairs, or basketball players shooting 3-pointers, or people hammering birdhouses together....
But where will we get all this training data? For spatial and physical reasoning tasks, we can leverage computer graphics to generate synthetic data. This approach is particularly valuable because simulations provide a controlled environment where we can create scenarios with known outcomes, making it easy to verify the model's predictions. But we'll also need real-world examples. Fortunately, there's an abundance of video content online that we can tap into. While initial datasets might require human annotation, soon models themselves will be able to process videos and their transcripts to extract training examples automatically.
Almost every video generator makes constant "folk physics" errors and doesn't understand object permanence. DeepMind's Veo2 is very impressive but still struggles with object permanence and qualitatively nonsensical physics: https://x.com/Norod78/status/1894438169061269750Humans do not learn these things by pure observation (newborns understand object permanence, I suspect this is the case for all vertebrates). I doubt transformers are capable of learning it as robustly, even if trained on all of YouTube. There will always be "out of distribution" physical nonsense involving mistakes humans (or lizards) would never make, even if they've never seen the specific objects.
Is that why the peekaboo game is funny for babies? The violated expectation at the soul of the comedy?
https://en.wikipedia.org/wiki/Object_permanence#Contradictin...
That contradicts this contradiction, unless there is another explanation.
"Why were you throwing the baseball at the running fan?" "I don't know...I was bored."
I don’t understand how you learned this. Who told you that’s what was going on in his head?
A paper from 2014 is not so sure this is a settled matter:
"Infant object permanence is still an enigma after four decades of research. Is it an innate endowment, a developmental attainment, or an abstract idea not attributable to non-verbal infants?"
and argues that object permanence is not in fact innate but developed:
"It is argued that object permanence is not innately specified, but develops. The theory proposed here is that object permanence is an attainment that grows from a developmentally prior understanding of object identity"
"In sum, it is posited that permanence is initially dependent on the nature of the occlusion; with development, it becomes a property of objects. Even for 10–12-month-olds, object permanence is still a work-in-progress, manifested on one disappearance transformation but not another."
Peekaboo is fun because fun is fun. When doing peekaboo the other person is paying attention to you, and often smiling and being relaxed.
They laugh just as much if you play ‘peekaboo’ without actually covering your face ;)
I said that learning visual reasoning from video is probably not enough: if you claim it is enough, you have to reconcile that with failures in Sora, Veo 2, etc. Veo 2's problems are especially serious since it was trained on an all-DeepMind-can-eat diet of YouTube videos. It seems like they need a stronger algorithm, not more Red Dead Redemption 2 footage.
Fair enough; you did indeed say that.
> if you claim it is enough, you have to reconcile that with failures in Sora, Veo 2, etc.
This is flawed reasoning, though. The current state of video generating AI and the completeness of the training set does not reliably prove that the network used to perform the generation is incapable of physical modeling and/or object permanence. Those things are ultimately (the modeling of) relations between past and present tokens, so the transformer architecture does fit.
It might just be a matter of compute/network size (modeling four dimensional physical relations in high resolution is pretty hard, yo). If you look at the scaling results from the early Sora blogs, the natural increase of physical accuracy with more compute is visible: https://openai.com/index/video-generation-models-as-world-si...
It also might be a matter of fine-tuning training on (and optimizing for) four dimensional/physical accuracy rather than on "does this generated frame look like the actual frame?"
As for object permanence, I don't know jack about animal cognitive development, but it seems important that all animals are themselves also objects. Whether or not they can see at all, they can feel their bodies and sense in some way or other its relation to the larger world of other objects. They know they don't blink in and out of existence or teleport, which seems like it would create a strong bias toward believing nothing else can do that, either. The same holds true with physics. As physical objects existing in the physical world, we are ourselves subject to physics and learn a model that is largely correct within the realm of energy densities and speeds we can directly experience. If we had to learn physics entirely from watching videos, I'm afraid Roadrunner cartoons and Fast and the Furious movies would muddy the waters a bit.
I found that when editing images of myself, the result looked weird, like a funky version of me. For the cat, it looks "more attractive" I guess, but for humans (and I'd imagine for a cat looking at the edited cat with a keen eye for cat faces), the features often don't work together when changed slightly.
Here's the best out of a few attempts for a really similar prompt, more detailed since Flash is a much smaller model "Give the cat a detective hat and a monocle over his right eye, properly integrate them into the photo.". You can see how the rest of the image is practically untouched to the naked human eye: https://ibb.co/zVgDbqV3
Honestly Google has been really good at catching up in the LLM race, and their modern models like 2.0 Flash, 2.5 Pro are one of (or the) best in their respective areas. I hope that they'll scale up their image generation feature to base it on 2.5 Pro (or maybe 3 Pro by the time they do it) for higher quality and prompt adherence.
If you want, you can give 2.0 Flash image gen a try for free (with generous limits) on https://aistudio.google.com/prompts/new_chat, just select it in the model selector on the right.
On the other hand, OpenAI's model, while it does seem to have some upscaling magic happening (which makes the outputs look a lot nicer than the ones from Gemini FWIW), also seems to perform all its edits entirely in latent space (hence it's easy to see things degrade at a conceptual level such as texture, rotation, position, etc.) But this is a sign that its latent space mode is solid enough to always use, while with Gemini 2.0 Flash I get the feeling when it is used, it's just not performing as well.
The cat's facial hair coloring is entirely different in your image, far more so than the OpenAI one. It has a half white nose instead of a half black nose, the ears are black instead of pink, the cheeks are solid colors instead of striped. Yours is the head of an entirely different cat grafted onto the original body.
The OpenAI one is like the original cat run through a beauty filter. The eyes are completely different. But it's facial hair patterning is matched much better.
Neither one of great though. They both output cats that are not recognizable as the original.
Hm, no, I’ve never had this thought.
I think it's decent of the author to provide a video for these people.
Of course the devil's always in the details and there have been real non-obvious advancements at the same time.
Mixture of experts tends to lock models up. Chain of thought tends to turn the models loopy. As in circling the drain even when asked not to. Tree of thought has a tendency to make the models unstable by flipping between the branches, and otherwise unpredictable complicating training ...
the author claims that visual reasoning will help the model solve this problem, noting that gpt-4o got the question right after making a mistake in the beginning of the response. i asked gpt-4o, claude 3.7, and gemini 2.5 pro experimental, who all answered 100% correctly.
the author also demonstrates trying to do "visual reasoning" with gpt-4o, notes that the model got it wrong, then handwaves it away by saying the model wasn't trained for visual reasoning.
"visual reasoning" is a tweet-worthy thought that the author completely fails to justify
The example you used to demonstrate is well done.
Such a simple way to describe the current issues.