Let me clear a huge misunderstanding
twitter.com
twitter.com
So it's strange to me to claim that it doesn't have an abstract representation.
But, maybe the latent space of a diffusion-VQVAE pipeline is fundamentally different from that of JEPA, I haven't read the relevant papers for that. Curious if someone could explain if they are different ideas of representation.
I’m a little oversubscribed at the moment so I haven’t downloaded the weights and played around with V-JEPA. The fact I’m pointing out that I could should make it pretty clear which way I lean on this: everyone wants to make money, let’s be real, but some seem resigned to “infinite, government-enforced monopoly via OpenPhilanthropy bribery” isn’t going to fly, so let’s at least prove how the sausage is made.
It’s fairly uncontroversial that things in the neighborhood of a “bottlenecked” VAE are often forced to exploit structure if it’s there.
This (claimed) result is about a way to exploit more of this structure/symmetry economy (greedy-ish optimizer) to pull the latent representations into yet a higher “effective regime” than yet demonstrated, with excellent properties around machine economics.
Representation learning isn’t new (though LeCun is an acknowledged pioneer in it), and constraints as a powerful tool isn’t new (causal making in attention architectures shouldn't ruffle many feathers).
Self/semi/unsupervised learning isn’t novel. But likewise, not are they k-equivalent synonyms: Dean et al. were distinguishing between Continuous Bag of Words and skip-gram in word2vec in 2013.
But it does cool shit (I like the I-JEPA reconstructions as a go to slide deck raster), and it’s available weight.
I think the truth shall soon emerge.
It is kind of ironic that researchers who claim LLMs lack adaptive intelligence seemingly refuse to adapt their intelligence to LLMs. If even GPT-3 can find logical holes or oversimplification in your arguments about GPTs, at one point this starts becoming embarrassing and unbecoming.
> The generation of mostly realistic-looking videos from prompts does not indicate that a system understands the physical world.
While arguably true, it also does not indicate that a system does not understand the physical world (reflections, collision detection, gravity, object permanence, long-term scene coherence, etc.).
If LeCun wants to argue it does not understand the physical world, he should do so directly. Not attack something that is not directly stated, but rather convincingly and tentatively demo'd (I myself find it hard to argue that a system that generates novel pond reflections has not memorized/stored in weights some generalization program to apply to realistic scene generation).
This demo shows it is not even a wild prediction to guess that soon (consumer tech) we will be able to discuss visual scenes with conversational AIs.
FWIW none of the video models released so far demonstrate any object coherence whatsoever, which suggests they don't have the higher level capabilities you mention yet.
In Sora, as soon as an object is obstructed by an obstacle or goes offscreen, it's likely to disappear or be radically transformed.
Or the video with the dog, that dog phases straight through those window shutters as if they weren't there and were rendered in layers rather than 3d. It doesn't understand the scenes it draws at all, it had shadows from those shutters so they were drawn to have depth, but that dog then were rendered on top of those shutters anyway and moved straight through them. You even see their shadows overlap since the shadow part is treated differently apparently, so it "knows" they overlap but also renders the dog on top, telling me that it doesn't really know any of that at all and is just based on guessing based on similar looking data samples.
And this in videos handpicked because they were especially good. We should expect the videos we are able to generate to be way worse than the demo in general. They didn't even manage to make a dog that moves between windows without such bugs, that was the best they got and even that was had a very egregious error for a very short clip.
Kind of made me wonder how these videos would look run backwards, but not enough to figure out how to make them run backwards.
EDIT: wow, the "backwards" physics is especially noticeable in the chair video[0]. Aside from the chair morphing wildly, notice how it floats and bounces around semi-physically. Clearly some issues grasping cause and effect.
https://www.youtube.com/watch?v=ezaMd4l_5kw
We also have the spontaneous creation and annihilation of wolves and the shape-shifting chair:
The more I watch the cherry blossom one, the more I see how wrong it is, even the fact there is Cherry blossoms in the middle of winter is just totally wack. I've seen it snow in Tokyo before during spring when the cherry blossoms were out, but you don't have a foot of snow on the roof like in the clip.
Edit: I know the prompt asked for the cherry blossoms in snow, but it's still a wild amount of snow which is somehow not covering the trees.
Someone correct me, but it seems like he's saying:
1- generative models don't understand the real world
2- generative models that work off of just pixels are more expensive and less useful than a model that represents the contents of the frame with abstract representations, particularly when related to higher-level actions like applying logic to a scene and not just interpolating between frames.
3- V-JEPA is such a model, and it performs particuarly well against generative when combining those representations as inputs with a much more easily trained and small model trained on specific tasks.
So go ahead and define it, in concrete terms, external to humans. It can't be equivalent unless there is a definite basis for equivalence. Cat videos don't cut it.
My point is that understanding, as we know it, only exists in the human mind.
If you want to define something that is functionally equivalent but implemented in a machine, that is absolutely fine, but don't point to something a machine does and say "look it's understanding!" without having a concrete model of what understanding is, and how that machine is, in concrete terms, achieving it.
Fine. Then I posit the existence of understanding-2, which is exactly identical to understanding, whatever it is, except for the fact that it can only be done by machines. And now I ask you to prove to me that AI doesn't have understanding-2.
This is just to show you the absurdity of trying to claim that AI doesn't have understanding because by definition only humans have it.
He said understanding is what humans do, not that only humans can do it. Stop arguing against a strawman.
Nobody would define understanding as something only humans can do. But it makes sense to define understanding based on what humans do, since that is our main example of an intelligence. If you want to make another definition of understanding then you need to prove that such a definition doesn't include a lot of behaviors that fails to solve problems human understanding can solve, because then it isn't really the same level of as human understanding.
Ok, so his argument is:
> Humans understand, models are deterministic functions
> Until you have a concrete definition of what understanding is you can't apply it to anything else
> Informal definitions of understanding by those who experience it aren't very useful at all.
Basically he says: "I don't accept you using the term 'understanding' until you provide a formal definition of it, which none of us has. I don't need such definition when I talk about people, because... I assume that they understand".
Which means: given two agents, I decide that I can apply the word "understanding" only to the human one, for no other reason that it is human, and simply refuse to apply it to non-humans, just because.
Clearly there is absolutely nothing that can convince this person that an AI understands- precisely because it's a machine. Put in front of a computer terminal with a real person on the other side- but being told it's a machine- he would refuse to call "understanding" whatever the human on the other side does. Which makes the entire discussion rather pointless, don't you think?
But they're very limited, and if you prompt with a relationship that isn't defined you get best-guess, which will either be quite distant or contaminated with other values.
If you ask Dall-E for "a woman made of birds" you get a composite that also includes trees and/or leaves. Dall-E has values for "made of" and "birds" but its value representation for "birds" is contaminated with contextual trees and branches.
Leonardo doesn't have a value for "made of", so you get a woman surrounded by bird-like blobs.
To understand a cat in a human sense the store would have to include the shape, the movement dynamics in all possible situations, the textures, and a set of defining behaviours that is as complete as possible. It would also have to be able to provide and remember an object-constant instantiation of a specific cat that is clean of contamination.
SORA is maybe 10% of the way there. One of the examples doing the rounds shows some puppies playing in snow. It looks impressive until you realise the puppies are zombies. They have none of the expressions or emotions of a real puppy.
None of this is impossible, but training time, storage, and power consumption all explode the more information you try to include.
Also the complain about 'made of' not being in the training data. Humans who never saw a bird can not draw a bird. Why is that saying something about the model?
I'm not saying that diffusion models act like humans. And I was talking specifically about image generation. My usage of the word understanding is in the task of image generation. I'm not even talking about 'made of', or 'birds'. Just 'cats' and 'hats'. If it can understand 1 thing, it can understand others, but they are not always in the training data.
This is all a non-problem. It kinda remind me of the discussion of what constitutes a 'male', or 'female'. All i want is to refer this one property that i observe in diffusion models. Which is what language is, reference. If you are so covetous of the word 'understand', then provide an alternative to refer to this property and i will gladly use it.
And how do you know that this is not what "understanding" is? To me, understanding the concept of a cat is exactly to immediately recall (or have ready) all the associations, the possibilities, the consequences of the "cat" concept. If you can make up correct sentences about cats and conduct a reasonable conversation about cats, it means that you understand cats.
No, plenty of humans can have reasonable conversations about things with zero understanding about them. We know they don't understand because when put in a situation in practice they fail to use the things they talked about. Understanding means you can apply what you know, not only talk about it.
How can Sora predict where the waves and ripples should go? Is it just "correlations not causal", whatever that means.
I'm no math wiz and my training in statistics is severely lacking but it feels like people need to review what they think it's possible with Generative AI because we are so far from understanding and AGI that my head hurts every time these words show up in a discussion.
1. The next frame is easy, but multiple frames is not
2. What works for text doesn't work for video.
Then Sora comes out and shows multiple frames and someone tweets gotcha.
He then tweets without saying he misspoke ..... goes on about the model doesn't understand physics.
And his project, V-JEPA, is the best
He keeps saying stuff about "sucks as a mental model" but doesn't say why that would not apply to text.
https://twitter.com/ylecun/with_replies
Me: If text doesn't need a mental model, I see no reason video needs it. His argument sucks or is badly worded.
That's a bit like saying the chatbots aren't actually intelligent. Sure, but there is at least a plausible illusion of it & even that has usefulness.
The same applies to physical world. e.g. Look at the reflections in the SORA demo. They're not right, but they're also not entirely wrong either. That to me suggests usefulness in approximating the physical world
Like, in general I think there's a lot of hype around AGI and that skepticism toward OpenAI's claims isn't completely unwarranted, but the amount of public attention on the topic lately has caused everyone to commit to very hard lines that lack nuance about the implications of various advances. It's in some ways cool that AI is no longer just an academic curiosity, but it makes for a lot of nonsense to slog through, even from top researchers
HN gets hung up on damning things that aren't perfect _right now_
You also see this with FSD...it isn't perfect today, so HN writes it off forever
Sora is a demo and a teaser of what will be a useful polished tool in three years, that's all
People have bought, owned and used, and sold their Tesla w/ FSD without once being able to use the feature.
Its not even any better or less relaxing - instead of driving you babysit a black box.
I think he's claiming that the misunderstanding is that SORA understands the physical world. He goes on to say that generating the next frame conditional on some action is much harder than generating an entire plausible video. This just doesn't make sense to me. Every frame in a plausible video clearly needs to be conditioned on the actions shown in previous frames.
The tweet is mostly incoherent and I'm left assuming he's upset about SORA and wants to say his ideas are/were better but hasn't managed to express why in a meaningful way.
It's like, the smarter someone gets, the more it trips them up when it comes to questions like the one Yann is trying to address. Because, "world modeling" doesn't have to mean that the system understands advanced physics and calculus. Look at our own brains – only analogies, I'm aware – and think about how much a child can infer about what will happen when you throw a ball at them.
Are they 'calculating' trajectories? Not consciously. That would be too slow, regardless. Their brains have become wired through experience and evolution. Formally explaining why the ball will do what it does? Well, that takes years of math and physics education.
Bottom line: in the months and years after a model comes out, research tends to uncover all sorts of things happening inside them that are unexpected. Best thing is to approach it with a beginner's mind, and say, "hey, it's doing _something_ interesting – let's try and understand what that is," rather than just forcing everything to conform to our existing worldview.
LeCun is the voice of reason trying to point out the limitations of the current technology and fundamental questions that need to be answered.
Until there is something more than cute demos, like an actual path forward that can implement what people are hyping, or better still working examples, it's all speculative nonsense.
I think this is the more likely true, but less popular take as well.
Why would it be pointless? I think there are many creative uses of video continuation model, considering the amount of control it gives you.
To me, it's clear that there's no real global understanding yet. Whether this person's preferred ML approach solves that, I'm not knowledgeable enough to tell.
Definitely felt like something was off to me.
Firstly his insistence on “self supervised learning” which is just a wrong and unhelpful rebranding of existing methodologies. Followed by talking about VicREG as if it’s a meaningful contribution and not just hacked together crap which is not only theoretically unfounded but plain nonsensical. Followed again by his “JEPA” work which again is just a rebranding of other works like BYOL and data2vec.
Its so frustrating to see this guy pop up and try to claim credit for an entire field of representation learning which he hasn’t made a single meaningful contribution to in decades.
I’ve met the guy a few times (briefly), and I’m aware of his general vibe. I don’t agree with him about everything, but “hack” is absurd, and I’m not posting about who is a hack (or even a crook) under a burner alt.
What’s your Fields medal for?
Many of us actually work in the field and are tired of the mindless hype and self promotion. LeCun is just upset people aren’t hyping his crap rather than OpenAIs.
But I’m sure you can give a good reason why he calls it self supervised instead of unsupervised and JEPA instead of BYOL or data2vec and what his actual contributions to the field of modern representation learning are.
Pauling had a Nobel, didn’t stop him being a crank about vitamin C.
I’m saying that it takes some cheek to say there’s nothing original about results that many experts (who are nothing to do with me, and me in a narrow way) think are pretty hot shit.
I try (and sometimes fail) to confine my strident claims requiring strident evidence to topics I’m willing and able to get into the nitty gritty on, and bring my CV and reputation to the party as a courtesy. Anyone can post biased-sounding shit and hand-wave over the substance of the claim under an 11-month burner named with malice aforethought to talk shit and be able to disavow it.
I gave you a rather firm tap on the shoulder about how borderline at best this move is, and under the assumption you’ve got an HN main, you know this.
I had hoped you might put the effort that went into doubling down into explaining the apparently obvious prior art, the whataboutism kick flip into the Shockley Race IQ trope made zero people more enlightened and steered a contentious thread even further off the rails. Nobel laureates do in fact go off the friggin rails from time to time, but I don’t think LeCun is an interesting comparison to Kary Mulis past “AM Radio yelling”. I’m calling that one out.
Maybe elaborate on a pretty debatable claim instead?
Like literally just read the papers I mentioned. Or don't, I don't care. I have literally no clue who you are so maybe take your creepy-as-fuck "firm tap on the shoulder" and wanting to see my "cv" shit somewhere else.
I generally appreciate the feedback that I’ve landed there.
This isn’t seeming to be going anywhere useful and I’m going to respectfully bow out. If you’ve got to have the last word, go to town: but do it by posting the links to the papers for those who haven’t read them.
I’ve pulled some dumb stunts on HN, even some recently, but I’ll stand behind there being no Navy Seal Copypasta flex here for the record.
Be well ml-anon.
What are those methodologies?
All these stupid new terms of trivial concepts confuse people who are new to the field. That's nothing new though, research has always been like this, people love to make up new phrases that they can own and sell as novel ideas to reviewers, even if they're just trivial renamings. It's nothing more than a PR game.
That's how language works.
In fact there is a direct formal equivalence between many so-called self supervised learning techniques and matrix factorization. And literally no one would claim that matrix factorisation is anything other than unsupervised.
Similarly Yann makes such a big deal about not doing contrastive learning when contrastive methods and his JEPA nonsense can also be shown to be formally equivalent. He’s a grifter like the rest of them. There’s a reason why he basically holds no power at Meta and hasn’t been in charge of AI research there in a long time.
We shouldn't censor ourselves and not discuss information that was posted on twitter.
So this type of content means that in order to follow the discussion I must also be a twitter subscriber.
Maybe people could consider first creating an archive.org snapshot that is then linked instead. Would be the kind thing to do and help the rest of us.
Yes, that would be the best way to handle this.
If it’s a whole thread though then just don’t bother. That approach to blogging is just dumb.
It made Trump president and gives Musk control of the media narrative though so it certainly works for some people.
This link opens the post for me, while paywalled news is semi-routinely posted.
What’s the objection?
In other words, paywalled news is a first party content provider, but Twitter is just a third party who inject their mandatory login between me (the reader) and the writer so they can make some profit by showing ads on the content.
That's fine, but authors should consider if they want to publish on a site like that when there are so many other options.
A big problem with machine learning so far has been the lack of some underlying, more abstract model of the subject matter. I don't understand much of this, but apparently something called "representation", a purely mathematical concept, is able to help with this.[2] This turns some kinds of abstract problems into linear algebra problems, which means matrix operations, the things GPUs do so well, help.
Code is available, from links in [2].
Would someone comment on this, please. Is this the beginning of a huge breakthrough? One where, instead of working in text token space or image pixel space, it's possible to work in an automatically generated more abstract space?
[1] https://scontent-sjc3-1.xx.fbcdn.net/v/t39.2365-6/427986745_...