I could be misreading this, but I hope the author doesn't think GANs are used in LLMs. They are cool, though.
I could be misreading this, but I hope the author doesn't think GANs are used in LLMs. They are cool, though.
But he’s made the error of trying to come at it from a technical perspective, when he clearly knows nothing about that side of things, which discredits the rest.
Huh? That quote appears to be the entirety of the "technical perspective" of the post and is an aside from his larger points that you have blessed as "fine". Literally nothing in the rest of the post relies on that incorrect statement.
Let me quibble with what is discredited here, given the entirety of your point is built upon an error.
> They're word-association mechanisms with no embodiment and no way to associate the text vectors they manipulate with real-world phenomena.
This is wrong too. RLVR grounds foundational models in reality.
> no way to associate the text vectors they manipulate with real-world phenomena.
Historically, that LLMs were text-only used to be a major argument for why they "lack access to meaning", see the Stochastic Parrot paper and the Octopus paper that it references. But even the authors of those papers have (grudgingly) conceded that the argument no longer holds due to multimodality.
AI is so good that if you see an hallucination as obviously wrong like this one it’s a sign it’s written by a human.
In general probably not much of a stretch to get rid of convolutions or recurrence by replacing with attention and see if it works hence the title of the original transformers paper.
Or "it's not 4-wheel-drive, it's diesel!", if familiar with cars.
Once Musk said he believes in Transformers instead of diffusion (or the other way around). When actually diffusion transformers are very popular and mainstream. What he meant was autoregressive inference vs diffusion. People like to use buzzwords while not knowing them. Old issue, it was the same decades ago.
To make it even more confusing, there are also diffusion LLMs, which typically but not necessary, use transformers also.
And independently of diffusion and transformer and llm or image generator, you can optionally put a GAN discriminator adversarial loss on any of them.
It's not like every model has the same amount of those shortcomings.
The first sentence seems to imply the average state. No?
The problem really is I haven't used these other AIs for finding links to things I am an expert in, so I cannot really say how likely they are to get things completely wrong, or at least have one wrong or misinterpreted fact per response.