New algorithm discovers language just by watching videos
csail.mit.edu
csail.mit.edu
We were unfortunately young and in poor company, so the idea didn't receive the appropriate attention.
[1]: https://en.wikipedia.org/wiki/Qualia#Inverted_spectrum_argum...
Blind people surely compensate for the lack of visual information by deliberately eliciting audio and tactile feedback that others don't need.
Also, watching others interact with the world is never the same thing as interacting with the world ourselves, because there's a crucial piece of information missing.
When we decide to act, we know that we just made that decision. We know when we made the decision, why we made it and what we wanted to achieve. We can never know that for sure when we observe others.
We can guess of course, but a lot of that guessing is only possible because we know how we would act and react ourselves. A machine that has never intervened and deliberately elicited feedback cannot even guess properly. It will need incredible amounts of data to learn from correlations alone.
But AI isn't an animal so the same constraints don't necessarily apply. I think you'd have to have a particularly anti-AI & pedantic bent to complain about calling this language discovery.
Feedback is fundamental to deep neural networks, it's how they're trained. And to be honest, all of the things 'astromaniak mentions can be simulated and made part of training data set too. While the "full experience" may turn necessary for building AGI, the amount of success the field had with LLMs indicates that it's not necessary for the model to learn language (or even to learn all languages).
Some of that isn't learned at all (other than through evolution). A newborn will flinch back from a growing circle on a screen.
There is a chance that abstract concepts can't be figured out visually ... although that seems outrageously unlikely since all mathematical concepts we know of are communicated visually and audibly. If it gets more abstract than math that'll be a big shock.
Ask someone with an articulation disorder. /s
Still, this is super reductive. Language can be feeling, it can be taste, it can be all sorts of things that are ineffable.
It is weird to me that the research into this assumes there is a baseline of repeatability somehow. At best, this is mimicry, imo.
It’s science done backwards which isn’t really science at all. Not that I think these models have no use cases, they’re simply being used too broadly because the people obsessed with them don’t want to admit their limitations.
> The idea of linguistic relativity, known also as the Whorf hypothesis, [the Sapir–Whorf hypothesis], or Whorfianism, is a principle suggesting that the structure of a language influences its speakers' worldview or cognition, and thus individuals' languages determine or influence their perceptions of the world.
Does language fail to describe the quantum regime with which we could have little intuition? Verbally and/or visually, sufficiently describe the outcome of a double-slit photonic experiment onto a fluid?
Describe the operator product of (qubit) wave probability distributions and also fluid boundary waves with words? Verbally or visually?
I'll try: "There is diffraction in the light off of it and it's wavy, like <metaphor> but also like a <metaphor>"
> If it gets more abstract than math
There is a symbolic mathematical description of a [double slit experiment onto a fluid], but then sample each point in a CFD simulation and we're back to a frequentist sampling (and not yet a sufficiently predictive description of a continuum of complex reals)
Even without quantum or fluids to challenge language as a sufficient abstraction, mathematical syntax is already known to be insufficient to describe all Church-Turing programs even.
Church-Turing-Deutsch extends Church-Turing to cover quantum logical computers just: any qubit/qudit/qutrit/qnbit system is sufficient to simulate any other such system; but there is no claim to sufficiency for universal quantum simulation. When we restrict ourselves to the operators defined in modern day quantum logic, such devices are sufficient to simulate (or emulate) any other such devices; but observed that real quantum physical systems do not operate as closed systems with intentional reversibility like QC.
For example, there is a continuum of random in the quantum foam that is not predictable with and thus is not describeable by any Church-Turing-Deutsch program.
Gödel's incompleteness theorems: https://en.wikipedia.org/wiki/G%C3%B6del's_incompleteness_th... :
> Gödel's incompleteness theorems are two theorems of mathematical logic that are concerned with the limits of provability in formal axiomatic theories. These results, published by Kurt Gödel in 1931, are important both in mathematical logic and in the philosophy of mathematics. The theorems are widely, but not universally, interpreted as showing that Hilbert's program to find a complete and consistent set of axioms for all mathematics is impossible.
ASM (Assembly Language) is still not the lowest level representation of code before electrons that don't split 0.5/0.5 at a junction without diode(s) and error correction; translate ASM to mathematical syntax (LaTeX and ACM algorithmic publishing style) and see if there's added value
"When CAN'T Math Be Generalized? | The Limits of Analytic Continuation" by Morphocular https://www.youtube.com/watch?v=krtf-v19TJg
Analytic continuation > Applications: https://en.wikipedia.org/wiki/Analytic_continuation#Applicat... :
> In practice, this [complex analytic continuation of arbitary ~wave functions] is often done by first establishing some functional equation on the small domain and then using this equation to extend the domain. Examples are the Riemann zeta function and the gamma function.
> The concept of a universal cover was first developed to define a natural domain for the analytic continuation of an analytic function. The idea of finding the maximal analytic continuation of a function in turn led to the development of the idea of Riemann surfaces.
> Analytic continuation is used in Riemannian manifolds, solutions of Einstein's [GR] equations. For example, the analytic continuation of Schwarzschild coordinates into Kruskal–Szekeres coordinates. [1]
But Schwarzschild's regular boundary does not appear to correlate to limited modern observations of such "Planc relics in the quantum foam"; which could have [stable flow through braided convergencies in an attractor system and/or] superfluidic vortical dynamics in a superhydrodynamic thoery. (Also note: Dirac sea (with no antimatter); Godel's dust solutions; Fedi's unified SQS (superfluid quantum space): "Fluid quantum gravity and relativity" with Bernoulli, Navier-Stokes, and Gross-Pitaevskii to model vortical dynamics)
Ostrowski–Hadamard gap theorem: https://en.wikipedia.org/wiki/Ostrowski%E2%80%93Hadamard_gap...
> For example, there is a continuum of random in the quantum foam that is not predictable with and thus is not describeable [sic]
From https://news.ycombinator.com/item?id=37712506 :
>> "100-Gbit/s Integrated Quantum Random Number Generator Based on Vacuum Fluctuations" https://link.aps.org/doi/10.1103/PRXQuantum.4.010330
> The theorems are widely, but not universally, interpreted as showing that Hilbert's program to find a complete and consistent set of axioms for all mathematics is impossible.
If there cannot be a sufficient set of axioms for all mathematics, can there be a Unified field theory?
Unified field theory: https://en.wikipedia.org/wiki/Unified_field_theory
> translate ASM to mathematical syntax
On the utility of a syntax and typesetting, and whether it gains fidelity at lower levels of description
latexify_py looks neat; compared to sympy's often-unfortunately-reordered latex output: https://github.com/google/latexify_py/blob/main/docs/paramet...
It bit me for a while later on in grade school when I took an English class and realized I didn't know anything about how language was constructed, it was all purely intuition. Formalizing language took a concerted effort on my part, and my teachers didn't understand how to translate the material... because I can't just be told things at face value, there is always a nested mass of "but why?" that must be answered or I fundamentally don't understand.
Once I finally got over that hill, it was smooth sailing again. It also made learning foreign languages very hard in school, I failed both foreign language classes I took, until I again took my own time to learn the fundamentals and now it all seems to stick.
If I understand the paper correctly, the model was trained on tens of thousands of video and audio clip pairs, randomly sampled 64 million times (80 samples/batch x 0.8 million training steps). For comparison, over the course of a baby's first two years of life, the baby will get trillions of sound samples (measured at 44.1KHz) and billions of frames of visual data (measured at 30fps). It's not too hard to imagine that a large AI model in the near future, trained on comparably vast quantities of audio and video data, can learn to recognize objects and words, like a baby.
The most exciting takeaway, for me, is that a bigger model inside a robot body should be able to integrate learning in this manner across all five human data modalities -- vision, hearing, smell, touch, taste -- as well as any other data modalities for which the robot has sensors -- radar, lidar, GPS position, etc. We sure live in exciting times!
Link to paper, code, demo, and data:
still several of OOM difference
I agree. Nothing I wrote above disagrees with your statement.
Is memory really an expected bottleneck in the long term?
There should be an app that pays people to dub popular shows in a broader way of language, charges viewers, and has some kind of sidebar or subtitle integration to hover/click and learn more, check what a word was and conjugation etc.
The focus on language learning would also allow the translators to make more appropriate choices for that use, whereas subtitles (and probably dubbing) for entertainment often has some variation that is confusing it to learn. I suppose that matters less if there's only one or the other, so no chance of conflict.
Because it's so stupid even some dumb code can understand the patterns. Wouldn't work with Seinfeld.
It doesn't really have much in common with LLMs. If the ideas were combined, the results might be interesting, although probably requiring very significant compute.
Funnily enough though, the comprehensible input movement for language learning has already started generating the type of content you're suggesting. Long first person view walks, following day to day activities and usually with a monologue about what they are seeing inbetween.
I think my son's learning Japanese by watching Anime on Crunchyroll.
Also, the ability to understand language is a spectrum, not a binary quality.
Algorithms necessarily contain some knowledge about the data they are working with so that isn't cheating. Neural nets for example have to have some architecture and that predisposes them to learning certain patterns.
Chomsky's argument is exactly the opposite, that vocabulary is learnt from scratch, but that the framework for grammar is innate.
Rather conclusive one already established in the ML field.