Are language models deprived of electric sleep?
blog.cbs.dk
blog.cbs.dk
Someone is eventually going to speed run 10k karma on HN and YouTube it. This year? Next? Who knows. But not a decade.
[0] https://en.m.wikipedia.org/wiki/Do_Androids_Dream_of_Electri...
The line is blurry, though. Google uses machine learning too, so it’s possible that nobody on the planet has ever seen your particular search results for a given query, which feels a bit like synthesis, although the building blocks are bigger. And GPT-3 isn’t really synthesising at all, of course, but it’s a convincing enough illusion that synthesis is a helpful mental model.
Or at least, a first-level approximation. Wikipedia articles can be wrong.
For those saying it can generate new things, well it is interpolating between existing data points in the corpus, but mushed through a bunch of math.
By not mixing the articles together, a search engine lets you see where the information came from.
These are both useful things to do, but one is more useful for fiction and the other for nonfiction.
I'm not sure "interpolation" is the right word, though, for these mixtures. Transformer output seems more creative than that.
It is a straight line in a higher dimensional space. We vibe with the space, or it chooses us by prior experience. Creativity is in some aspect, in the eye of the beholder. I think the creativity we see in ML models is not unlike the creativity we prescribe to humans.
Interpolation is the right word and the right concept and metaphor. A mixing between two concepts to create a path between the two.
E.g
Ask "is pi disjunctive" and you'll get the top result of
> = 3.1415926535897932384626433832795... is a disjunctive number.
But if you click into the article (what percentage even will?) you'll find it's part of a larger statement:
> It is not known whether π = 3.1415926535897932384626433832795... is a disjunctive number.
Yes, some have argued that Google can’t make content and from a technological standpoint that’s right, but here’s the fulcrum: from the perspective of “I asked a question and got a result that seems right though I’m not sure”, there’s shockingly little difference between GPT-3 generated content, and low-effort SEO-optimized content found on the first page of Google.
often, this is also mechanically generated
When your game plan is take a multi-head attention transformer and overfit it on as much of the Internet as we have A100s to fit it on, you’re going to get a distribution, and if you sample from it, it’s going to sound like the Internet.
People read too much into this.
The technology is obviously powerful borderline dangerous, but it’s still just maximizing P(this-sentence-was-on-the-internet).
So it’s transformers plus let’s force some loss into the super-resolution pipeline.
“Residual” sounds a lot fancier than X = fn + X, and “Diffusion” sounds a lot fancier than “let’s do deconv repeatedly and jam a loss term in there”.
My point is all of this stuff is on Huggingface. You/I/We just don’t have 5k A100s for 6 months, so we can’t play.
Nonetheless, this approach of taking a fairly complex writing prompt and simply regurgitating contextually similar text from the internet often yields 100% meaningful results. This suggests that our intelligence is a lot more mimetic than we'd like to believe, which is perhaps an even more unpopular opinion.
Homo Sapiens has a serious agenda around being something other than the smartest chimpanzee.
If we just admit the possibility that the math was right all along we win, the math works, and we lose “I am a unique, distinct entity.”
I’m on team math.
I tend to lean towards Sean Carrol’s approach of the interpretation is important and we should spend real effort thinking about the different impacts of different interpretations. After all, Einstein didn’t get to general relativity just by calculating. It first required deep thought and consideration to decide what to calculate. And I’m not convinced that the interpretations will forever remain in the realm of philosophy (though, I accept I could be totally wrong about this).
[1] I don’t really disagree with Feynman on this. I think it’s perfect advice for people getting started in really learning QM (which is beyond where I am!). Being very familiar with the mechanics and being able to calculate fluently are probably a good starting point for people learning QM. But I do think that it’s worth experts in the field spending some time beyond the mathematical fundamentals.
I agree with this sentiment. quantum mechanics (stemming out of statistics mechanics) start from regarding all particles as indistinguishable.
IMO, this theory is more mathematics than it is physics.
Given a wide enough application, humans end up ourselves as particles, essentially we are commoditifiying our individuality.
Each time I mention this online I state it more confidently, in the hope that some day someone will Cunningham's law me and change my mind. If I haven't committed some gross misunderstanding it's disappointing that so many physicists fall for such obvious bunk. I've already seen the attempted arguments listed at [1] and none of them are remotely convincing.
[1] https://en.wikipedia.org/wiki/Many-worlds_interpretation#Pro...
I’m not a physicist but I am an interested layman, I’m pretty sure I’ll catch wind of the Nobel Prize you’d win for proving him wrong about everything.
The mystery to me is why the obvious flaw in the idea hasn't killed it dead for so many smart people, such as David Deutsch. Instead they engage in bizarre logical contortions to try and recover the Born probabilities and some shadow of a meaning for them. If you think there's some actual merit in their attempts, listed at the wiki link I pasted, I'm all ears.
Fully deterministic, you see the part of the wave function that you see. You’re entangled with the apparatus.
Deutsch, and his protege Marletto go the whole way: no free will, no arrow of time, no subjective human experience at all.
Now? No measurement problem. No interpretations of collapse.
These people have decided that the subjective experience of observing an experiment is of secondary importance to clean math.
I appreciate that “you and I don’t exist in any way we’d recognize it” is a big pill to swallow, but I find the idea that humans looking in microscopes mutates the universe a bigger pill.
And the question here is one of interpretation. You say that MWI gives you "clean math" but this isn't about the maths. The maths of quantum mechanics is what it is regardless of how you interpret it. Whether the Schrödinger equation describes the evolution of the system state and observation is a state vector collapse, or the Schrödinger equation describes the evolution of a multiverse and the apparent state vector collapse is an artifact of distinct outcomes decohering, the maths is exactly the same. MWI's claim that it's somehow "mathematically cleaner" or "just what the maths is saying" is nonsense. It's there in the name. It's an interpretation.
QM is a highly effective theory (or rather framework for theories) that has been tested to the umpteenth degree. The question of intepretation is a philosophical one more that a mathematical one, and certainly fascinating. To my mind a "good" interpretation would have to give some actual meaning to the functional components of the theory. If QM predicts something will happen 30% of the time and I test it and it happens 30% of the time, an interpretation that fails to give a meaning to that predictable, observed fact is no use to me.
You seem to think the problem others have with your preferred interpretations is that they are too small-minded to accept that reality is the way these ideas imply. This isn't the case at all. I don't find "free will" an interesting concept and I'm perfectly happy to contemplate a multiverse, deterministic or otherwise, if that's what the scientific method leads us to. What I'm not happy to do is to launch off into a world of wacky, unfalsifiable notions that remove meaning from our existing theories rather than adding it.
I could equally well make claims about the psychology of MWI's adherents, suggesting emotional explanations for why they seem bound to make ever more absurd claims in defence of their idea rather than accepting that it's flawed. I won't, though, because doing that is arrogant and presumptuous.
I’ve got no agenda around Everett’s initial idea and monograph being the end state. To your point, there are some known flaws with his initial formulation.
I find the idea that QM needs an interpretation absurd, MWI was a step on the path to that way of thinking. At a high level, the notion that all possible outcomes of an observable-producing operator on the wave function are equally “real” is clarifying.
I tend to agree with Deutsch and Marletto that “that which is admissible is admitted”.
It’s possible we’re in violent agreement.
QM, and more specifically QEM and QCD make wildly accurate predictions. As long as no one is talking about a causal structure unique to “consciousness” or some ridiculous narcissism like that, I’m very satisfied to rely on the falsifiable and well-tested laws.
But when doing actual research, we check our work against the real world. For example, that's how you get a list of real references rather than fake ones.
Suppose we played a guessing game: given a title, does the Wikipedia article exist or not? You could fairly confidently say that "Apple" exists and "wjifdvq" does not, but given a plausible-looking name of a person or place that you don't recognize, you'd have a harder time. It's not a problem in practice though, because you can look it up.
It generates comments arguing that ML models are just parrots.
This strikes me as a crucial part of what intelligence (whether in humans or other creatures) means.
In case of an algorithm, there is no intent of its own—except the intent and the minds of humans who trained it, ran it and supplied inputs.
IMO the onus is on AGI believers to prove that material world is somehow the source of consciousness or intelligence (it wouldn’t hurt defining the terms first either). Otherwise it’s a philosophical position, and while one is entitled to hold their own one is not entitled to force it onto others.
True, although much less so for (pseudo)anonymous online communication, as we’re having here.
>It always matters who said it and why; we never take the substance independently from the agenda and the mind behind it, the context in which that mind existed and its relationship with our own, and so on.
I have no idea who you are, what your agenda is, or anything about how your mind operates. There is essentially zero relationship between your mind and my own. The only context I have about you is the text of your post.
>This strikes me as a crucial part of what intelligence (whether in humans or other creatures) means.
Totally agreed. This is why GPT-3 can do a fairly good job emulating anonymous online discourse, but cannot convincingly emulate a person we actually know.
Let’s imagine a slightly more down-to-earth exchange and raise the stakes in a way. For example, say it is still mostly a philosophical discussion here on HN, but in which someone is pushing a position on how much potential a certain company or store of value has; would you not question (at least to yourself) whether that person is invested and seeking to profit from it in short term, which would taint the motive? Or say someone was arguing in support of a controversial policy of a particular government known for its strong control over access to information and freedom of expression; would you not wonder whether the person is in fact a citizen of that country being misled by own government (and/or motivated to support it rather than seek truth)?
Yes, the substance of what they say may or may not be true independently of that context, but if we want to function socially and exhaustively validating every claim being made is not an option we have to take shortcuts, and I think we do it all the time even without realizing it. (I’m not writing that lightly since it seems similar to profiling which is ethically icky, but it is my conclusion upon introspection.)
In these cases we can at least imagine possible motive mismatch (known unknowns); in case of a GPT3-like thing instead of a motive you get a scary abyss or much more obscured motives of its human creators. I can’t imagine it having no impact on how I participate in an exchange.
[0] Even still, you can see how elsewhere in the thread there are warranted accusations of the motive being tainted by human exceptionality bias.
But extremely serious scientists, very smart people, are still drawing epicycles on blackboards studying “consciousness”.
Studying, even measuring the capabilities of an animal is science.
Justifying a soul is the purview of spirituality, not science. (Nothing against spirituality, I have a spiritual life, I just don’t confuse it with science).
Yeah, it's hard to quantify and isolate and experiment on, but that just speaks to either current limitations of human science, or possibly to limitations that cannot be surpassed. Given how much mileage certain philosophical movements have gotten out the common intuition that emerged during the Elightenment that everything is scientifically tractable, I understand the resistance to accepting these limitations and opening the door to all of the philosophical consequences of that intuition failing. But sorry, reality doesn't care about your philosophical attachments.
You really think that we’re a special case, that a difference in degree has become a difference in kind?
I personally experience a feeling that I’m conscious subjectively, but I have no evidence that I’m any more or less motivated by pleasure or pain or community than a dolphin is.
Where do we draw the line? What’s the acid test for “yup now we’re dealing with consciousness”?
Descartes was a genius, but he was no Alan Turing, and Alan-fucking-Turing got it wrong on the most famous thing named after him (among the lay population at least). The Turing Test was a great idea, but it’s now trivially useless.
Humans are special to (mostly) themselves and (substantially) other humans.
They are not special to the universe. We’ve had this argument, it was called the “Inquisition” at least once, and we eventually cleared up once and for all what celestial body rotates around the bigger one.
We do on the other hand know for a fact it's possible to run an instance of consciousness in a volume of about a liter that consumes like 20 watts (aka your average human brain), so there's something probably wrong with our general approach to the matter. GPT-3 already uses about twice as many parameters as our organic counterparts do, with much worse results. And it even doesn't have to process a ridiculously large stream of sensor data and run an entire body of muscle actuators at the same time.
This isn't accurate. GPT-3 has 175B parameters. The human brain has ~175B cells (neurons, glia, etc.) The analog to GPT-3's parameter count would be synapses, not neurons, where even conservative estimates put the human brain at several orders of magnitude larger. It's likely that >90% of the 175B could be pruned with little change in performance. That changes the synapse ratios since we know the brain is quite a bit sparser. In addition, the training dataset is likely broader than the majority of Internet users. Basically, its not an apples-to-apples comparison.
That said, I agree that simply scaling model and data is the naive approach.
If you can get GPT-like performance out of a 17B model, you should publish that.
Retrieval models (again, lots of published examples: RETRO, etc.) that externalize their data will bring the sizes down by about that order as well.
And the real problem is the way humans decide sentience in the first place - by how much the machinery acts as if it were sentient. There is no other information - just our perception. If we imagine two agents, A and B, one sentient and one insentient, but both acting identically, we theoretically couldn't decide which one is which.
So whether or not there is sentience in a machine then becomes a question as unanswerable as whether there is something that exists outside the universe we can perceive. We cannot know what we cannot possibly perceive.
It’s so unpopular that literally every post on HN/proggit related to language models has the same comment ;)
I used to do this for a living.
But… I take your meaning. I think I was subconsciously channeling OPT because that’s what the thread is about.
I agree it’s the kind of thing a fine-tuned HN chat bot would say all the friggin time.
But are we doing any different?
2. It was far from obvious a priori that this could work at all.
The existence of the sentiment ‘lol you trained a big model on X so obviously it produces a good probability distribution on X’ only exists because big models proved to be extraordinarily more effective than anyone expected.
It’s not a certainty that you’re overfitting but the burden is to show otherwise.
Besides, if your corpus is asymptomatically everything, why wouldn’t fitting it perfectly fitting it be the goal?
It would probably be more accurate to say that the bias/variance tradeoff loses meaning as the training set goes to infinity.
Grapevine is that the big actors are going to 5x their flops at any price in the next three years.
What does validation loss even mean at that scale?
That's called Google, and it serves a different purpose. Perfectly learning the distributional properties of text instead is not overfitting, that's just fitting.
No, the "popularity" of that sentiment exists because of their "effectiveness". That sentiment was existing and being voiced 10 years ago.
Otherwise brilliant and rational people just go all mystical when we start talking about meaningless words like “intelligence” or “consciousness”.
An animal or a human or a big matrix perform X well on Y task. That’s quantitative and objective.
All this “is it smart” bullshit is thinly-veiled “how does the world still rotate around my subjective experience given this thing writes better Tolkien fan fiction than I do”.
Performance on tasks. Everyone else shuffle over to the Philosophy department.
As you start to get to the size and comprehensiveness in the corpus that’s going on now, novel model outputs approach being something that might not have been in the training set, but likely will be in the future.
Copilot spitting out a function out of its training data with changed variable names to match those in your file(s) - and no one is actually testing what proportions of results those are - is still regurgitation.
The N in the NLP training set and arch implies a fuzzy match.
Why anyone would rather start with a plausible but broken-in-a-subtle way buffer rather than a blank one is beyond me.
The idea that models could possibly overfit the training data is hardly a new idea. It's standard practice to test for that. Check section 7 of the PaLM paper for example. https://arxiv.org/abs/2204.02311
Being Mr Who Ever Heard of This Guy myself, I took it as a compliment. :)
@veedrac on the other hand, who is smart, pretty clearly got the worse on that friendly little skirmish.
(Sorry, can’t resist.)