Self-supervised learning: The dark matter of intelligence
ai.facebook.com
ai.facebook.com
This somewhat humorous but also worrying paper comes to mind: https://youtu.be/Lu56xVlZ40M
I think we really need to consider how we can mitigate the risk of raising completely ununderstandable yet scarily capable machines.
I do agree that more parameters and a bigger training set makes the models more opaque (or at least more expensive to understand), but the self supervision is not the reason.
Code, Data, Model, Use-Case, ...
For AI models, that isn't really the case. We don't know how each individual model will act in the real world, and certainly not after having spent millions of hours training against itself or in a simulated environment. And as these models get more and more powerful, this growing uncertainty is something that I believe is worrisome.
We don’t know how brain works, why sleep exist, most of humans cultures and languages are not documented, hundreds if not thousands of psychological and medical things are unknown, etc. Machines and their applications are a few more magnitudes more understood than any human matter, by the sheer fact we created them. Their complexity is ridiculously low in comparison to human, their biology, psychology, etc.
but all of that being said, i think it's also worth considering what a non-human entity with more power than us might be able to do. after all, even the worst examples of humanity are confined to at most "only" causing millions of deaths, but not the eradication of all life on Earth for example. it's not difficult to imagine that a non-human entity, with non-human goals, and with super-human powers, might act in a way that is contrary to our interests, or even to our existence.
AI is just the fruiting body of something that has already happened.
How many people have been in US prisons again? Oh right, about 3%. Clearly a high percentage of these people did something both unexpected that seriously damaged others.
So this is not a real difference between AIs and humans. One might even say, not a single AI has been convicted yet, so it might not be true for them. For humans, serious, bad, violent, illegal, even when totally irrational, ... behaviors are very common indeed.
(and frankly, if there ever is a true Human <> AI conflict, not allowing any kind of errors for AIs seems to me a strong contender for the casus belli)
The computer can do whatever it wants. People can do whatever they want. The question will be what level of security access will they have.
The key difference today is people are really good at making rationalizations for individual decisions. Computers are not.
Sometimes decisions are generally important, when they are important they require trust, and if these models never generate a method of demonstrating success/trust they won't be adopted.
And therein lies the rub. Do people actually want to know the truth? They invent methods to obtain it, but is that their aim or is it to confirm their current understanding of the world? Unlike a human being that can be forced into silence through coercion or manipulation, a computational model, once proven with certainty, is never going to go back in Pandora's box. This contradiction of pursuing truth but suppressing the inconvenience of its conclusions hasn't disappeared for some reason despite greater and greater knowledge of ourselves. Will it ever?
Several attempts at training with real world models have produced results that have been attacked for being "algorithmically biased". This isn't because there was anything experimentally wrong with the dataset or choice of model on their own, but because of the result's contradiction of a certain worldview. While, I agree with you that people come closer to the truth in the long-term, in the short- and medium-term, I expect certain people to pad their own data and quash the different findings out of ideological motivation for some time.
I'd consider The ability to output a rationalization in human understandable text (ie: english in my case) will determine success /failure.
How is it harder to understand a model that was trained to predict which two images crops come from the same source image (a la contrastive loss for example), than a model that was trained to predict which class an image belongs to?
How do you do that with humans? You watch their behavior, language and you test them.
If you make them believe that they choose another decision, they will create another explanation on the fly.
Slow thinking - such as reasoning - is more expensive so we only use it when we are faced with a new situation for which we have not 'cached' the answer. It entails the creation of a mental model and imagining outcomes. The exact mix of the two modes is tuned by evolution.
If all we had was the slow system we'd be much less efficient at survival.
Language models that are not given concepts of part-of-speech, dependency trees, grammars, develop representations of their own and people have found ways to inspect them.
https://arxiv.org/abs/1104.5466
Annotation doesn't scale. You are not going to get to real intelligence by labelling billions of data points.
I expect that computer vision for autonomous vehicles will become good enough for practical use quite soon after the AV companies start doing large scale self-supervised learning (I don't like that term). Maybe Tesla and Waymo and others are doing that already.
Another great application area is medical imaging. Progress in this area has been seriously hampered by the limitations of the annotation paradigm. Once people start to do SSL, there should be rapid progress.
Fine tuning labels black-box style is a terrifying concept to most who are working in fields where great risk must be managed to avoid unintended and disparate impact. Facebook making an oopsies suggesting my friends face in a photo instead of mine with a SSL trained model may seem trivial, but this non-human oversight is concerning for other applications.
If you were rejected by a bank for a home loan, you would want to know why your creditworthiness wasn’t evaluated positively (and banks must explain precisely why to regulators). And if a self-driving car made a decision in a collision, or a CV model performing cancer screening in X-rays that gave a bad result - ...or recommending politically divisive content... - or a thousand other real world examples that have human impact.
*Of course I am talking about tech based on the current DL approaches. If we had AGI, then the question does not even need to be asked.
Perhaps there's something in the black box nature of our own self understanding similar to these hyper complex function approximators.
Judges make harsher rulings when they're hungry. A regulator might decline a loan and justify it b/c they had a bad experience with someone that reminds them of the person they're dealing with. AI makes bad decisions in boundary conditions or when they're trained on biased data. I think it's good to strive for opacity in every case but in complex decision spaces I'm not sure if we'll ever get fully satisfactory explanations.
Which is a meandering way of making the non-point that models making judgments have made me re-evaluate what I trust and why, and I think in many cases I'd rather trust a black box if I know what's gone into it and what's come out.
Actually, that finding is highly improbable: https://nautil.us/blog/impossibly-hungry-judges
> If you were rejected by a bank for a home loan, you would want to know why your creditworthiness wasn’t evaluated positively (and banks must explain precisely why to regulators).
This is why we don't use humans to evaluate credits, but precise algorithms. Humans are just applying the algorithms, they are not evaluating themselves with their gut feeling. I don't see why this would change with AI. Regulations prevents unexplainable tools to be used there, so deep learning black box models will not be used, similarly to why humans are not used today.
But in cases where performance is required, but not explainability, then deep learning will strive. And I believe that cases where explanation is required is only a very small subset of areas where AI could be useful.
> And if a self-driving car made a decision in a collision
Would you prefer a car that crashes once every 100 million miles but is not explainable, or a car that crashes once every 100 miles but is interpretable can explain why it crashed ?
We're not talking about "Not Hotdog" here. Application of a model in a real-world at-scale scenario is a lot more than running inference and walking away.
At a bank credit decisions are evaluated by humans, frequently and often. These reviews are conducted in the forms of sampling audits, control processes, and other scenarios that involve internal bank employees and external regulators. In each case, humans will inspect the details of what occurred. This would be impossible with any type black-box model (SSL, deep NN, etc.).
Bank credits are auditable, because there are some precise algorithm around them. Humans created the algorithm specifically so that it is interpretable. Then they must apply it the same way to everybody. They can't just say "oh, I think this person is more likely to actually reimburse, I play golf with him and I trust him !"
This is a specific case where no black box will ever be used (at least I hope).
Of course, for cases where explainability is needed, either required by law (such as banking), or by common sense, then black boxes will not be deployed.
That's why I like things like this https://arxiv.org/abs/2102.12627 "How to represent part-whole hierarchies in a neural network" by Hinton.
But as a serious answer: of course, there is no major economic use of trying to make sure it has an internal subjective experience, but that’s not what people are aiming for when referring to AGI. The goal is that it be able to accomplish general tasks and goals. Like, what can you do with it? Everything you can do at all. That’s what people are aiming at. What tasks of reasoning and planning could a person do for you? An AGI would, basically by definition, be able to do those same kinds of things. (Of course, something could be an AGI but not as intelligent as a typical human, so long as it could still reason about all the same kinds of things. But at that point it would presumably just be a question of scaling things up.)
The fundamental problem is whether AGI is outer-directed or inner-directed - i.e. whether it sets its own goals, or whether you tell it what to do and it improvises a solution.
AGI is most useful when it's outer-directed with limited agency but some improvisational autonomy. You can give that kind of AGI specific problems and it will solve them in useful but unexpected ways. Then it will stop.
AGI is most dangerous when it's inner-directed with full independent agency. Not only will humans have no control over it, there's a good chance humans won't even understand what it's aiming for.
Agency is almost entirely unrelated to symbolic intelligence. You can have agency with very limited intelligence - most animals manage this - and no agency with very high symbolic intelligence.
This is not a Boolean. But there will be a cutoff beyond which inner-directed behaviour predominates, initially driven by programmed "curiosity", leading to unpredictable consequences.
I'm not quite sure what you mean by "sets its own goals" (for "inner-directed"). I assume you don't mean "modifies its own goals" (as, why would that help further its current goals?), but I'm not sure what it would mean. Maybe you mean like, if it acquires goals and preferences in the way that humans do, with shifting likes and dislikes that aren't consistent across time? Yes, that would certainly be quite dangerous (unless the AGI was via like, just emulating a human's mind, which might not be so dangerous, in which case only somewhat dangerous).
I believe I see your point in the safety features in having it receive a specific short term task, achieve the task, and then stop. Once it stops, it isn't doing things anymore, and therefore isn't causing more problems. But it seems like it might be difficult to define that precisely? Like, suppose it takes some actions for a period of time, but complicated and anticipated-by-it-but-not-by-us consequences of its actions continue substantially after it has stopped?
Unlike software, which can be mathematically airtight, there's no reason to assume that physical devices built by humans are unexploitable in the face of the massive intelligence of the AGI. So the AGI will just ask itself which is easier: delicately performing surgery on its own sensors, or building tons of nukes.
Honestly, that people are even worried about AGI doomsday shows to me the power of narrative. Everybody has heard of SkyNet and the sorcerer's apprentice, nobody has heard of the AGI who outsmarted itself by creating digital porn for its objective function. Therefore, through the miracle of the availability heuristic, AGI doomsday it is.
Here are some words that I read but did not grasp firmly: learn, data, task, train, intelligence, general, model, skill, label, language, understanding, reality, observation, predictive, objects, concepts, act, hypotheses, knowledge, 'common sense', 'dark matter', 'artificial intelligence', teaching, classify, supervision, autonomous, 'self-supervised learning', recognize, patterns, representations, processing, systems, pretrained, vision, real-world, helpful, promising, 'energy-based models', prediction, uncertainty, 'joint embedding methods', 'latent-variable architectures', reasoning, 'predictive learning', 'supervisory signals', signals, structure, unobserved, property, input, 'co-occuring modalities', 'unsupervised learning', feedback, reinfrocement, 'downstream tasks', meaning, syntactic, word, associate, probability, vocabulary, 'convolutional network', network, 'prediction uncertainty', 'predicting missing words', computing, softmax layer, probability distribution, energy, incompatible, computer vision, and so on.
Rather than list "ideas" or "concepts" I try to make a record of what appeared and the response I had to its appearance.
How are we to solve difficult problems if we stop mentioning them or their difficulty?
(Context https://en.m.wikipedia.org/wiki/Escape_room )
Wow, are these so-called ai scientists really that daft?
Sure on paper we drive 20-40 hours "practicing" then hit the roads, but we've been back-seat driving and driving via video games or TV since the day we're born.
While a 2-year old doesn't know the intricacies of driving my 3-year old can definitely yell hey dad the lights red slow down.
Full-immersive life experience. Perhaps ai's need to be put into a real-world "birth" simulation (perhaps we already are this experiment) to learn as they grow.
The problem I see is the narrowness, if you're training on a narrow subset then you're gonna get narrow results. I don't know the best way of doing it but you need to start thinking of an ai's "brain" like that of a child's and how it absorbs things - the human brain is remarkable sure, but I don't doubt it's duplicatable in silicon.