Learning to think critically about machine learning
news.mit.edu
news.mit.edu
And that is in itself a dangerous moral and ethical lapse.
To be honest it would be a morally and ethical less dangerous world if we could get our feet back on the ground in relation to digital technologies in general.
> fundamental limits that no one talks about seriously.
I am starting to touch and stumble into the invisible cultural walls that I think make people "afraid" to talk about limitations. I am not yet done analysing that, but suspect it has something to do with the maxim that people are reluctant to question things on which their salary depends. That seems to be a difference between "scientists" and "hackers" in some way.
Going back to Hal Abelson's philosophy, "magic" is a legitimate mechanism in coding, because we suppose that something is possible, and by an inductive/deductive interplay (abduction) we create the conditions for the magic to be true.
The danger comes when that "trick" (which is really one of Faith) is mixed with ignorance and monomaniacal fervour, and so inflated to a general philosophy about technology.
I once worked on a team that spent a lot of time building models to optimize parts of the app for user behavior (trying to intentionally remain vague for anonymity reasons). Through an easy experiment I ran I ended up (accidentally) demonstrating that the majority of DS work was not adding more than minimal improvements, and so little monetary value and it did not justify any of the time spend on this.
I was let go not long after this, despite having help lead the team to record revenues by using a simple model (which ultimately was what proved the futility of much of the work the team did).
Just a word of caution as you
> start to touch and stumble into the invisible cultural walls that I think make people "afraid" to talk about limitations
Competences work at multiple levels, visible and invisible. Being good at your job. Showing you're good at your job. Believing in your job. Getting other people to believe in your job. Getting other people to believe that you believe in your job... and so on ad absurdum. Once one part of that slips the whole game can unravel fast.
I find this hand a little over played.
It depends on the degree of fidelity we demand of the answer and how deep we want to go questioning the layers of answers. However, if one is happy with a LOL CATS fidelity, which suffices in many cases, we do have a good enough understanding of SGD -- change the parameters slightly in the direction that makes the system work a little bit better, rinse and repeat.
No one would be astonished that using such a system leads to better parameter settings than ones starting point, or at least not significantly worse.
Its only when we ask more questions, ask deeper questions that we get to "we do not understand why SGD works so astonishingly well"
The question then becomes: why does this generalize [4], given that the classical theory of Vapnik and others [5] becomes vacuous, no longer guaranteeing lack of over-fitting?
This is less well understood, although there is recent theoretical work here too.
[1] Lee et al (2019). Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent. https://proceedings.neurips.cc/paper/2019/hash/0d1a9651497a3...
[2] Allen-Zhu et al (2019). A convergence theory for deep learning via over-parameterization. https://proceedings.mlr.press/v97/allen-zhu19a.html
[3] Du et al (2019). Gradient Descent Finds Global Minima of Deep Neural Networks. http://proceedings.mlr.press/v97/du19c.html
[4] Zhang et al (2016). Understanding deep learning requires rethinking generalization.
[5] Vapnik (1999). The nature of statistical learning theory. https://arxiv.org/abs/1611.03530
BTW your randomized algorithm with a minor tweak is surprisingly (unbelievably) effective -- randomize the weights of the hidden layers, do a gradient descent on just the final layer. Note the loss is even convex in the last layer weights if matching/canonical activation function is used. In fact you dont even have to try different random choices, but of course that would help. The random kitchen sink line of results are a more recent heir to this line of work.
I suspect that you already know this and the fact that the noise in SGD does indeed regularize and the way it does so for convex function has been well understood since the 70s, so I am leaving this tidbit for others who are new to this area.
I think it’d have to be related to the huge number of dimensions it works on. But I have no idea how I’d even begin to prove that.
How much of that magic is smoke and mirrors? For example, the First Tech Challenge (from FIRST Robotics) used Tensor Flow to train a library to detect the difference between a white sphere vs a golden cube using a mobile phone's on-board camera.
The first time I saw it, it did seem pretty magical. Then in testing realized it was basically a glorified color sensor.
I think these things make for great and astonishing demos but don't hold up to their promise. Happy to hear real-world examples that I can look into though.
As for a formal analysis, I just can’t imagine there existing a formal analysis of ML that can describe the distinctly qualitative aspects of it. It’s like coming up with physics equations to explain art.
The vision model was tolerably decent at tracking incremental updates to object positioning, but for some reason would take 2+ seconds to notice that a valid object was now in view (which is quite a lot, in the context of a 30s autonomous period), and frequently identified the back walls of the game field as giant cubes.
I mean this is an extremely difficult thing to disentangle in the first place. It is very common for people in one breath to recite that correlation does not equate to causation and then in the next breath propose causation. Cliches are cliches because people keep making the error. People really need to understand that developing causal graphs is really difficult, and that there's almost always more than one causal factor (a big sticking point for politics and the politicization of science, to me, is that people think there are one and only one causal factor).
Developing causal models is fucking hard. But there is work in that area in ML. It just isn't as "sexy" because they aren't as good. The barrier to entry is A LOT higher than other type of learning, so this prevents a lot of people from pursuing this area. But still, it is an necessary condition if we're ever going to develop AGI. It's probably better to judge how close we are to AGI with causal learning than it is for something like Dall-E. But most people aren't aware of this because they aren't in the weeds.
I should also mention that causal learning doesn't necessitate that we can understand the causal relationships within our model, just the data. So our model wouldn't be interpretable although it could interpret the data and form causal DAGs.
That is literally the point of the thought experiment: https://en.wikipedia.org/wiki/Wave_function_collapse
It isn't just our models that can't explain it, there are real physical limits which mean that _no_ model can predict what state the cat is in.
The only reason why cats are a more outrageous example than electrons is that we see cats behave classically all the time.
The only vaguely plausible explanation why cat states are impossible in general is that large quantum system become spontaneously self decoherent at large enough numbers of particles.
Perhaps you meant to say "...state the cat will be in when observed"?
Otherwise, an important nitpick applies: superposition means that the system is not in any single state, so there's nothing to "predict" - it's a superposition of all possible states.
Prediction comes in when one asks what state will be observed when a measurement is made. As far as we know, that can only be answered probabilistically. So no model can specifically predict the outcome of a measurement, when multiple outcomes are possible.
You can't be liable for anything, you were just doing what the computer told you to do, and computers aren't fallible like people are.
2015 called, they want you back! Now seriously, "just" does an amazing amount of work for you. How do you "just" make logistic regression write articles on politics, convert queries in SQL statements? or draw a daikon radish in a tutu?
Humans are "just" chemistry and electricity, and the whole universe just a few types of forces and particles. But that doesn't explain our complexity at all.
In other words, we too can't do three digit multiplication in our heads reliably, but can do it much better on paper, step by step. The problem you were mentioning is caused by the bad approach - LMs need intermediate reasoning steps to get from problem to solution, like us. We just need to ask them to produce the whole reasoning chain.
- Chain of Thought Prompting Elicits Reasoning in Large Language Models https://arxiv.org/abs/2201.11903
- Deep Learning for Symbolic Mathematics https://arxiv.org/abs/1912.01412
Whatever makes you think it’s necessary for AGI, when we don’t have it?
Even if it had 80% accuracy (optimistic) it would still he too mediocre to be used at any serious scale.
Most of the stupid crap that people give about degenerate cases where deep learning doesn't work (cartpoll in reinforcement learning, sine/infinite unbounded functions) are showcasing how bad gradient based training is - not how bad deep learning is at solving these problems. I can within seconds solve cartpoll with neural networks using neuroevolution of weights....
Because, with RELU activation, I’m fairly confident that the latter, at least, is possible.
(Where inputs are given using digits (where each digit could be represented with one floating point input), and the output is also represented with digits)
Like, you can implement a lookup table with neural net architecture. That’s not an issue.
And composing a lookup table with itself a number of times lets one do addition, etc.
... ok, I suppose for multiplication you would have to like, use more working space than what would effectively be a convolution, and one might complain that this extra structure of the network is “what is really doing the work”, but, I don’t think it is more complicated than the existing NN architectures?
If the weights are designed, and the network architecture allows something to hold the information needed, then there is really no obstacle to having it get multiplication entirely (not just 90%).
Now, would that be learnable? I’m not so sure, at least with the architecture one would use if designing the weights.
But,
I see no reason a transformer model couldn’t be trained on multiplication-with-work-shown and produce text fitting all of those patterns, and successfully perform multiplication for many digits that way.
And, by “showing all work” I don’t necessarily mean “in a way a person would typically show their work”, but in a easier-for-machine way.
Neither can people, for the most part.
I have more expertise in deep learning than anyone else here and the delusions of the incoming transformer winter will be painful to watch. In the meantime, enjoy your echo chamber.
I... I guess it's possible?
No, you don't. Looking at your experience, there is simply no way that you are the foremost expert in DL on HN.
What I know of your experience shows a low number of years of experience, a lack of papers, and a lack of true hands-on experience at the small number of companies in the world that have the resources to truly investigate large models. How can you know so much about LLMs without ever having the resource to train one?
I'm obviously not going to dox you, so you can easily just dismiss what I'm saying. But even just reading through your HN comments shows arrogance in your own knowledge (across multiple domains).
A specifically memorable quote is:
> I frequently create unique on the internet [words]
This is very true. Your erudition is apparently only matched by the uniqueness of the words you use when on the internet.
Meaning?
They're not magic - nothing is, but what are they?
> but it also have fundamental limits that no one talks about seriously
What are these fundamental limits? 20 years ago I imagine skeptics in your camp would have set these "fundamental" limits at lower than DALL-E 2, GPT-3, AlphaStar etc. Or are you talking about limits today? In which case, sure, but I think "fundamental" is the wrong word to use there given they change continuously.
> It's mathematically related to all prior signal processing techniques (mostly a proper superset)
And human brains are what if not signal processing machines?
Emergent magic.
Just because something is difficult to analyze doesn't mean it has limitless power.
Agreed. Is it an original lapse, or derivative though? When researchers/engineers oversell their story to get the funding they wouldn't otherwise get, where is the collapse? With the engineers/researchers? Or with the forces that built a system where that was the only way forward for them?
When a hungry thief steals to eat, is the thief morally bankrupt? Or is those that engineered the shortage?
Suddenly it all becomes a lot more palatable that many don't know how it works.
More importantly, it can be a dangerous business lapse.
https://ocw.mit.edu/courses/res-tll-008-social-and-ethical-r...
More generally, what does it mean for a model to be "fair"?
LIT Company’s Definition of Fairness (Group Unaware): The company believes that a fair process and, therefore, a fair model, would not account for gender or race at all.
Advocacy Group's Definition (Demographic parity): An advocacy group believes that a model is fair if the distribution of outcomes for each demographic, gender, or other subgroup is the same among those that applied and those that were accepted. For example, in the example above, 30% of the applicants for loan applications come from women. In the demographic parity definition of fairness, this means 30% of the approved loan applications should come from women.
I feel like the course content is somewhat slanted here. It is missing the definition of "fairness" in which you treat race and gender just like any other feature. Many systems work this way in practice - for example car insurance charges you differently by gender, because the statistics for genders are different. Ad-matching by gender and race is a longstanding practice. And any new system that you just train from scratch, by default it will not know to treat gender or race different from anything else.
It is an interesting question, though. The main problems, I think, are practical ones - large enough AI models cannot be "race-blind" because if you remove race as a feature, they will be able to infer it anyways from proxy features. Whereas the only real way to enforce a system achieves the same percentage results for different groups is to add a "quota system" where you explicitly use different thresholds for different groups. So the practical alternatives often become "quota" or "nothing".
So, yeah, for the Diversity Committee definition of fairness, you need quotas.
The real issue here, of course, is whether they are just like every other feature:
Certainly in our society they are not perceived that way. People perceive very serious issues and have very strong feelings around race and gender. We see that right here on HN, of course.
There is also, of course, a lot of discrimination by humans based on race and gender. If we want an unbiased, fair (and accurate) system, we have to correct for that. And the discrimination creates higher order effects: If there is discrimination against group X in K-12 education funding, then fewer of X will go to college, and fewer will have higher-paying jobs. If we then select blindly for income, we incorporate that bias (which might be appropriate if studying income by group, but not if we use it as a proxy for intelligence or effort).
> the practical alternatives often become "quota" or "nothing".
Those aren't practical alternatives, they are logical extremes creating a Manichean choice. Those are alternatives or a political debate, not for practical problem-solving.
"Fairness" is a technical term-of-art meaning "the outcome _should not_ be effected by inputs X, Y and Z", and the collected science around making a system behave that way. It closely but not quite matches the colloquial meaning, ie "that's not fair!" - kinda like how "a fair coin" has a specific meaning that mostly tracks with how people use the word, but not quite, and with a lot more specificity.
It's typically applied when you want to correct for real-world "unfair biases" in the training data; which in practical application is typically race, gender and the other legally protected categories - but AFAIK is just whichever inputs you decide you want to not have an impact on the outcome.
AFAIK what you get out of the AI/ML "fairness" science is a way to measure, and correct for, dependency on the inputs that you (exterior to the system) have decided that you want to _not_ impact the outcome.
Interestingly, car insurers in the EU are no longer allowed to differentiate premiums by gender. But since this law came in, the difference between average male and female premiums has widened, based on correlated risk factors, such as occupation and model of car. https://www.theguardian.com/money/blog/2017/jan/14/eu-gender...
> large enough AI models cannot be "race-blind" because if you remove race as a feature, they will be able to infer it anyways from proxy features
In theory using a gradient reversal layer and an adversarial classifier you could do just that, to an extent. It could hurt the model's performance, which is exactly your point (should we ignore features with signal because they can be used to discriminate.)
> “It is not someone else’s job to figure out the why or what happens when things go wrong. It is all of our responsibility and we can all be equipped to do it. Let’s get used to that. Let’s build up that muscle of being able to pause and ask those tough questions, even if we can’t identify a single answer at the end of a problem set,” Kaiser says.
Hopefully the actual course content is more concrete than this. But this kind of language strikes me as encouraging people to feel the "correct" way about a problem, but not really emphasizing coming up with concrete, actionable solutions.
And without actionable solutions, I feel like the value of this content would be very limited.
While I don't think a majority really thinks current systems are conscious, SOTA results are absolutely astounding (check out DALL-E 2 if you haven't seen it already). Whether or not an agent is conscious doesn't really matter from a practical standpoint (but obviously a moral one) in the long run - it is intelligence that matters with these agents, and they're getting absurdly more intelligent by the half-decade
Human consciousness isn’t just a brain. It’s a system of which the brain is a part, occurring through time.
The idea consciousness emerge proportionately with the accuracy of your mental isomorphic world representation is cute however we don't become more conscious by becoming more erudite, and the most intense magical qualias, such as e.g orgasms are accessible to the simplest mammals and are unrelated to activities in the higher cognitive regions of the brain. Even a newborn that has no understanding of its surrounding experience qualias.
Very large languages model hide to the layman that they are the gigantesque failure in NLP ever by showing they improve the state of the art in zero or few shot learning. Who cares this is so cringe. Full size learning is what matter the most and even full size learning do not yield satisfying accuracy on most Nlp tasks (but close enough) Therefore the only use of PALM is to have mediocre (70-80%) accuracy which is better than previous SOTA, only for tasks that have no good quality existing datasets. And 530 billion is close to the max we can realistically achieve, it already cost ~10 millions in hardware and underperform a 300 million model in full size learning (e.g dependency parsing, word sense disambiguation, coreference resolution, NER, etc)
It's crazy people don't realize this gigantic failure but as always it's because they don't care enough
Most ML is stats and phase-change to find parameters to work on. Optimisation is not a black art, its Operations-Research by another name.
I blame the rise of Kurzeweil. Magical realism was a great literary trend, but applying it to machine learning and consciousness and the uplift is not helpful.
And instead of understanding what's going on, and solving the problem efficiently, instead it's "throw GB's or TB's of data at an algo that we tweak and hope for the best".
Sure, it gets results for massive processing times and data. And sure it's "fuzzy" and can fail on really stupid stuff and provide bad answers confidently. But it'll get the next VC funding line, won't it?
- Classification in computer vision
- Protein folding (AlphaFold)
- Image generation (Dall-E 2)
- Answering general language queries (GPT-3)
It's unclear how any of these applications could have been tackled _without_ machine learning. We had no solid grasp on these problems before ML, with the exception of protein folding. I think you are being cynical. Of course not every ML project will result in a home run success. Should that make us sceptical of the entire field?
In the end, ML is not that different to what we are doing with all models, ie use a set of data to create a crude model of the process that we are trying to learn/figure out.
ML doesn't have deduction though, it just has induction (and that's usually the training to make the model, not the model itself).
However, anyone who thinks this tech couldn't go off the rails in the hands of nefarious actors should go read Ed Black's "IBM and the Holocaust" or Josef Teboho Ansorge's "Identify and Sort".