This level of incomprehensibility is worrisome.
This level of incomprehensibility is worrisome.
For now we have some degree of accountability, models need to be explainable for some small subset of the possible applications. But the overall mechanism seems to be that if there is a utilitarian perceived benefit that the solution is greenlit, without knowing whether this is just a statistical fluke which will steer us to longer term unintended consequences or whether it is a structural improvement. And once the decision is made to accept the model the ability to measure a baseline is destroyed so any 'creep' will likely go undetected. Such longer term consequences could be really damaging.
Whereas with regards to tech, society is currently at the level of "computer says foo" (so it must be true). Requiring explanatory output to justify decisions is a likely path forward for making the use of LLM's accountable. But given how little we've been able to make traditional human-guided tech companies accountable for things like sharing scoring formulas and abusing personal information, I'm not hopeful.
In essence, dealing with plausible but potentially misleading justifications is something that we have had to do forever, and will still have to do for future artificial agents as well.
When using ChatGPT, I can't help but thinking "this sounds like a high school essay". But really, it's not. Rather, it's someone who has college+ level of reading and studying, but has never had to have their ideas tested or scrutinized. A user of ChatGPT kind of mitigates this by asking for clarification, which is a form of learning in the specific session. But in the real world, ChatGPT-as-student would be adjusted with that feedback and then incorporate it into its overall model for serving the next user.
[0] For example, the SSC post about "predictive processing" resonated strongly with me
An LLM can sometimes be prompted to give a response that looks like an explanation for how it came up with a previous response, but it is important to realize that the actual process by which it came to make the original response bears no relationship to the “explanation” it gave. Because everyone's experience with language has been with human-generated language, which is often not completely rational but, when it is not, often deviates in ways we intuitively understand, it is difficult to see the responses of current LLMs as being fundamentally different, even when we know they are.
The fact the brain reaches decisions before we are consciously aware of them does not mean the brain does not tag decisioms with motives and other information necessary to explain them, though obviously we have to trust this information, and it may sometimes be wrong, but it has a pretty good track record and there’s no equivalent facility for AI today.
This holds especially true for the kinds of actions taken by AI today which are usually deliberated thought processes for humans.
This isn’t the same thing is saying all actions are explainable, but it is quite different to being a black box. In general, one only needs to question the explanations of a person if there is a good objective reason to do so.
> It is well known with experimental evidence that when people justify their decisions with the reasoning leading to it, it does not necessarily have anything to do with the actual reasons for the decision
Out of interest, what’s the highest quality evidence you are aware of in this area?
- fundamentally limited
- rough edges sanded off in the ball tumble mill of evolution
I've only done a small amount of statistical work professionally, but I remember that it was basically a nightmare to try and collect the right data and use the right variables to get a useful inference. Other engineers on my team would try to improve the models by throwing more variables at them or increasing the number of parameters in the model and I was always pushing back on this. I felt a bit crazy pushing back on this--surely, using more variables and more parameters will make the model fit historical data more closely, and a model which fits historical data more closely will fit future data more closely, and a model which fits future data more closely will be more useful?
Explaining why this is wrong is really hard. I have deep respect for people who are able to explain the principles of statistics and statistical modeling.
> Explaining why this is wrong is really hard.
Seems straightforward. Points 1 and 3 are correct; point 2 is false.
Statistics is not intuitive.
1. No cross-validation, and
2. You have the ability to solve for additional parameters without losing accuracy.
Statistics is not intuitive.
They haven't demonstrated that, it remains a conjecture. They show some evidence that they claim supports their conjecture. This evidence comes from "probing" (the classifier-based method you describe) which is a heuristic method that lacks the power to demonstrate anything conclusively. This is as the authors point out:
It’s believed that the higher accuracy these classifiers can get, the better the activations have learned about these real-world properties, i.e., the existence of these concepts in the model.
"It's believed" as in, we don't know, therefore we can't use to demonstrate.
I think this assertion is incorrect. Since the classifiers they used are very simple, they constrain the form of the model quite strongly; it must represent different tiles by different directions in the vector space of the internal state, otherwise the classifiers wouldn't be able to work. The representation for the whole board is then a sum of representations for each individual tile.
The experiments with modifying individual tiles in the model also show at a high level how they influence the predicted move. The researchers could've also looked at how that is implemented at the level of individual weights. It's not in this write-up and maybe they didn't even try, but that doesn't mean they have literally no idea.
The worrisome part is that it's performance isn't perfect, so there are bugs, and you might actually need to know all the low-level details to identify and fix the bug.
(I'm thinking more chess than othello, since I'm much better at the former, but I assume the thought process is similar)
Wars and disasters are usually pursued by minority groups in charge of means of violence, and manipulation is a way for them to trap you into letting a crack open long enough for them to invade it, but if we cant resist an AI, they deserve the Earth more than us :)
Now if a new species of humans suddenly emerged with a brain not only has capable as our but faster, more adaptable and extensible, it would disrupt the world.
(In terms of tangible effect on the world today and what seems on track for the few decades to come, AI is still far behind coal - to pick one - when it comes to concrete negative externalities tho)
Suppose (by way of analogy; this is not meant to be directly addressing AI) that in the next year some ingenious physicists finally get quantum gravity figured out, and that soon after the papers are published someone notices that the theory implies that one can use common household materials to make a device capable of destroying a planet.
(Seems unlikely, but I don't think anyone would have predicted that the last revolutionary theory of gravity we worked out would imply that one can use not-so-common materials to make a reasonably small device capable of destroying a city, but it turns out it does.)
It seems fairly likely that (1) the knowledge of how to do this could not be suppressed for ever, and that (2) if making the device turns out not to need all the fancy apparatus and hard-to-get materials that e.g. hydrogen bombs require, it wouldn't be that long before some idiot actually does it and destroys the earth.
And yet, these new discoveries would be "merely an aspect and consequence of the logical progression of that continuity", as you put it.
So, whatever the actual situation is with AI, I think it is demonstrably not the case that we can be confident AI won't somehow kill us all or destroy the world merely because AI is a thing we made and the world has had thousands of years to adapt to us.
Maybe everything will be fine, maybe not, but figuring out which will require looking in more detail rather than handwaving about how ecosystems always adapt.
Saying it’s only going to be as bad as the industrial revolution isn’t a pleasant prediction.
Why would we want more of this?
I mean, we've been struggling for milenia against our biases. I'd let some time pass until we can decide on what these ML models can get subtly wrong.
An LLM upends that paradigm, and worse, the closer it comes to appearing to reason and discuss, the more it implies that the human brain is nothing more than a biological computer, which is worrying for a lot of people.
If humans are nothing more than biological computers, do (or can) silicon computers have souls? If a collection of diodes can have a soul, what makes us special?
The brains of other people are just as incomprehensible.
"Curve Fitting" is the objective, the function encoded in the weights is the solution, and not actually well understood. See work from Anthropic[1] and Google[2] that explores this.
As an analogy consider applying the same argument to the AlphaGo value function. It's "just" fitting a bunch of curves to the statistics of millions of self-played games. However, to effectively capture those statistics the network needed to develop a bunch of heuristics. Needless to say these heuristics are not understood (else we'd already know the principles needed to play at AlphaGo's level), and are not just exhaustive lists of statistical trends but more like strategies[3].
Recent work[4] strongly suggests that "grokking" (a striking but not unnatural[5] form of generalization) involves networks transitioning from memorized statistics/solutions to a general solution. The curve fitting perspective would totally miss all this for a comfortable but misleading story: "the objective is curve fitting so it's just interpolating data points".
[1] https://transformer-circuits.pub
[2] https://arxiv.org/abs/2212.07677
[3] https://www.pnas.org/doi/10.1073/pnas.2206625119
Depending on how the model is set up, we'd say 'set of basis functions', 'language', 'strategy'.