Those mechanisms only explain next word prediction, not LLM reasoning.
That's an emergent property that no person, as far as I understand it, can explain past hand waving.
Happy to be corrected here.
"hey it's got an irrational preference for naming its variables after famous viking warriors, lets change that!"
But worse, it's not that you can't change it, you just don't know! All you can do is test it and guess its biases.
Is it racist, is it homophobic, is it misogynistic? There was an article here the other day about AI in recruitment and the hidden biases. And there was a recruitment AI that only picked men for a role. The job spec was entirely gender neutral. And they hadn't noticed until a researcher looked at it.
It's a black box. So if it does something incorrectly, all they can do is retrain and hope.
Again, this is my present understanding of how it all works right now.
But overall in my opinion if devs are able to rebuild it from scratch with a predefined outcome, and even know how to improve the system to improve certain aspects of it, we do understand how it works.
Regarding the car, if you know how to build a car, you understand how a car works. A driver is more like someone using and llm, not a developer able to create an llm.
Sorry to inform everybody doing their Ph.D. on LLM interpretability that they’re just wasting their time.
Yes! loads! (: I want to be able to say statements like "this model will never ask the user to kill themselves" and be confident, but I can't do that today, and we don't know how. Note that we do know how to prove similar statements for regular software.
Common misconception, MoEs do have different "experts", but the model learns when to send input to different experts, and the model does not cleanly send coding tasks to the coding agent, physics tasks to the physics agent, etc. It's quite messy, and not nearly as intepretable as we'd want it to be.
In a non-linear system the former is often easier than the latter. For example we know how planets “work” from the laws of motion. But planetary orbits involving > 2 bodies are non-linear, and predicting their motion far into the future is surprisingly difficult.
Neural networks are the same. They’re actually quite simple, it’s all undergraduate maths and statistics. But because they’re non-linear systems, predicting their behaviour is practically impossible.
The study of LLMs is much closer to biology than engineering.
I don't want to paste in the whole giant thing, but if you're curious: [0]
[0] https://drive.google.com/file/d/1D5yICywmkp24YajboKHdYFcBej0...
This article describes how Belgian supermarkets are replacing music played in stores by AI music to save costs, but you can easily imagine that the ai could also generate music to play to the emotions of customers to maybe influence their buying behavior: https://www.nu.nl/economie/6372535/veel-belgische-supermarkt...
What is it that we don’t understand?
The ML field has a good understanding of the algorithms that produce these floating point numbers and lots of techniques that seem to produce “better” numbers in experiments. However, there is little to no understanding of what the numbers represent or how they do the things they do.
And inspecting each part is not enough to understand how, together, they achieve what they achieve. We would need to understand the entire system in a much more abstract way, and currently we have nothing more than ideas of how it _might_ work.
Normally, with software, we do not have this problem, as we start on the abstract level with a fully understood design and construct the concrete parts thereafter. Obviously we have a much better understanding of how the entire system of concrete parts works together to perform some complex task.
With AI, we took the other way: concrete parts were assembled with vague ideas on the abstract level of how they might do some cool stuff when put together. From there it was basically trial-and-error, iteration to the current state, but always with nothing more than vague ideas of how all of the parts work together on the abstract level. And even if we just stopped the development now and tried to gain a full, thorough understanding of the abstract level of a current LLM, we would fail, as they already reached a complexity that no human can understand anymore, even when devoting their entire lifetime to it.
However, while this is a clear difference to most other software (though one has to get careful when it comes to the biggest projects like Chromium, Windows, Linux, ... since even though these were constructed abstract-first, they have been in development for such a long time and have gained so many moving parts in the meantime that someone trying to understand them fully on the abstract level will probably start to face the difficulty of limited lifetime as well), it is not an uncommon thing per se: we also do not "really" understand how economy works, how money works, how capitalism works. Very much like with LLMs, humanity has somehow developed these systems through interaction of billions of humans over a long time, there was never an architect designing them on an abstract level from scratch, and they have shown emergent capabilities and behaviors that we don't fully understand. Still, we obviously try to use them to our advantage every day, and nobody would say that modern economies are useless or should be abandoned because they're not fully understood.
This example from software doesn't meaningfully hold for neural networks. It's a bit like trying to watch an individual COVID virus duplicate and then attempting to predict the pandemic. It's incredibly complicated and we haven't yet built the tools to help us understand
Similarly, two adult humans know what to do to start the process that makes another human, and we know a few of the very low-level details about what happens, but that is a far cry from knowing how adult humans do what they do.
[1]: https://www.reddit.com/r/slatestarcodex/comments/1o6n5ne/why...
That is what the parent meant.
AI sits at a weird place where it can't be analyzed as software, and it can't be managed as a person.
My current mental model is that AGI can only be achieved when a machine experiences pleasure, pain, and "bodily functions". Otherwise there's no way to manage it.