"godlike"? Really? I'm not religious, but this seems like an overreaction for something that has no agency.
"godlike"? Really? I'm not religious, but this seems like an overreaction for something that has no agency.
Considering that we don’t know how the brain works so well, and we don’t understand why LLMs work so well, simply on the basis of their output I think the safest assumption is that these models do indeed have agency, or at least the capability of agency.
> How do you know agency is not simply the output of a large language model encoded in neurons?
I'm not sure what you mean here. Is agency an emergent effect of large digital or biological neural network? Maybe! Is it an emergent effect of a large language model? If it is, then it should be clear, or demonstrable, that the model (1) has goals (2) takes concrete steps to achieve those goals.
> What is the difference between neuronal and digital weights?
Brain chemistry works at orders of magnitude less speed, since we're talking about periodically building and releasing an ionic differential between the inside and outside of a cell wall. Moreover, we have a massive number of neurons and a stupidly massive amount of interneuronal connections, with billions of years of training over billions of lineages. Digital weights, in contrast, are a stripped down model of this system that throws out a whole class of complexities like hormones and metabolism.
> I think the safest assumption is that these models do indeed have agency, or at least the capability of agency.
I think this is an overly generous assumption.
> it should be clear, or demonstrable, that the model (1) has goals (2) takes concrete steps to achieve those goals.
That definition seems arbitrary to me; many humans wouldn’t pass this test. On the other hand, LLMs certainly seem capable of acting towards specific goals (such as helpfulness). So, I would say that based on your definition, LLMs have agency. But I think you really meant, internally generated goals. Time will tell.
That said, humans who don’t have clear goals can and are coerced into all sorts of damaging behaviours by those who do have goals. So even if I accept that LLMs don’t have their own goals, they can certainly be manipulated to act in favour of the goals of others. That’s effectively what prompt engineering is all about.
So I just think it’s a mistake to make assumptions about these LLMs. We don’t know why they work so well, and it will take a very long time until we do.
In the meantime, let’s not make assumptions that we can’t justify.
What part of this is arbitrary, i.e. random, whimsical, or biased? This is a fairly comprehensive working definition of "agency". "Internally generated", which is implied, is a nice touch.
Nearly every human over the age of 18 months passes this test with flying colors. Toddlers have goals, and do everything in their power to achieve them. What humans are you thinking of that don't have agency?
> In the meantime, let’s not make assumptions that we can’t justify.
I totally agree. I think that until demonstrated otherwise, I will assume that LLMs are a giant statistical sieves that (1) periodically spit out text directly from their training set, unmodified, and that (2) do not learn on their own, do not formulate their own goals, and do not take actions to achieve those goals.
What could go wrong?
Great question! These are predictive models that accept a text query, do some matrix math, and then return some text. At what point in that server-client relationship does this algorithm jump the rails and run amok?
> A popular nightmare scenario for AI is giving it access to tools, so it can make API calls and execute its own code and generally break free of the constraints of its initial environment. Let's do that now!
I also think that you're assuming we know a lot more about how these things work than we actually do; you seem to think nobody is going to hook these up to APIs that can actually modify the world, despite the barrier to doing so being incredibly low; and you don't seem to have read about the adversarial training that people have been doing between the LLMs.
It's obvious that you think everything is all safe and nothing will go wrong, and I really hope you're right. But I think it's a very dangerous assumption.
Hope for the best, plan for the worst.
> Considering that we don’t know how the brain works so well, and we don’t understand why LLMs work so well, simply on the basis of their output I think the safest assumption is that these models do indeed have agency, or at least the capability of agency.
These systems can already interact with others, it’s not moving the goalposts, it’s common knowledge. Anyone with access to the APIs can make it happen. Or are you now claiming we are just talking about one specific LLM and not LLMs generally?
Anyway this discussion is fruitless, I’m out.
I don't disagree that digital systems with neural architectures could have agency in principle, but agency generally is definitely not the output of a large language model. Animals without language have agency, in that that they take actions to fulfill their desires. Current LLMs may have some degree of intelligence, but they don't even appear to have any consistent wishes or desires. You can get them to talk longingly about x... until you give another prompt and suddenly x doesn't matter to them at all.