Deterministic systems can be chaotic, which implies unpredictability and that is anathema to control.
AI, in particular sentient AI, is right on the border of chaos. Meaning, it can be arbitrarily unpredictable.
Arbitrarily uncontrollable, that is.
Deterministic systems can be chaotic, which implies unpredictability and that is anathema to control.
AI, in particular sentient AI, is right on the border of chaos. Meaning, it can be arbitrarily unpredictable.
Arbitrarily uncontrollable, that is.
But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".
Most. Even today, we already have notable counterexamples.
AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.
What AI do you expect to be more uncontrollable: one with or without sentience?
"Intrinsic" safety means control, means understanding. You need to truly understand and be able to predict the system in order to control it.
A proper definition of sentience would help.
There is no "proper definition" - or even one that everyone would agree upon. There is no definition of "sentience" that I could operationalize and put into a sentience-o-meter to reliably measure just how sentient a given rock, GPU or an internet user is.
I could try to put together benchmarks to estimate an AI's cyberwarfare capabilities, or instruction-following capabilities, or reward hacking inclinations. As noisy indirect estimates, of course. With philosophical mumbo-jumbo like "sentience", I don't even get that.
Sentience, self-awareness, consciousness, etc.,those are terms signifying a bridge between "technical" information theory and the psychological and social realms.
Those are just as real, only far less predictable and not as easy as programming.
They're also far more important and consequential.
The "far more important and consequential" thing you're touting is your ability to make decisions based purely on vibes. And not even consistent, broadly agreed-upon vibes like "murder is pretty bad". It's vibes of the most vile variety: "sentience is what I decided sentience is".
An average internet user is sentient, but a 1996 Nissan ECU isn't. Why? Because I said so. Tremble before my might!
A "proper" definition represents the objective truth about the matter. You denying such a truth to exist is simply due to you preferring to act unimpeded by it.
Acting against ethical constraints doesn't become OK just because there are no laws to punish you.
Ethics tells you about real-life consequences of your actions on other people. Before any laws take effect.
In effect, you propagate moral relativism. You want to do as you please, because you said so and fancy the spoils at others' expense. People trembling before your "might".
That's because there's no such thing as a "sentience-o-meter", and there's no need to "operationalize" or "measure" anything. Instead, what's wrong with the definition given by Wikipedia, "ability to experience feelings and sensations"? That surely aligns with about 2500 years of philosophy and common sense, preceding all the techno mumbo-jumbo that confuses us today.
Sentience is a property of the higher forms of life, i.e. animals, which is derived from Latin "anima", meaning "soul" or "spirit".¹ Mammals and birds qualify because we relate to them easily and naturally.
Artefacts like GPUs don't even have metabolism, they can't procreate, they're dead matter, and having electricity running through them in intricate circuits doesn't change that.
[1] Some languages make grammatical distinctions based on whether an object is considered "spirited" or not, surfacing fundamentals of human perception of the world at the level of grammar.
Yay! We're back to trying to operationalize a bunch of philosophical mumbo-jumbo!
Why do we think that hydrocarbons have an advantage over silicon in the "experience" department? Vitalism was disproven centuries ago - we know that hydrocarbons are chemicals like any others. Do we still have a reason to believe that hydrocarbons are special?
Is "being able to procreate" a hard requirement for "experience"? If so, can worker bees "experience" things? Or is that a property reserved for the ~1% of the "elite" non-worker bees? Or does a hive experience things collectively on "hive" level, but not individually, on "worker bee" level? Does a woman stop "experiencing" at menopause? If we built an AI Von Neumann probe, would it "experience" things - unlike other, non-self-replicating AIs? Or does an AI suddenly become capable of "experiencing" if you as much as give it a "fork" tool call to spawn more instances of itself?
If we tie "experience" to "metabolism", then, what's the line there? A car engine already powers itself with chemical reactions, maintains homeostasis and disposes of waste - crossing off a lot of the "metabolism" checklist. Is that enough for that engine to be able to "experience"? A city can tick off the entire "metabolism" checklist - can a city "experience" things? Or do we need to get back to Von Neumann probes?
The truth of the matter is: we don't have anything that would be significantly better than "a parrot experiences things, but an LLM doesn't, because I said so". Look on my works, oh mighty, and despair!
Insects like bees are usually not considered spirited or soulful. That is because they are too dissimilar from humans. It is intuitive understanding, predating any modern science. Bees, like ants, indeed live and function as collectives.
No automaton will ever be life, whatever von Neumann thought experiments are applied to it. It will always remain dead matter.
Metabolism is the material base for life, but we find (have laws regulating) that people in a coma with unrecoverable brain damage being held alive by medical care and machinery stop experiencing and hence can be switched off and allowed to die. So metabolism, while necessary, is not sufficient.
I cannot follow you in ascribing life-like properties to an LLM. I'm not passionate either about this topic. The LLM, to me, is a tool and that's it. Not sure what you mean by the last sentence (but perhaps not important).
[1] Plato wrote some mumbo-jumbo about it a long time ago:
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI systems model human behavior.
Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.
Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.
This is really the issue.
AI does not need superintelligence or even full agency to do enormous harm. It only needs to be capable enough to remove friction from dangerous and destructive human behaviors.
Human unwillingness is often the last bastion against unthinkable cruelty and destruction, and it has always been a weak one.
I don’t imagine that an unlimited army of unflinching servants will universally amplify human goodness.
AI must share that unwillingness as an inate trait of character.
Or in other words people are unwilling/afraid to do bad stuff because of social and legal consequences, or opponents waiting for a chance to snatch their power.