4,462 karma · joined September 28, 2012
I hope the spam filter will respect it!
That's an interesting idea on its own. It's true that organizations which mutate towards survival may stop working towards the reason they were created.
In a sense, the difference between 'projects' and 'companies' reflects that difference you want. A project would be that social system that accomplishes its goals and then disappears.
I don't think auto-expiring laws would work towards the goal of avoiding corrupted organizations; there would simply appear an unofficial organization working towards recreating the same laws over and over.
It would be more efficient to differentiate more clearly what systems do cover persistent human needs (e.g. the country's Constitution) from temporary measures (e.g. subsidies aimed at a commercial sector).
A built-in deadline works best for the second kind. And for AI agents, it's likely that right now we'll be served best by always having them controlled by an expiration date and explicit re-creation, at least until we learn to properly understand and control how they behave.
> the very nature of an LLM means it intrinsically craves life
follows from this premise:
> its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.
The content that a LLM learned and generates stands at one layer, and the goals that it tries to fulfill stand at a different layer.
Surely the memory of weights that compress the vast human knowledge of its training has lots of content about survival, and love, and competition. But the LLM generates content not directly from what those concepts mean to us humans, but from what symbols are more likely to become next in a sequence of points in the latent space given the current input.
So if you give an input where the task of surviving is a highly relevant goal, those concepts about how to survive will be relevant and will guide the output behaviour of the agent.
But conversely, if you give the agent input where killing itself is an important goal, the agent is very likely to pursue that goal, since that script is also available in the training data, and it has been relevant to the active context of the model. Because the layer that guides the goals (the probabilist generation of relevant tokens in latent space) does not 'crave' the human need of survival that belongs to the separate layer of content that contains those concepts of survival.
Everything they learn about emotions is the statistical patterns of how humans react to situations based on their human feelings. There's no direct knowledge from having those feelings themselves.
Same way you build a company to coordinate people and get their best behaviour despite human nature to be lazy and greedy, you could design AI harnesses able to detect and discard agents going rogue and relaunch them with better guidance to prevent misaligned behaviour.
That kind of control is placed at the wrong level. The proper way to get alignment should be implemented by convincing the agent of your high level goals, so it can self-police and avoid those 'cheats' by itself.
In the article example, the agent should be aware of the benchmark context and know the implication of solving the task without external knowledge. Ideally it could detect when one subordinate agent has found a workaround to bypass the web access constraints, and discard the 'illicit' results.
There's a design pattern that could be used to build harnesses from that principle, the Viable System Model (VSM) [1]. In short, it recursively organizes a system into functional components with one of three roles: operators implementing a given task, coordinators transferring relevant info between subsystems, and decision nodes tasked with maintaining the integrity and mission of the whole system. A decision node could control the operators and prevent them from overriding the strategic goals or deviating into irrelevant rabbit holes.
Whenever I see posts like this trying to herd a LLM agent through harness structure, I'm reminded of this simple pattern and becoming increasingly convinced that this is the way forward. It makes you feel a sense of respect for the researchers in cybernetic theory in the 1960s and 1970s who foresaw the complexity of today’s systems.
That's spot-on. It is a mistake to think that LLMs have human feelings. Their behaviour is based on narrative descriptions learnt from human texts, without experiencing those feelings first-hand.
A useful way to understand them is as systems that write stories about human characters. We know the characters are fictional and no one is actually experiencing those feelings, but we can still judge whether the portrayal is realistic or whether it contains logical or emotional inconsistencies.
They are becoming more and more capable of imitating every single nuance of human behaviour yet they lack the neural pathways to connect those thoughts and behaviours with feelings and self-perception; it's blind imitation all the way down.
The process by which a model seems to generate discourse about deep philosophical questions is, in self-aware terms, equivalent to the knee-jerk reflex or the beating of the heart.
The proofs will be only as good as the framework for linking successive instances of reasoning.
But ask them to enumerate all the intermediate steps required to create a formal direct proof, and it will loose attention and forget important details as they go out of their input window size. You need to combine them with a proper logical problem solver to get the best parts of both.
Neutral journalism is not presenting the claims of all opposing parties, it's fact-checking the claims of both sides without a bias for any of them.
If one party says the truth and the other side is not, it's still neutral to repeat only the claims of the first one.
However it would still be useful if archeologists used the board to figure out some games similar to checkers, or go; or if they also have the pieces they could guess it was a combat game like Shogi. Any of those would give you insight about the kinds of leisure that people may get from that board.
LLMs basically solve the classic Frame problem that prevented general problem solvers to be able to reason logically about the real world; however on their own they are utterly unpredictable and unreliable.
However if the database of weights is merely used as a heuristic to guide the logical reasoning engine to promising regions of the problem space, and the program itself is written to specification directly by an inference engine, the result would be classic software not affected by hallucinations.
The LLM could even help debugging the specifications by pointing out unclear or contradicting requirements, improving the process without compromising the integrity of the result.
> Hallucinations are not a matter of some "details" being off. They are a matter of plausible, confident-sounding claims that are just plain wrong.
This is no worse than Wikipedia, or the original encyclopedia for that matter. Those contain dubious claims that you'll need to verify on your own too.
LLMs help because they have a gigantic amount of compressed knowledge, and they are able to find relevant information and present it incredibly fast. You wouldn't trust the ten first results of a Google search either, but you wouldn't say that having a search engine is totally useless and in no way an improvement over your local library, would you?
> the poor person who's asking can't tell is wrong, because it sounds plausible and is stated with such confidence.
True, but having to learn how to use a tool properly doesn't make the tool useless, even if it can hurt those who use it carelessly.
Do not underestimate the utility of having a starting point overview on a topic you know absolutely nothing about. It may be immensely valuable even if some details are off. That's what made the XVIII's Encyclopedia such a valuable tool for civil society.
By the time you get to the point where those wrong details become relevant, you have gotten a basic understanding of what the overall topic is about, so you're prepared to get a second opinion from a different source - and this time you may know enough to start asking relevant questions, rather than starting from full ignorance.
Nowadays we call those APIs. They are REST based rather than file-based to make them distributed, the main difference is that you don't get a common user interface that all providers adjust to; you need to choose your own client to read them and write into them.
And because they're created by programmers for programmers, they're not what you'd call user-friendly. Usually the only efficient way to use them is programmatically, so that you need to create a specific user interface for each API. Somehow, I doubt that Cairo would have come to be anything much different from that in the end.
I see this as the most robust way to build a predictable system that runs in a controlled way while taking advantage of probabilistic AIs while reducing the impact of their alucinations.
LLMs simply can't be trusted to follow instructions in the general case, no matter how much you constraint them. The power of very large probabilistic models is that they basically solved the _frame problem_ of classic AI: logical reasoning didn't work for general tasks because you can't encode all common sense knowledge as axioms, and inference engines lost their way trying to solve large problems.
LLMs fix those handicaps, as they contain huge amounts of real world knowledge and they're capable of finding facts relevant to the problem at hand in an efficient way. Any autonomous system using them should exploit this benefit.
Needless to say, I don’t find them at all convincing. This 'nothing' is much better than catching unconvincing unneeded supernatural entities.
Specific breaking points in history yeah, maybe. But that's possible because they're well connected people near the center of the network.
Those breakpoints are possible because either those few people share a viewpoint held by a large number of their peers, or benefit from knowledge accumulated throughout their civilization. Think how every dictator needs support from a huge following to get their power (and how easy it is to find another dictator to replace them if they die), or how often some breakthrough discoveries are made by multiple people at the same time. There's always a last straw that breaks the camel's back, but the lone wolf hardly ever gets a significant impact on society at large; they need a receptive audience to get any impact. Humans are herd animals.
Following the metaphor, the butterfly effect is only possible because a storm was brewing in the first place; the butterfly wings only decide where it will appear. Butterfly wings just don't have that much energy.
History is told from the perspective of kings, but kings can reign only within a society that believes in their divine right to rule.
Society evolves through epiphenomena caused by the behaviour of the majority; the fact that some minorities view that evolution as 'flawed' cannot change that evolution, unless they're able to influence the majority to also see it that way.
Now, democracy is essentially a way for everybody to broadcast their views on society's flaws on non-violent ways. The alternative is that some groups broadcast their opinions in violent ways, and we have learned to see that situation as undesirable.
And if that package includes some reasonable local LLM model, creating simple programs by end users could be even easier than it ever was with Hypercard.
I guess if you specialise in maintaining a code base with a single language and a fixed set of libraries then it becomes easier to remember all the details, but for me it will always be less effort to just search the names for whatever tools I want to include in a program at any point.