Concrete AI Safety Problems
openai.com
openai.com
IMO, a better approach to AI safety research is to focus on securing the first channels that a malicious AI would be likely to exploit. Like spam, and security. Can you make communications spam-resistant? Can you make an unhackable internet service?
Those seem hard, but more plausible than the "Watch out for paperclip optimizers" approach to AI safety. It just feels like inventing a way to build a nuclear weapon that can't actually explode, and then hoping the problem of nuclear war is solved.
We should also be working to secure the channels an AGI might exploit, but most of those are already tied to an economic incentive to invest in security, and are already getting large investments compared to the relatively tiny field of AI safety.
That is precisely the approach that does not make sense to me. By the time friendly developers can launch AGIs, unfriendly developers will not be far behind. And AI safety seems likely to be an issue far before the development of AGI - a non-general malicious AI that can only hack internet services or trick humans into running arbitrary code via conversation is already quite a serious problem.
So to me focusing on "unfriendly developers building narrow AIs" seems more logical than focusing on "friendly developers building AGI".
So I don't think we need to worry about someone maliciously making an evil robot. That's already a problem and we already spent thousands of years figuring out systems of society to protect us from those dangers.
We all know strong AI wouldn't be some sort of robot running around shooting people like a movie, this would be an extinction event worse than a rather large asteroid if it fell into the wrong hands.
The first country or corporation to create strong AI capable of self evolution will either be able to immediately take control of the world as we know it, or worse create something capable of destroying humanity as we know it.
That's not what's going on at all for anyone else that at first glance though that (I'll admit I am guilty of thinking that after glancing at the article as well).
Goal 1: Measure our progress Goal 2: Build a household robot Goal 3: Build an agent with useful natural language understanding Goal 4: Solve a wide variety of games using a single agent
That's all they are trying to accomplish with this (so far anyway). None of which requires strong AI.
TLDR they are formulating some laws of robotics that are a bit more granular for dumb AI which are incapable of self improvement and very task oriented.
In other words the 99% of paperclip optimizers are going to start making virtual paperclips not paving over the universe with actual paperclips. Hacking your reward function is easer than solving hard problems.
PS: Don't do drugs :0
I think making friendly or unfriendly AI is similarily hard (you depend on the same variables just with different THEN clauses).
Making accidentally homicidal well-meaning AI is much easier than both IMHO.
The stuff you talk about is farther in the future, and a much tougher problem.
It would still be a "mistake", but not the "oops what does this button do" sort of mistake, but rather the "I could have sworn they fired first" sort of mistake.
That is why I think the main goal of AI safety should be how to defend against intentionally malicious AI. We should be thinking about drone swarms and hacker-AIs rather than paper clips.
> As long as AI is an open technology there will always be some criminals who just want to see the world burn.
Thankfully we have yet to see terror attacks with nuclear weapons, but the lower barriers of entry for potentially catastrophic AIs undoubtedly is alarming.
I often wonder if analog solutions will not be critical. To continue with your example, a nuke that must be armed with a physical lever would be a rather great hindrance to any AI. Internet communications, etc. will be more tricky, due to the high frequency activity, but the point of having meatspace 'firewalls' on mission critical activities is a nice short term kludge.
Access to that kind of data would help an AI determine the people that are more susceptible to manipulation. Add in records on health care, as you mentioned, and information on debts and you have data that can help an AI gather as many human minions as it needs.
This distinction is very important when comparing the threat of AI with other significant threats. Before nuclear bombs were built we could not tell you what the difficulty was in creating one. Now that difficulty is well defined, and we can use that knowledge to prevent them from being built by most nations, except the most well funded.
If the barrier for entry for AGI (then ASI) is lower than we expect, then the threat of AI is significantly different than if AGI/ASI can only be created by nation states.
The reason I am framing things this way is we need to be very careful here because we are starting to turn towards speculation.
Science fiction turns out to be true when physical reality agrees that it can be true. This again, is why we have a global communications network and personal wireless devices connected to it. This is also the reason we do not go faster than light.
The reason we don't have flying cars is they are completely possible. They are also terribly dangerous and expensive and a complete waste of energy.
The reason we don't have AGI is not that it is impossible, again if nature can create it, we can recreate it. Since we don't have a good understanding of the networked nature of emergent intelligence we cannot create a power optimized network that would allow us to create a energy efficient version. AGI itself is a complete waste of energy at this point. We already have many types of AI that are energy efficient and used in products now.
This is a ridiculous argument. Furthermore, even if it were true, it tells us nothing about the timeline. It could take 10,000 years for all we know.
I don't think you need AGI to cause a catastrophe. A narrow AI specializing in cyberattacks could be catastrophic, and is probably possible with current techniques.
Evolution does generate highly optimized systems but generally only when those systems have been around for tens of millions of years. Human-level intelligence has only been around for what, 50k - 100k years? We're probably still more in the 'just works' phase rather than the 'streamlined and optimal' phase.
That's how cold war remained cold.
Maybe. Can you make a human that can't be suborned by a superhuman evil AI?
Google Post: https://research.googleblog.com/2016/06/bringing-precision-t...
It was a pleasure for us to work on this with OpenAI and others. John/Paul/Jacob are good friends, and wonderful colleagues! :)
I think the scariest part of AI security is when the program itself becomes unfathomable. By that I mean, we can't just look at the source code and go "Ah! There's your problem". Now, your paper assumes a static reward function, but we can imagine the benefits of an AI that could dynamically change its reward function, or even its own source code.
In fact, the most powerful tool I can think of to train a multi-purpose agent is through evolutionary methods, and genetic algorithms. Take for example the bigger ideas behind https://arxiv.org/abs/1606.02580 [Convolution by Evolution: Differentiable Pattern Producing Networks] and http://arxiv.org/abs/1302.4519 [A Genetic Algorithm for Power-Aware Virtual Machine Allocation in Private Cloud], and determining the fitness of agents by the global accuracy on a large number of broad ML tasks. But I digress...
Given enough computing power and time, these have the possibility of ending in an "outbreak-style" scenario. [This exercise is left to the reader]. And the way AI ideas and methods are so rapidly disseminated and readily available, it's safe to imagine that it could happen in a relatively short time span.
Here's my question: I know you're with Google Brain, but do you know if OpenAI is actively researching these avenues of "self-determined" agents? For their first security-related article, I was expecting security measures along the lines of: safety guidelines for AI researchers, containment and exclusion from the Internet, shutdown protocols for the Internet backbone, etc. I get the impression some of these issues might rear their ugly heads before our cleaning robots become cumbersome.
P.S. Looking at your CV, it's funny to see that you once interned at Environment Canada. I'm also working there presently, during which time I can perfect my knowledge in ML to eventually transition careers. Small world...
Edit: Grammar.
The hard part is actually figuring out what you care about, particularly in the context of a truly universal optimizer that can decide to trade off anything in the pursuit of its objectives.
This has been a core problem of philosophy for 3000 years - that is, putting some amount of rigorous codification behind human preferences. You could think of it as a branch of deontology, or maybe aesthetics. It is extremely unlikely that a group sponsored by Sam Altman, whose brilliant idea was "let's put the government in charge of it" [1], will make a breakthrough there.
I don't actually doubt that AIs would lead to philosophical implications, and philosophers like Nick Land have actually explored some of that area. But I severely doubt the ability of AI researchers to do serious philosophy and simultaneously build an AI that reifies those concepts.
> The hard part is actually figuring out what you care about, particularly in the context of a truly universal optimizer that can decide to trade off anything in the pursuit of its objectives.
This seems basically equivalent to what they are saying. A reward function that rewards "what we actually care about." This might seem vague, but that's fine because these are only proposed problems.
The goal is avoiding unsafe AI. The reason such pointless efforts are wasted on this approach is we don't have a good alternative. The only thing I can think of is delaying it's creation indefinitely, but that's also a difficult challenge. For example, in the Dune books, the government outlaws all computers. That might work for a while.
Statements are adding noise and less than nothing of value if they just consist of telling people they are working on the wrong thing... and not proceeding to tell them what they should instead be working on, and giving clear positive reasons why (instead of negative reasons someone should not be working on something).
Incidentally this is a broader problem with HN discourse.
One might wish to point out that the emperor has no clothes and yet have no desire to plan his majesty's outfits for the next 6 months.
"How do I make a program make beautiful music" is a CS problem, but only after you have some notion of aesthetics in the first place.
In the context of a universal optimizer, "how do we make this program behave reasonably without bad side effects" is maybe a CS problem, but it's predicated on "how do we codify our notion of reasonable behavior", which is analytic philosophy with probably a bit of social science thrown in.
Problem-posing is itself difficult and how a lot of philosophical breakthroughs are made. If you want rigorous problem-posing where the solution would be handy for AI, hiring a philosopher might be a good start. Very few of us are equipped to do this kind of work, certainly not here in the comments section.
In fact, I'm surprised that there doesn't seem to be any reference in the article to previous work on these philosophical implications, e.g. the stuff that has been written by Nick Bostrom or MIRI. Perhaps there are some in the paper?
I think that for the forseeable future, we will inevitably end up with two of the problems that various philosophers have outlined over the last few years:
(1) How do we ensure that an AI agent does exactly what we want it to do and
(2) What do we ultimately want if we can desire anything?
I think that any developer trying to approach this will be doomed to hack around these two issues. We can probably come a long way in AI capabilities without having the optimal solution to this, but the core problem will remain for a long time and haunt those who are cautious.
In particular, I'm talking about verification and validation testing. I'm curious why generally these approaches are not being leveraged to ensure quality of output here.
I suspect this is because of the persistent belief that AI will annihilate humanity with one mishap, but I'm suggesting that we approach this much more like traditional engineering problems, such as building a bridge or flying a plane, whereby rigorous standards of are continually applied to ensure the system behaves as designed.
The resulting system will look much more like continuous integration with robust regression testing and high line coverage than it will be the sexy research ideas presented here, but I can't help but think it will be more robust. These systems are too complicated to treat them as anything but a black box, at least from a quality assurance standpoint.
Yes. That's why I was critical of an academic AI effort which attempts automatic driving by training a supervised learning system by observing human drivers. That's going to work OK for a while, and then do something really stupid, because it has no model of catastrophic actions.
[0] http://www.wireheading.com/ - David Pearce's ideas are.. interesting.. to say the least ;)
But let's change the question up a bit... There are billions and billions of flying intelligences on this planet. Birds, insects, even mammals. Nature has already created that. We've created things that are even better at flying fast and carrying more weight. So simply looking at 'flying cars' and saying they didn't happen so AI can't happen is at the least, very ignorant.
If nature can create something randomly, we can create something directed in a shorter period of time (well, we don't really have another 4 billion years to try). AGI is an eventuality.
We have evidence of intelligent species (ourselves) invading. We also have evidence spaceflight is possible.
An alien invasion is speculative at this point because there is no alien species on to which we can base any scientific refutation. It is not impossible, it is simply impossible to define any probably of occurrence.
"if done right" glosses over quite a bit, yes.
I'd tend to think the reverse: the idea that we can recreate general intelligence requires some degree of speculation, but once we have general intelligence running on a computer, it takes a fairly contrived set of conditions for it to not scale up exponentially.
I assume this probably already worked into the current prototypes. Does anyone have references to discussions about this in current gen self driving car prototypes?
Patrick Lin (October 8, 2013). "The Ethics of Autonomous Cars". The Atlantic. http://www.theatlantic.com/technology/archive/2013/10/the-et...
Tim Worstall (2014-06-18). "When Should Your Driverless Car From Google Be Allowed To Kill You?". Forbes. http://www.forbes.com/sites/timworstall/2014/06/18/when-shou...
Jean-François Bonnefon; Azim Shariff; Iyad Rahwan (2015-10-13). "Autonomous Vehicles Need Experimental Ethics: Are We Ready for Utilitarian Cars?". arXiv.org. http://arxiv.org/abs/1510.03346
Emerging Technology From the arXiv (October 22, 2015). "Why Self-Driving Cars Must Be Programmed to Kill". MIT Technology review. http://www.technologyreview.com/view/542626/why-self-driving...
To me this is the toughest nut in the lot. Training a Pac-man agent to avoid ghosts and eat pellets, in a world of infinite hazards and cautions! Any strategies?
I don't think these techniques transfer easily to the AI field. While I might be able to prove that the state machine that controls my nuclear power plant always rams in the control rods in case something bad happens, it's a lot harder to show that some fuzzy system like a neural network doesn't exhibit kill-all-humans behaviours.
However -- many of these techniques do transfer to the AI field -- albeit with some tweaking and careful thought.
Requirements are still utterly critical. Phrasing the requirements right is important and requires more than a passing thought -- particularly as concerns testability.
A lot of it boils down to requirements that get placed on the training and validation data sets; and the statistical tests that need to be passed: how much data is required and how you can demonstrate that the test data provides sufficient coverage of the operating envelope of the system to give you confidence that you understand how it behaves.
The architecture is critical also -- how the problem is decomposed into safe, testable and understandable subsets -- which has much more to do with how the system is tested than how it solves the primary problem.
Oh brother. Avoiding negative side effects is a wasteful proposition. Learning from those side effects, however, is priceless.
Also the related stories that explore this idea: http://www.fimfiction.net/group/1857/the-optimalverse
One problem isn't nonexistent or irrelevant just because there's another problem that you regard as worse or more urgent. It's not even like solving the problems of AI safety requires the same kinds of people or the same kinds of resources as solving the problems of land mines; if you tell people not to think about AI safety it's not really going to make them go away and solve the land mine problem.
I get it; we don't have land mines in first-world countries, and we will have AIs, so AIs are more interesting to talk about. That's why we continue to have land mines all over the world I think. Not our problem.
All the issues surrounding implacable AI killers on the loose are only something to talk about, if you haven't lived with them for generations already. Want to get real answers to sophomoric questions about robot killers? Just ask the people who already know.
It sounds as if that's intended to be an objection to something I wrote, but I've no inkling what. I certainly didn't mean to deny that there are a hell of a lot of them out there.
> we don't have land mines in first-world countries
I think the chances of a productive discussion would be greater if you didn't leap straight to assuming bad faith on the part of the people you're talking to.
Land mines are a big deal. They're a problem that needs solving. But you're not merely saying that; you're jumping into a discussion of something else and saying "you shouldn't be talking about this at all as long as there are land mines".
Which would be at least somewhat consistent (albeit rude), if that were your response to every HN discussion of things less important than land mines. But it isn't. By the advanced technique of clicking on your username, I see that you've been quite happy to participate in discussions of "table-oriented programming", mobile phone headphone jacks, and off-by-one errors in audio programming, and that you work in embedded software development. Are those things, unlike AI safety, more important than land mines?
I doubt you think that headphone jacks are more important than land mines. So why do you react to a discussion of headphone jacks by talking about headphone jacks, and to a discussion of AI safety by saying it's ridiculous and sophomoric to ask about AI safety when there are millions of land mines out there killing people?
You're trying to make out that the reason is that land mines are the same kind of things as hypothetical unsafe AI systems because they are human-made machines that kill people. But you're an intelligent person and surely you can't possibly really believe that. To deal with land mines we need treaties to stop them being deployed, we need ways of finding them that are cheap enough to deploy in quantity and effective enough to be worth deploying, we need ways of disarming them with the same qualities, and we need effective help for people who get blown up by them. None of these bears any resemblance to anything we might do about AI safety. To an excellent approximation, there is no overlap between the people who can do useful work on AI safety and the people who can do useful work on land mines. And the dangers don't arise in the same way: land mines are dangerous because they are put in place with the specific intention of killing anyone who passes, whereas in the scenarios AI safety people worry about no one intends the AI systems to cause trouble.
So that can't really be it, I think.
Why do you object to discussing AI safety but not to discussing mobile phone headphone jacks, really?
http://xkcd.com/1696/ (the current xkcd)