https://www.iihs.org/research-areas/fatality-statistics/deta...
8,244 karma · joined October 11, 2022
https://www.iihs.org/research-areas/fatality-statistics/deta...
There's another theory that says the best way is by putting a big spike in the driver's steering wheel.
So. I guess, if you believe that the only viable solution is model alignment, rather than relying on technical barriers to exfiltrating weights, then this is a decent steering wheel spike.
The main influential people I think of in that crowd who steer the conversations are Eliezer Yudkowsky and Scott Alexander. They're both doing fine financially at this point, I think, but not "buy politicians and laws like Elon Musk" fine.
There are some very rich people who might be loosely associated, like Dario Amodei (maybe?) or perhaps some VCs, but I'm not sure anyone really thinks of them as being a major influence among rationalists, more like people who hopped on the bandwagon. Certainly not ruling it.
I would understand "popular".
I am ten thousand percent in agreement that this is needed. And yet, I don't know how we ensure and incentivize that these precautions are taken by the actors involved, and as I understand it, even many of the top researchers in the field agree that they don't know what precautionary measures would even be effective, let alone sufficient. It seems that game theory has so far been pushing the AI companies to build it anyway without a robust solution for preventing harms, even while they express worry in public about the potential for those harms.
Again we can hold corporations or people responsible for damages after the fact, but many examples can be cited to show how that tends to be insufficient.
Don't you think that AI has by nature at least a little more potential for unpredictable results and unexpected harms than your typical technology? It seems like a major oversight to completely disregard this aspect.
We can maybe hold someone responsible when things go wrong and harm is done, and blame them for negligence as though the harm were intentional, but from a preventative safety perspective that's not usually sufficient nor even always helpful for preventing accidents.
I for one would like nobody to build the AI, is that an option? How do we get it on the menu?
Should we trust the people who say "nothing could possibly go wrong and there's no reason to worry or think about safety measures, if there are any trifling problems along the way we'll just figure it out as we go?"
As far as I'm aware, lots of them are actually deeply concerned about AI killing everyone.
What is their trigger condition? Will they get fired for pulling the plug? Do they get a bigger bonus if the servers keep running? Whose approval do they need? What response time is acceptable? How will they detect that the incident is happening?
It's easy to hand-wave "someone can just pull the plug" but there's an entire history of industrial accidents that happened because of the above problems of incentives, detection, procedures, not being taken seriously in advance. Someone could easily have pulled the plug on Chernobyl but nobody did, at least not before it was too late.
Faced with such a scenario, is the prudent next move:
a) blow one up and see what happens, or
b) do whatever you can to be sure it won't happen before conducting the first test, and make sure the confidence in the calculation is very very high
because I vote for b, and so did Teller.
Not sure how long we'd survive such a scenario, even sheltering underground. But surely it couldn't happen to us.
(However, it now seems like the AI might get us first.)
For example, if the head of cyber security at your company suggested there's no need to worry about hacker infiltration or worms because one can always unplug one's computer as the primary defense mechanism, you might find that a little lacking. Will you be able to unplug the computer before the damage is done? Will it spread to other systems before you detect it? How will you unplug the computer if the attack is from an external facility? What if an attack happens but the boss says the computers have to keep running because an important customer is monitoring uptime? What if the attack goes unnoticed because it looks like a benign service?
Now imagine the head of cyber security answers by saying "actually you don't even need to unplug them, you can just wait for the computers to overheat, thus solving all concerns."
Here's my answer, as a non-superintelligent human: "see to it that the humans on top of the situation have a compelling financial interest in the systems not disconnecting".
In nuclear engineering, where safety is taken seriously, it's not enough to end the conversation at "the humans in charge can always simply shut down the reactor during a meltdown" or "a meltdown has never happened before, so we don't have to design safety systems before one does".
Also, if you had to pick a year for when we see the first AI-controlled company with at least $100M in assets under its control, what date would you pick?
I'm reminded of the 2010s when people said an AI would never be able to break from its box and be connected to the internet, and then in the 2020s the first thing the AI companies did was offer AIs as online services. I expect in the 2030s fund managers will be eager to hand control of their billions to AI, so how the AI gets control of funds and a corporation is hardly an obstacle.
For construction work the AI would probably be hiring other companies.