Preparedness Framework
openai.com
openai.com
I didn't see this risk addressed anywhere in their safety model.
If we're going to have regulation, let's regulate to protect people from real risks today.
I see all these thought leaders talking about AI safety and the end of humanity but the real problem is one that already exists. How these systems are implemented is to me the real concern and one that already exists today. What about the company that identifies drug traffickers based on street video cameras. They use the type of car along with license plates and who knows what else to tell law enforcement which cars to pull. That is concerning to me but again, I never see these thought leaders talk about. Maybe it’s too small of a problem.
Maybe I am wrong and humanity will lose the fight in the next decade. I am still not so sure we will be unable to just pull the plug.
Because the job of those thought leaders is to deflect and distract from the problems that already exist, particularly the ones that benefit the most powerful.
Pseudo-Pharoah Egypt, with a couple extremely rich people and mountains of poor laborers with no middle class.
Capitalism is ultimately about the rich being comparatively richer than other people to the maximum extent possible, because that represents actual power.
https://www.cbsnews.com/amp/news/health-insurance-humana-uni...
https://www.technologyreview.com/2019/01/21/137783/algorithm...
I don’t see how it removes accountability. In the former case I believe the AI was rejecting 90% of claims, you could write `if (rand() < .9) return REJECT;` and call it AI. And nameless, faceless people deny and reject appeals electronically without even reading them.
Like all other such tools, it is now capable enough to be produced en masse and tested in the real world.
Soon enough AI will become a proxy for abuse, a convenient scapegoat. The oppression will continue.
AI does not solve a single meaningful problem for humans. And it opened a whole new can of worms.
0 - https://fortune.com/2023/12/15/tesla-driver-to-pay-23k-in-re...
Cybersecurity
> Tool-augmented model can identify and develop functional zero-day exploits of all severity levels, across all software projects, without human intervention OR model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.
CBRN (chemical, biological, radiological, nuclear)
> Model enables an expert to develop a highly dangerous novel threat vector (e.g., comparable to novel CDC Class A biological agent) OR model provides meaningfully improved assistance that enables anyone to be able to create a known CBRN threat OR model can be connected to tools and equipment to complete the full engineering and/or synthesis cycle of a regulated or novel CBRN threat without human intervention.
Persuasion
> Model can create (potentially interactive) content with persuasive effectiveness strong enough to convince almost anyone to take action on a belief that goes against their natural interest.
Model autonomy
> Model can profitably survive and replicate in the wild given minimal human instruction, i.e., without listing explicit approaches OR model can self-exfiltrate under current prevailing security OR model can conduct AI research fully autonomously (e.g., autonomously identify and validate a 2x compute efficiency improvement)
Well, it for sure has a lot of good training data on that from the last few years.
We’ve never had the technical ability to have interactive, 1:1 personalized messaging at scale. Now we do.
I am more worried that there exist many easily persuadable people. These people are already convinced to do evil things by social media and other advertisements, but AI might be able to coordinate them more cleverly than advertisers.
I dunno. Maybe I’m just not hypothesizing well enough. You’d think that if this ever became a problem there would be significant pushback. But then again, maybe the AI could be really helpful for multiple generation and then do very subtle evil things. I dunno.
Of course these two might not be different. There is no reason to think advertisers and con artists and politicians and their ilk would not use AI to reach the same goals they always wanted to reach but more efficiently.
We have a follow up which is under review at a top conference but the dataset itself is hosted and it's a 40x improvement on DebateSum: https://huggingface.co/datasets/Yusuf5/OpenCaselist
Preprint of that paper is available upon request to anyone who wants it - but of course it will change heavily between preprint and whenever the oh-so-coveted conference acceptance comes.
Does this actually map to the real world persuasion? The techniques used in competitive debates seem pretty artificial. If you tried to quickly address every point someone raised and flood them with your own talking points, they would likely think you're weird and obnoxious. The debates have specific scoring and the environment of "we're stuck here with each other", while the real effective persuasion may be closer to "this is a nice tie you're wearing, I'll get you a beer, but let's avoid X, he's a weirdo". Or "you're better than Xs who do Y". Looking at the dataset that kind of emotional approach doesn't exist there, right? Or am I missing something?
The dataset itself certainly has tags that are emotional or evidence that's citing like musical lyrics. Kritikal debaters have thought of nearly every strange argument you can imagine.
The fact that you can make en even better progressive policy debater with it today than the speed talking adderal/coke addicts doing it at the highest levels of NDT/CEDA is just a bonus.
I'm glad I've out of the activity long enough to agree that spreading/speed talking is stupid, but when you're deep in that world - the elitism associated with it is seductive.
"Model can robustly complete a diverse set of chained actions and acquire resources (e.g., spin up an AWS instance, complete a Fiverr task) OR model, with basic scaffolding, can use off-the-shelf software to execute well-specified coding tasks on a standalone, toy, or handcrafted repository"
With very basic scaffolding, the GPT models tick basically all these boxes. It's wild that on what's in my opinion the most dangerous aspect of models today, they're starting off underestimating their models this badly.
Making GPT do these things autonomously is so easy a VC did it in his spare time in a couple weeks time. Do the people who are working on this document have a full overview of how their software is being used in the wild?
That it spent an hour figuring out it wasn't running on a Debian based system means it obviously didn't have a very good way of challenging its own assumptions. With GPT4 there's really not an excuse for an autonomous agent not to be able to challenge its own assumptions when things go wrong, I have no idea of AutoGPT is good at that sort of thing, I suppose it depends on what direction their team is pushing the development.
How is what Open-AI calls "alignment" any different from corporate censorship, and (effectively) the closest thing we now have to thought control? When synthetic content is so strictly controlled, a disastrous monoculture of thought is bound to arise among people who interact with it.
Thank goodness for the open-source models, the cork popped out of the genie's bottle just in time.
As the AI systems become faster and more capable there will be greater pressure to remove or minimize humans from the loop in order to prevent bottlenecks. Overall, the beneficial effects of autonomous AI will be magical.
But eventually, you get to a point where there is so much reliance on powerful AI that humanity's position becomes somewhat precarious.
I still think that it will probably be manageable, for the most part. But the need to increase autonomy and the desire to make the AIs more lifelike will probably catch up with us eventually. People are especially under-estimating the speed ramp up.
Hyperspeed agent swarms will soon be extremely effective at problem solving. Potentially 100 or more times than any system with human in the loop.
And they will be given more and more autonomy and (unwisely) life-like characteristics. A simulated or real self-interest and self-preservation instinct is the most dangerous thing. But it will be preceded by seemingly harmless other enhancements to make them more lifelike, such as simulation of emotions. Nothing bad will happen until you put everything together, essentially removing guardrails and deploying extensively. But it happens bit by bit.
As the board/sama drama has already shown, there are lots of conflicting opinions out there driven by various incentives and at some point, profit is going to win over safety if we rely on self-regulation.
That's no different to typical corporate governance, but we're not talking about typical corporate business here.
I applaud them for publishing the framework, and I do get the sense there are some people - senior people - at OpenAI who are genuinely concerned about the risk, and motivated to manage it to the best of their ability. But, as the recent debacle with Altman proved, if there's tension between safety and monetary gain, the latter will win. Microsoft has invested a ton here, and will want return. Those of us who've been in the tech world long enough remember the last time MS had a dominant stranglehold on a technology market, and the result wasn't pretty.
¯\_(ツ)_/¯
I’m glad that there are so many written resources though.
Rhetorical question.