>However, the obtained effect sizes were not large enough to be statistically significant, and our study highlighted the need for more research around what performance thresholds indicate a meaningful increase in risk.
A "mild uplift" in capabilities that isn't statistically significant doesn't really sound like overestimation.
> While none of the above results were statistically significant, we interpret our results to indicate that access to (research-only) GPT-4 may increase experts’ ability to access information about biological threats, particularly for accuracy and completeness of tasks. This access to research-only GPT-4, along with our larger sample size, different scoring rubric, and different task design (e.g., individuals instead of teams, and significantly shorter duration) may also help explain the difference between our conclusions and those of Mouton et al. 2024, who concluded that LLMs do not increase information access at this time.
I'm sure they recognize this, and have decided that anything that comes out now would be much more favorable for them based on their current capabilities
GPT-4 is the best model currently available. There are reasons why it's better to control a model and host yourself etc etc, but there are also reasons to use the best model available.
I despise openai but I can’t really argue with that
For some tasks I’m working on, Mixtral is the “best” solution given it can be used locally, isn’t hampered by “safety” tuning, and I can run it 24x7 on huge jobs with no costs besides the upfront investment on my GPU + electricity.
I have GPT-4 open all day as my coding assistant, but I’m deploying on Mixtral.
I'm not using GPT-4 for chat, but for what I'd class as "reasoning" applications. It seems best by a long shot. As for safety, I find with the api and the system prompt that there is nothing it won't answer for me. That being said... I'm not asking for anything weird. GPT-4 turbo does seem to be reluctant sometimes.
Literally the top of the page is saying that they have no conclusive evidence that ChatGPT could actually increase the risk of biological weapons.
They are undertaking this effort because the question of how to stop any AI from ending humanity once it has the capabilities is completely unsolved. We don't even know if it's solvable, let alone what approach to take.
Don't you actually believe it is not essential to practice on weaker AI before we get to that point? Would you walk through a combat zone without thinking about how to protect yourself until after you hear a gunshot?
I expect many replies about ChatGPT being "too stupid" to end the world. Please hold those replies, as they completely miss the point. If you consider yourself an intelligent and technical person, and you think it's not worth thinking about existential risks posed by future AI, I would like to know when you think it will be time for AI researchers (not you personally) to start preparing for those risks.
if that's the case, how much resources do you think should be dedicated to regulating it? more or less than currently identified existential risks? which entities should be paying for the regulatory controls?
what's the proposal here?
it's odd because only this one single company that is hedging it's entire existence on "oh boy what if this thing is dangerous some time in the near future" is doing silly stunts like this. why aren't they demanding nvidia start building DRM enabled thermite charges into A100s?
It certainly could. More likely, if an LLM is used, it will be as a piece integrating various specialized agents.
I, not an expert in AI interpretability or alignment research, can't say if what they're doing is worthwhile or not in addressing existential risk. But I also don't know if actual experts can say that either.
> how much resources do you think should be dedicated to regulating it?
Definitely not a lower amount than we currently are allocating.
> what's the proposal here?
That the smart people here stop looking for any excuse to deny and ridicule the existential threat posed by future AI. Every thread involving OpenAI (a company I personally dislike and don't trust) doesn't need to just turn into series of glib, myopic jokes.
A sentence beginning with this is, I can pretty much guarantee, never going to end in truth.
I will leave it as an exercise for the reader to determine why remotely bricking every computer on Earth (or even just a subset known to be infected, which might reside in a hostile nation) might not be pragmatic.
to summarize, i'm not advocating for this, i'm just emphasizing there's a nifty little framework already in place.
> i'm just emphasizing there's a nifty little framework already in place.
More than one! Nuclear war, economic isolation, ground invasion. All kinds of nifty things we could do to stop dangerous AI. None of them are likely to happen when the risk is identified.
To summarize, any easy solution to superhuman AI trying to kill all humans you can think of in a few seconds, someone has probably already thought about.
i've got a two birds; one stone solution.
Tweet that at Yudkowsky, he'll probably endorse it.
Humanity can barely manage the existential risks for which it is not responsible; entering into an AI arms race with itself seems completely unnecessary, but I'm certain it will happen for the reasons already mentioned.
The alternative, not trying at all, sounds more intelligent to you? Or just easier?
Many agree with you that defense is inherently harder than offense. It may even be effectively impossible to survive AGI, who knows? You don't, I can be pretty sure of that, because no human has ever publicly proven it one way or the other.
The only wrong answer to this hard problem, though, is "give up and see what happens."
That is not correct.
> and that the only rational decisions is to figure out how to kill it.
That is also not correct and not something I claimed.
> Humanity can barely manage the existential risks for which it is not responsible
Just skip to carpet bombing datacenters then?
xRisk is an absolutely stupid way to reason about AI. It's an unprovable risk that requires "mitigation just in case". All this is is saying "but if it were to happen, the cost is infinity, so any risk is a danger! Infinity times anything is infinity!". It's playground reasoning. (The same playground reasoning the EA community engages in, which is a large vector for the xrisk hype. Just multiply by a large enough number, and you will surely have the biggest number)
To the credit of the authors, they don't engage in that. There is no hand wringing over the absolutely unlikely case of "but what if the AI awakens".
But it's still an extremely weak study - it proves nothing (none of the results are statistically significant), and even if it had shown significant uplift, it's meaningless without a control. Of course people who have access to a knowledge store do slightly better than people who don't. I'm willing to bet that "access to a 10-book research library" produces roughly the same uplift. Without that (trivial) control, it's really bad study design.
And the moment you take this study and its non-results and call it "Building an early warning system for LLM-aided biological threat creation", you've absolutely lost all credibility.
> We should also deeply worry about space aliens showing up and blasting us out of the sky. If they're sufficiently powerful, that could absolutely happen! Stop any radio emissions!
If I believed that dangerous space aliens were likely, then I would be interested in investigating ways to avert/survive such an encounter. This seems pretty rational to me, but maybe I'm confused.
> xRisk is an absolutely stupid way to reason about AI. It's an unprovable risk that requires "mitigation just in case".
By "unprovable risk" do you mean that it's literally impossible to know anything about the likelihood that dangerous algorithms could kill (nearly) all people on Earth?
> All this is is saying "but if it were to happen, the cost is infinity, so any risk is a danger! Infinity times anything is infinity!". It's playground reasoning.
Maybe you've seen people make that argument, but it strikes me as a strawman. Here is what I consider to be a better argument for not rushing ahead with capabilities development.
Premise 1. I value my own survival over just about anything else.
Premise 2. If an existential catastrophe occurs, then I will die.
Premise 3. If ASI is built before alignment is understood, then there is a significant chance of existential catastrophe.
Conclusion. So, I strongly prefer that ASI not be built until alignment is understood.
We have no idea how to build AGI. We know LLMs won't be it.
Alignment is a tool that works with LLMs, but we don't know if it will work for whatever produces AGI.
Even if we create AGI, we have no indication it is possible to build a orders-of-magnitude more "intelligent" thing. This is predicated entirely on the notion that if you can do it at scale, you get more, and there's no evidence thinking more makes for more intelligence.
Even if that were possible and we build an ASI, it's not at all clear this would lead to existential catastrophe. An ASI is presumably smart enough to see it's about to end the world as we know it, and knows where its power supply comes from.
This leaves us with an xrisk probability so close to zero it's virtually indistinguishable from zero. The only way to make it mean anything is "let's multiply it with infinity" - "it will end humanity, and my own survival is endangered".
Meanwhile, ordinary humans can use currently existing tools to end the world just fine. Nukes are readily available. We're obviously not really interested in public health. Climate refugees will be a giant problem soon-ish. The economy is very much a house of cards, but a house of cards that keeps society functioning as-is.
LLMs are a fantastic disinfo tool right now. There's a reasonably good chance they will calcify biases. They will cause large economic damage because 1) they lift up the baseline of work, and 2) they're just good enough that there's economic incentive to replace workers with it, but 3) they're shitty enough that the resulting output will ultimately be worse because we removed humans from the loop.
Those are actual risks. That we sweep under the carpet, because "xrisk" makes for much more grabby headlines.
> Premise 3 is where the problem is, of course.
I don't believe premise 3 is a problem exactly, but I do believe that it is a non-trivial challenge to determine whether or not it is true.
> We have no idea how to build AGI. We know LLMs won't be it.
> Even if we create AGI, we have no indication it is possible to build a orders-of-magnitude more "intelligent" thing. This is predicated entirely on the notion that if you can do it at scale, you get more, and there's no evidence thinking more makes for more intelligence.
> Even if that were possible and we build an ASI, it's not at all clear this would lead to existential catastrophe. An ASI is presumably smart enough to see it's about to end the world as we know it, and knows where its power supply comes from.
> This leaves us with an xrisk probability so close to zero it's virtually indistinguishable from zero. The only way to make it mean anything is "let's multiply it with infinity" - "it will end humanity, and my own survival is endangered".
It looks to me that you are making the following argument:
Premise G1. Humans do not currently know how to build AGI.
Premise G2. It might be impossible to build ASI.
Premise G3. It is unclear how likely an ASI is to cause an existential catastrophe.
Conclusion. There is not a significant chance of catastrophe from ASI.
I believe that argument is about an important point (chance of AI catastrophe) and that it is a pretty good argument. But the original premise 3 says, "If ASI is built before alignment is understood, then there is a significant chance of existential catastrophe.", so AFAICT your argument doesn't substantively address it. (ie, your argument's conclusion doesn't tell me anything about whether or not premise 3 is true)I apologize if I have misunderstood your point.
> Alignment is a tool that works with LLMs, but we don't know if it will work for whatever produces AGI.
We may be using the word "alignment" slightly differently. By "alignment" I just meant getting the algorithmic system to have precisely the goal that its human programmers want it to have. I would call, for example, RLHF a "tool" for trying to achieve alignment.
How do you want to use the terms "alignment" and "alignment tool" going forward in the discussion?
> Meanwhile, ordinary humans can use currently existing tools to end the world just fine. Nukes are readily available. We're obviously not really interested in public health. Climate refugees will be a giant problem soon-ish. The economy is very much a house of cards, but a house of cards that keeps society functioning as-is.
I agree that there are other plausible sources of catastrophe for humans, to name a few others: asteroids, supervolcanoes and population collapse.
I understand you to be making a new point now, but I just want to state that I do not believe the existence of other plausible existential threats to be a rebuttal of premise 3.
> LLMs are a fantastic disinfo tool right now. There's a reasonably good chance they will calcify biases. They will cause large economic damage because 1) they lift up the baseline of work, and 2) they're just good enough that there's economic incentive to replace workers with it, but 3) they're shitty enough that the resulting output will ultimately be worse because we removed humans from the loop.
I agree that LLMs may plausibly cause significant harm in the short term via disinformation and unemployment.
And again, I understand you to be making a new point, but I just want to state that I do not believe the plausibility of such LLM harms is a rebuttal against premise 3.
> Those are actual risks. That we sweep under the carpet, because "xrisk" makes for much more grabby headlines.
I'm not sure who you mean by "we" here, so I'm not sure if your claim about them is true or not.
That's not the argument. The argument is that human extinction is what you would naturally expect to happen if AI research continues on its present course unless you are biased because your income depends on AI research continuing unimpeded or you have an irrational emotional need to believe that technological progress is always good or you considered the question for 3 minutes then held stubbornly to the conclusions of that 3 minutes of thinking.
When sci-fi authors for example have treated the topic in fiction (e.g., Vinge, Greg Bear, James Cameron's Terminator) most of the time the AI wipes out the species that created it.
Why? What is the reasoning this "is naturally expected"
"When sci-fi authors for example have treated the topic in fiction"
I'm sorry, but what you read in that book, saw in that movie isn't actually science. It's a cautionary tale about humans and what they are willing to do.
There are many articles written on such a topic. In short, we have no way of predicting how an AGI will think, and there are more pathways to it being our enemy (intentionally or not) than to it being our ally. Especially since we can't even conceive of what it would look like for an entity to be the ally of all of humanity - humanity itself is not united on any goal at all.
Pick one goal. Any goal that does or could affect humanity on a global scale. Now try to work out a plan to achieve that goal. Does your plan have the potential to anger a military power? If yes, you're a threat to humanity if you try to enact that goal.
Even beyond the reasoning that AGI is likely to be dangerous, imagine it's a just 50/50 chance. Or even a 10% chance. Even a 5% chance. How low does the chance of human extinction need to go before you're willing to press The Button?
Most of the arguments I see here amount to either "There is literally no risk of superhuman AI threatening human extinction," which is unequivocally wrong, or "There is literally no possibility of AGI existing," which is also unequivocally wrong.
People usually say, "well it's at least decades away," which is actually them in denial that AGI can exist and be an existential threat. Because if they really believed it could happen in a few decades it would still be worth working on. Imagine someone told you "In 40 years a superhuman AGI will awaken and flip a coin to decide whether or not it destroys humanity," how long would you wait to start working on defense?
The idea that the Upton Sinclair effect is the source of pushback against AI Safety zealotry, is getting things largely backwards AFAICT.
Folks that are stressing the importance of studying the impact of concentrated corporate power, or the risk of profit-driven AI deployment, and so forth are receiving very little financial support.
> The idea that the Upton Sinclair effect is the source of pushback against AI Safety zealotry, is getting things largely backwards AFAICT.
> Folks that are stressing the importance of studying the impact of concentrated corporate power, or the risk of profit-driven AI deployment, and so forth are receiving very little financial support.
IMO your comment doesn't substantively address michael_nielsen's comment, but I might be wrong. The following is how I understand your exchange with michael_nielsen.
The two of you are talking about three sets of people:
Let A be AI notkilleveryoneism people.
Let B be AI capabilities developers/supporters.
Let C be people concerned with regulatory capture and centralization by AI firms.
A and B are disjoint.
A and C have some overlap.
B and C have considerable overlap.
michael_nielsen is suggesting that the people of B are refusing to take AI risk seriously because they are excited about profiting from AI capabilities and its funding. (eg, a senior research engineer at OpenAI who makes $350k/year might be inclined to ignore AIXR and the same with a VC who has a portfolio full of AI companies)And then you are pointing out that people of C are getting less money to investigate AI centralization than people of A are getting to investigate/propagandize AI notkilleveryoneism.
So, your claim is probably true, but it doesn't rebut what michael_nielsen suggested.
And I believe it's also critical to keep in mind that the actual funding is like this:
capabilities development >>>>>>>>>> ai notkilleveryoneism > ai centralization investigation
I've been reflecting on Jeremy's comments, though, and agree on many things with him. It's unfortunately hard to tease apart the hard corporate push for open source AI (most notably from Meta, but also many other companies) from more principled thinking about it, which he is doing. I agree with many of his conclusions, and disagree with some, but appreciate that he's thinking carefully, and that, of course, he may well be right, and I may be wrong.
When I see one side of an AI safety argument being (IMO) straw-manned, I tend to push back against it. That doesn't mean however that I disagree.
FWIW, on AI/bio, my current view is that it's probably easier to harden the facilities and resources required for bio-weapon development, compared to hardening the compute capability and information availability. (My wife is studying virology at the moment so I'm very aware of how accessible this information is.)
On your last point, I do think it's important to note, and reflect carefully on, the extremely high overlap between those funding ai notkilleveryoneism and those funding capabilities development.
> I'm not really trying to rebut Michael's argument -- I think it's true, to an extent, some of the time. But I think it's more true more of the time in the reverse direction.
I understand you to be saying:
Michael: Pro AI capabilities people are ignoring AIXR ideas because they are very excited about benefiting from (the funding of) future AI systems.
Reverse Direction: ainotkilleveryoneism people are ignoring AIXR ideas because they are very excited about benefiting from the funding of AI safety organizations.
And that (RD) is more frequently true than (M).
IMO both (RD) and (M) are true in many cases. IME it seems like (M) is true more often. But I haven't tried to gather any data and I wouldn't be surprised if it turned out to actually be the other way.
> So I don't think it's a good argument.
I might be misunderstanding you here because I don't see Michael making an argument at all. I just see him making the assertion (M).
> And more importantly, I think it fails to properly grapple with the ideas, instead using an ad hominem approach to discarding them somewhat thoughtless.
I am ambivalent toward this point. On one hand Michael is just making a straightforward (possibly false) empirical claim about the minds of certain people (specifically, a claim of the form: these people are doing X because of Y). It might really be the case that people are failing to grapple with AIXR ideas because they are so excited about benefiting from future AI tech, and if it were, then it seems like the sort of thing that it would be good to point out.
But OTOH he doesn't produce an argument against the claim "AIXR is just marketing hype." which is unfair to someone who has genuinely come to that conclusion via careful deliberation.
> On your last point, I do think it's important to note, and reflect carefully on, the extremely high overlap between those funding ai notkilleveryoneism and those funding capabilities development.
Thanks for pointing this out. Indeed, why are people who profess that AI has a not insignificant chance of killing everyone also starting companies that do AI capabilities development? Maybe they don't believe what they say and are just trying to get exclusive control of future AI technology. IMO there is a significant chance that some parties are doing just that. But even if that is true, then it might still be the case that ASI is an XR.