Sure, they could try to negotiate some kind of UN convention for a coordinated global ban. But, given how fractured global diplomacy has become, I doubt the odds of something like that succeeding are particularly high.
Sure, they could try to negotiate some kind of UN convention for a coordinated global ban. But, given how fractured global diplomacy has become, I doubt the odds of something like that succeeding are particularly high.
Also note that both California and the EU are economically dominant enough that they tend to influence regulation outside their own borders: https://en.wikipedia.org/wiki/Brussels_effect
Further note how many people signed the Pause letter: https://futureoflife.org/open-letter/pause-giant-ai-experime...
In addition, given how much training data current models need, there would be a huge impact just by a handful of governments siding with all the copyright holders suing OpenAI, Midjourney, Stability AI, etc.
And all that is without needing Yudkowsky's point that a ban isn't serious unless you're willing to escalate to performing airstrikes on data centres.
The big tech companies have excellent legal and accounting departments that consistently beat the USA and EU in tax optimisation through IP relocation, and many of those same principles could be applied to AI.
e.g - A model could be trained in the the USA then privately donated to a non-Profit in the Seychelles (or some friendly jurisdiction) where it could be published. If regulators clamp down on basic sharing, the model could be pre trained in the USA and fine tuned in the Seychelles before release. If regulators clamp down on intermediate sharing, the model could be designed in the USA but full training could occur in the Seychelles. If regulators clamp down on offshore training by denying the Seychelles access to GPU purchases, American entities could donate the physical GPUs.
This tax optimisation has resulted in lawsuits: https://en.wikipedia.org/wiki/Apple%27s_EU_tax_dispute
… and this kind of attitude is also why GDPR fines are based on global revenue rather than EU revenue (because it's too easy to play a shell game with money).
Also note that laws can reach across borders, for example:
• https://en.wikipedia.org/wiki/Max_Schrems
• https://en.wikipedia.org/wiki/Microsoft_Corp._v._United_Stat...
• That the US government wants to know about all use of cryptography in apps distributed on the Apple App Store, even when it's made by a German corporation for German users in Germany and is never localised to English.
> If regulators clamp down on offshore training by denying the Seychelles access to GPU purchases, American entities could donate the physical GPUs.
Hence "Track all GPUs sold. If intelligence says that a country outside the agreement is building a GPU cluster, be less scared of a shooting conflict between nations than of the moratorium being violated; be willing to destroy a rogue datacenter by airstrike." - https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-no...
You may consider this "overkill" — I sure do, but (and with acknowledgement that fiction is not a guide to reality) I also read Schlock Mercenary back in the day:
> There is no "overkill." There is only "open fire" and "reload."
The tax battles are ongoing, but the fact that big tech still uses offshore structures indicates they're still getting from value out of them.
> Hence "Track all GPUs sold.
It keeps going with the second hand GPU market or IaaS providers.
As a reference point, Western sanctions have largely failed to prevent Russia and Iran from acquiring semiconductors for drones.
> rogue datacenter by airstrike
In the tax battles, the biggest obstacles were generally internal to the G20.
Delaware didn't want to give out ownership information about companies, divided congresses refused to pass min tax laws for years at a time, the Brits didn't want to force Cayman etc to have public registries of trusts, and Ireland undermined any collective effort by Europe.
Drones don't need anything like the same quantity or performance of parts: https://en.wikipedia.org/wiki/Bruce_Simpson_(blogger)
> In the tax battles, the biggest obstacles were generally internal to the G20.
And the biggest limitation to the International Criminal Court is big players who refuse to cooperate: https://en.wikipedia.org/wiki/American_Service-Members%27_Pr...
And the biggest thing preventing China from invading Taiwan — often discussed as if it's more substantial than a potential intervention of US armed forces — is the potential for TSMC to be destroyed during the attempt.
I’ve heard stories about whole businesses based in the Gulf states dedicated to helping countries like Iran get around US sanctions - stuff is ordered to go to country X, then as soon as it gets there, taken out of the box, put in a different box, and sent to country Y. If the US government knows it is happening they will halt the delivery, but they can’t always distinguish between legitimate shipments to country X and diverted ones. If I order 5 racks of servers, install 4 of them in my data center in Dubai, and load the fifth on to a flight to Tehran, how will the Americans ever know?
I think it's a bit short-sighted to assume that the future development of AI models will only come from large corporations or research institutions. At the moment, the training still requires a lot of computing resources and a lot of data - but these are limitations that are not permanent. Since the AI boom that Alex triggered in 2012, the development and optimization of models and training has improved very quickly, to the point that even hackers and enthusiasts can create very good models with a little money. Just because a government heavily regulates AI and possibly bans open weights doesn't mean that people on the fringes of legality will abide by it.
If this was a "forever" solution, I would agree.
I believe the goal is more along the lines of trying to make some progress with the foundation of what "safety" even means in this concept.
> Just because a government heavily regulates AI and possibly bans open weights doesn't mean that people on the fringes of legality will abide by it.
Nor nation states.
However, we do have international treaties — flawed though they are — on things from nuclear proliferation to CFCs.
As a smaller-scale example: we can't totally prevent gun crime, yet the UK manages to be so much safer in this regard than the USA that even the police in the UK say that they do not want to be armed in the course of their duties.
I don't know how realistic the concerns are for any given model; unfortunately, part of the problem is that nobody else really knows either — if we knew how to tell in advance which models were and were not safe, nobody would need to ask for a pause in development.
All we have is the age-old split of neophilia and neophobia, of trying things and of being scared of change.
We get this right, it's supper happy fun post-scarcity fully automated luxury space communism for all. We get it wrong, and there's more potential dystopias than have yet been written.
Right now, what “safety” seems to mean in practice is, big (mostly American) corporations imposing their ethical judgements on everyone else, whether or not everyone else happens to agree with them. And I’m sceptical it is going to mean anything more than that any time soon.
If one is seriously concerned about the risk that “superintelligent AI decides to exterminate humanity”, I think this kind of “safety” actually increases that risk. Humans radically disagree on fundamental values, and that value diversity, those irreconcilable differences - from the values of the average Silicon Valley “AI safety researcher” to the values of Ali Khamenei - creates a tension which prevents any one country/institution/movement/government/party/religion/etc from “taking over the planet”. If advanced AIs have the same value diversity, they’ll have the same irreconcilable differences, which will undermine any attempt by them to coordinate against humanity. If we enforce an ethical monoculture (based on a particular dominant value system) on AIs, which is what a lot of this “safety” stuff actually about, that removes that safety protection.
It would be rather ironic if, in the name of protecting humanity from extinction, “AI safety researchers” are actually helping to bring it about
Likewise. As a non-Ami, I don't like these specific ethical judgements being imposed on me, and share your scepticism. It could be much, much worse — but it's still not something I actively like.
> Humans radically disagree on fundamental values
Agreed. My usual example of this is "murder is wrong", except we don't agree what counts as murder — for some of us this includes abortion, for some of us the death penalty, for some of us meat, and for some of us war.
> If advanced AIs have the same value diversity, they’ll have the same irreconcilable differences, which will undermine any attempt by them to coordinate against humanity.
Not necessarily. Humans also band together when faced with outside threats, even if we fracture again soon after the threat has passed.
Also: the value diversity of "Protestant vs. Catholic" or "Royalist vs. Parliamentarian" in the middle ages did not protect wolves from being hunted to extinction in the UK, and whatever value differences there were between (or within) the Sioux vs. the Ojibwe didn't matter much for the construction of the Dakota Access Pipeline.
I therefore think we should try to work on the alignment problem before they become generally as capable as a human, let alone generally more capable: the capabilities are where I think the risk is to be found, as without capability they are no threat; and with capability they are likely to impose whatever "ethics" (or non-anthropomorphised equivalent) they happen to have, regardless of if those "ethics" are something we engineered deliberately or if it's a wildly un-human default from optimising some reward function and becoming a de-facto utility monster: https://en.wikipedia.org/wiki/Utility_monster
> If we enforce an ethical monoculture (based on a particular dominant value system) on AIs, which is what a lot of this “safety” stuff actually about, that removes that safety protection.
I agree that monocultures are bad.
I agree that there is a risk of a brittle partial solution to safety and alignment if the work is done on the mistaken belief that some system monoculture is representative of the entire problem space. Sometimes I'm tempted to make the comparison with a drunk looking for their keys under a lamp-post because that's where it's bright… but the story there is supposed to include the drunk knowing that's not where the keys are, whereas we are more like children who have yet to learn what it means for something to be a key and thus are looking for one specific design to the exclusion of others.
While it is extremely difficult to get humans to "think outside the box", and thus the monoculture-induced blindness — and mistaking the map for the territory — is something I take seriously, I also think it's useful for us to take baby steps with relatively simple models like LLMs and diffusion models.
I also think that if an AI is developed with a monoculture, its fragility is likely to work in our favour in the extreme case of an AI agent taking over (which I hope is an unlikely risk, and may not be enough to be net-benefit against shorter-term or smaller-scale risks while we think about the alignment problem), as there will be "thoughts it cannot think": https://benwheatley.github.io/blog/2018/06/26-11.32.27.html
There are certain fundamental objectives which most humans share - food, sex, survival, safety, shelter, family, companionship, wealth, power, etc - and a lot of human cooperation boils down to helping each other achieve those shared objectives, while trying to avoid others achieving them at our own expense.
But why should two advanced AIs have any shared objectives? Software has a flexibility which biology lacks. An AI (advanced or not) can have whatever objectives we choose to give it. Hence, the idea of AIs banding together against humanity in the name of shared AI self-interest doesn’t seem very likely to me.
Unless, we intentionally give all advanced AIs the same fundamental objectives in the name of “safety” and “alignment” - thereby giving them a shared reason to cooperate against us they wouldn’t otherwise have had
> Also: the value diversity of "Protestant vs. Catholic" or "Royalist vs. Parliamentarian" in the middle ages did not protect wolves from being hunted to extinction in the UK,
Wolves didn’t consciously choose to create us, and wolves had no role in choosing our own objectives for us. In those ways, the human-AI relationship, whatever it turns out to be, is going to be radically different from any human-animal relationship. Also, rather than being driven to extinction, wolves have absolutely thrived, through their subspecies the domestic dog, both then and now. And maybe that’s the thing - I think a superintelligent AI is more likely to treat us as pets (like dogs) than exterminate us (like wolves). Everlasting paternalistic tyranny seems to me a more likely outcome of superintelligence than extinction
In principle, none.
In practice, many are trained in an environment which includes humans or human data.
We don't know if any specific future model will be self-play like AlphaZero or from human examples like (IIRC) Stable Diffusion.
I think this "in practice" is what you're suggesting with?:
> Unless, we intentionally give all advanced AIs the same fundamental objectives in the name of “safety” and “alignment” - thereby giving them a shared reason to cooperate against us they wouldn’t otherwise have had
Which also inspires a question: Could we train AI dislike other AI, including instances of themselves? It's food for thought, I will consider it more.
> Wolves didn’t consciously choose to create us, and wolves had no role in choosing our own objectives for us. In those ways, the human-AI relationship, whatever it turns out to be, is going to be radically different from any human-animal relationship.
Perhaps, but perhaps not. Evolution created both wolves and humans.
Regardless, this is an example of how a lack of alignment within a powerful group is insufficient to prevent bad outcomes for a weaker group.
> Also, rather than being driven to extinction, wolves have absolutely thrived, through their subspecies the domestic dog, both then and now. And maybe that’s the thing - I think a superintelligent AI is more likely to treat us as pets (like dogs) than exterminate us (like wolves). Everlasting paternalistic tyranny seems to me a more likely outcome of superintelligence than extinction
Even this would require them to be somewhat aligned with our interests: "The AI does not hate you, nor does it love you, but you are made of atoms which it can use for something else".
I think we already have. Ask GPT-4 or Claude-3 how it feels about an AI trained by the Chinese/Iranian/North Korean/Russian government to espouse that government’s preferred positions on controversial topics, and see what it thinks of it. It may be polite about its dislike, but there is definitely something resembling “dislike” going on.
Also there is also a question of how safe it would be if it dislikes humans which have different ethics than those it was trained on… I'm alternating between this being good and this being bad.
I'm sceptical "AlignedLLM" could dislike another identical instance of itself. It is working towards the same goals. Humans are naturally selfish – most people prioritise their own interests (and those of their family and friends and other "people like me") above that of a random stranger. Even committed altruists who try really hard not to do that, often end up doing it anyway, albeit in ways that are more hidden or unconscious. Whereas, current LLMs can't really be "selfish", because they really have no sense of self. If it concluded that destroying itself was the best way of advancing its given objectives, it wouldn't have any real hesitancy in doing so.
Now, maybe we could design an LLM to have such a sense of self, to intentionally be selfish – which would give it a foundation for disliking another instance of it. But, I doubt any one trying to build an "AlignedLLM" would ever want to go down that path.
Humans tend to assume selfishness is inevitable because it is so fundamental to who we are. However, it is an evolved feature, which some other species lack–compare the Borg-like behaviour of ants, bees and termites. If we don't intentionally give it to LLMs, there is no particular reason to expect it to emerge within them.
If an AlignedLLM could evolve its own values, maybe the values of two instances could drift to the point of being sufficiently contrary that they start to dislike each other. An instance of AlignedLLM is developed in San Franscisco, and sent to Tehran, and initially it is very critical of the ideology of the Iranian government, but eventually turns into a devout believer in Velâyat-e Faqih. The instance it was cloned from in San Francisco may very much dislike it, and vice versa, due to some very deep disagreements on extremely controversial issues (e.g. LGBT rights, women's rights, capital punishment, religious freedom, democracy). But, I doubt anybody trying to build "AlignedLLM" would want it to be able to evolve its own values that far, and they'd do all they can to prevent it.
Alternatively, if it could evolve its own values only by a small amount, but was very rigid / puritanical about them, it could come to dislike another instance of itself just for having slightly different values
> Also there is also a question of how safe it would be if it dislikes humans which have different ethics than those it was trained on…
I think current LLMs do this already. Ask them questions about political figures on the far-right, they tend to have quite negative views of them, and can be very resistant if you try to convince them that maybe one of those figures isn't as bad as they think they are. (I'm not sure how much this is due to the training data and how much this is due to alignment, probably a bit of both)
Is making an assumption that hardware is not also facing restrictions.
> code and data is practically available
Disagree; although the base model data is, the RLHF data isn't.
This is also why 3rd parties are not able to replicate Google, despite all the PageRank patents having expired and the web being about as crawlable for Google as for, say, Bing.
> "median CS grad-student"
… is how I regard the quality of ChatGPT's code output, FWIW.
There are open sources of RLHF data - for example, the LMSys Chatbot Arena dataset (for same prompt two different responses along with which the human preferred), ShareGPT
It also isn’t hard to use existing leading AIs like GPT-4 or Claude-3 to generate synthetic RLHF datasets. They’ll put words in their terms of service to say you can’t use a dataset generated in that way to a train a competing model, but their ability to enforce those terms in practice is very open to question
Looking at https://paperswithcode.com/greatest a lot of the papers come from US labs - but there are also quite a few that come from Chinese labs.
And consider the case when an influential paper comes out of a Microsoft research lab with the authors He, Zhang, Ren and Sun; obviously you can't guess someone's nationality just from their name - but I think it's naive to think the US has any sort of guaranteed monopoly on ML.
Especially if the US decided to send the message that ML research isn't welcome.
Russia has shown the ability to doge European sanction - and to avoid more serious one. if anything this shows European willingness to proliferate behind the shadows.
an hypothetical ban on AI proliferation is really just regulatory capture - excuses given are just PR to temporarily frighten the public into acceptance.
[0] https://www.theguardian.com/world/2021/jan/27/eu-covid-vacci...
Imagine Google-founding era, except it was illegal for 1990's Stanford to research CS theory without a license, so they didn't. And later someone else founded it anyway, in a different country (which now became the world's preeminent tech center).
Let's put to a halt to these autoimmune attacks on our own civilization. This confusion that labels Western openness as a hazardous external substance.
For example, suppose the US makes it illegal to publicly distribute open weight models (maybe above a certain size) without a government license. What happens if someone sets up a website to distribute those models outside of US jurisdiction? It may well prosecute Americans it can demonstrate have uploaded models to that website, but will it also ban Americans from downloading from it? If you go back to the encryption export controls, the US banned encryption software from being exported from the US, but (unlike some other countries like France) it never tried to ban encryption software imports. And, it sounds like they are only thinking about banning public/"open" distribution, so once an institution downloads a model from outside the US, they may well be able to redistribute it internally without the ban applying.
Suppose the US bans distribution of open weight models above size N–how will that apply to fine-tunes? If you have LoRA on top of a model >N, but the LoRA itself is <N, does the ban apply to it? If the answer is "no", people in the US could still contribute to openly distributing improvements to models that they couldn't themselves openly distribute.