From where I stand, it looks like the pandora box has already been opened anyway. The era of hugginface and/or Llama2 models is only going to grow from there.
From where I stand, it looks like the pandora box has already been opened anyway. The era of hugginface and/or Llama2 models is only going to grow from there.
I’m not in the piracy scene, but my impression was they routinely pass full res movies around the Internet without much barrier to discovering and downloading them, at least to technically competent users. Is that still true?
The difference is that a few people with lots of resources take on legal risk. In the piracy example many people with few resources take on risk, which works out since no one wants to sue people with no money.
Plus, while it's definitely the case that with sustained interest, old data tends to linger around... the moment the interest wanes, it's gone. I've been on the internet for a while and there are so many hobby sites, forums, and software projects from the early days that are simply gone for good (and not on archive.org).
Well yes, that’s precisely why they are lobbying for it.
If open source catches up, Meta can ride the tailwind and also catch up. I'm sure Meta will flip its position once it doesn't feel outclassed by the competition.
It's a homesteading land grab, plain, simple and pure.
There is a competitive landscape, first-mover advantages, incumbent effects, et al that are being anchored in e.g. Sam Altman's interests and desires at this very moment. If you want a vision of the future of garage AI, imagine a boot stomping on Preston Tucker's face over and over. The current AI industry's goal is to preserve an image of openness and benefit while getting ready to pull the ladder up when the time is right.
With Llama2 we're at Meta's mercy, as it cost 20M to train. No guarantee Meta will continue to give us next-gen models. And even if it does, we're stuck with their training biases, at least to some extent. (I know you can fine-tune etc.)
I'll argue that between stable diffusion and llama 2, there is nothing highly specific that prevents [very] large amount of people from adopting these models and specializing for them own needs.
The tragedy would be if those went away.
That situation will change as technology evolves. We'll eventually reach a point where a normal desktop PC can train AI. The wealthy will always be able to do it faster, but the gap will shrink with time.
The trick is making sure that laws aren't put in place now that would restrict our ability to do that freely once the technology is there, and that massive data sets are compiled and preserved where the public has free access to them.
Netflix cheats. They send non-supercomputer boxes out to ISPs to install locally. If I could convince every ISP to install a bunch of my media servers people could watch my shows from anywhere in the US too.
I don't think compute cost has dropped by 1000x since 20 years ago. Maybe by 10 to 50x. And if you add in the demand for higher quality, the cost has probably increased. Like encoding a video for streaming 20 years ago, at that standard, may have cost roughly the same as it does today, or more, when you factor in the increases in resolution and quality.
My prediction is that training the latest model will continue to cost millions to tens of millions for a long time, and these costs may even increase dramatically if significantly more powerful models require proportional increase in training compute.
Unless of course we have some insane algorithmic breakthrough where we find an AI algorithm that blows llama2 out of the water for a small fraction of the compute.
Comparing Netflix (and most highly profitable computer businesses) to the world of producing AI models by training is not going to be fruitful. Netflix takes a lot of effort to operate but you can do on the small scale what Netflix does, quite directly. You can't replicate an AI model like ChatGPT-4 very easily unless you have all the data and huge compute that OpenAI does. Now, once the model has been produced, you can operate that model on the small scale with maybe less amazing results (see llama.cpp, etc) but producing the model is a scale problem like producing high quality steel. You can't escape the need for scale (without some serious technological developments first).
By the way, I'm not sure how easy it will be to stop bad actors since barriers to entry are exponentially lower to developing a malicious AI tool than, say, developing a nuke.
The wrong hands will have the same access to whatever "superpowers" AI gives regardless of what regulations are or are not put in place. Regulations can't and won't stop potential bad actors with state-level resources, like China, from using any technology they decide they want to use. So trying to regulate on that basis is a fool's errand.
The real question is, what will put the good actors in a better position to fight the bad actors if it ever comes to that: a big tech monopoly or decentralized open source anarchy? The answer should be obvious. No monopoly is going to out-innovate decentralized open source.
> I'm not sure how easy it will be to stop bad actors since barriers to entry are exponentially lower to developing a malicious AI tool than, say, developing a nuke.
Since some bad actors already have nukes, the answer to this should be obvious too: it's what I said above about the wrong hands getting access to technology.
On a very good day, as many as three of those might simultaneously not be bad actors.
They aren't bad actors whose access to AI technology is likely to be meaningfully impacted by regulation (but, for certain of the non-US ones, that hasn't stopped the US from trying before), but that's a different issue.
Seriously? You don't think China, Iran, and North Korea are bad actors? What planet are you on?
But I don't see China or North Korea firing nukes or even blowing up western buildings. They are limited by the threat of response.
A rogue wacko in his basement can make Kim Jong Un look like Theodore Roosevelt.
Examples:
In the Iran Hostage Crisis you had a constrained actor (Iranian government) making somewhat rational choices to use hostage taking as a negotiation tactic.
In the Oklahoma City Bombing, you had unconstrained actors (Timothy McVeigh and Terry Nichols) blowing up a building with a vehicle borne improvised explosive device for personal reasons.
All of the AI danger propaganda being spread (see [1], for example) has the purpose of regulatory capture. You could have said all the same things about PageRank if it had come out in 2020. A malicious AI tool is harder to assemble than straight up cracking. The people who can do it are highly-trained professional criminals taking in millions of dollars. Those people aren't going to be stopped because the source is closed. (I'm thinking of that criminal enterprise based in Israel that could manipulate elections, blackmail any politician anywhere in the world... etc. They were using ML tools two years ago to do this.)
The ML tools are already in the wrong hands. The already powerful are trying to create a "moat" for themselves. We need these models and weights to spread far and wide because the people who can't run them will become the have-nots.
There are no superpowers, and the wrong hands are the ones least effected by any effort at restricting distribution by “strong regulation”.
The ones in control of the models also control what sentences are sanctioned, this is a problem the more widely LLMs are used. To add insult to injury, while we are not allowed private use of the models, governments and ad-tech surveillance capabilities will skyrocket.
Do you see the problem here? The capabilities of opensource models are not anywhere near high enough to justify such a cost, now or anytime soon.
And it won't end there. As the march of progress continues, we will see the AI doom crowd agitate for tighter surveillance of money flows, limits on private compute, bandwidth limits to homes, tracking what programs we run on our computers, on who is allowed to read the latest in semiconductor research and on and on.
Big companies are easier to regulate.
But the problem isn't regulating the big companies, or the smaller companies, or underground entities. The problem is state-level adversaries like China who might misuse a technology, whether it's AI or anything else. Such adversaries can't be regulated by laws or executive orders or UN declarations; they have proven that many times in the past. The only way to control them is to have sufficient counter-capability against whatever assets they have. And government regulation is a terrible way to try to achieve that goal.
We have tried this with China, going back to Nixon opening up trade relations in the early 1970s. It hasn't helped.
The idea that there's only one important problem is a fallacy.
I have made no such claim.
The people advocating for regulating AI are claiming it will solve all the relevant problems--i.e., that it will prevent AI from doing great harm. So pointing out a problem that the regulations will not solve is refuting the claims the advocates of regulation are making. That was my point.
I don't see why not. The whole point of regulations is to regulate, i.e., to keep the regulated activity within some particular bounds. If the regulation won't accomplish that, then it is pointless. Unless, of course, the actual purpose of the regulation is not the same as the purpose that is publicly stated--which is exactly what happens with regulatory capture.
The US has shown time and time again it’s complete incompetence when it comes to meaningful regulation of large companies.