1,890 karma · joined September 24, 2014
Bugs usually include business logic edge cases, and significant problems include someone realising while reading the PR that we can actually create a much better solution to the underlying problem. Ideally that realisation would happen prior to the PR being offered, but that's also not really how humans work
Example: During a PR review for a graphical feature for a game, someone reading it realises that we can actually have a significantly better solution to the underlying problem. You can't catch this in testing, because it doesn't even make sense conceptually to test it. It also sucks that it happened after someone put in a lot of work, but with graphics development you expect a lot of what you write to get canned and replaced with a better solution, because the technology evolves over time. The work is iterative towards the final goal anyway
Its also very common to miss subtle edge cases with graphics hardware, eg someone misunderstood the intricacies of GPU hardware, or a team member has relevant experience that someone else does not have. Or they missed a problematic memory access pattern on some hardware for example. Or something simple like they've technically forgotten a barrier, that the validation layer doesn't report for some reason
Eg: The function "tanh" is broken on some AMD GPU hardware, and should never be used under any circumstances. The actual GPU implementation of it is just screwed. Ideally everyone would know this, but its a very common function to crop up during specific graphics algorithms (as it smoothly remaps the range [-inf, +inf] -> [-1, 1]). So occasionally I've spotted that, and had to explain that we need to use an approximation instead, and then now everyone knows . It rarely gets caught during testing setups, because people don't know they need to include that hardware in their tests in the first place
This kind of sandboxing is not complex to do, especially for a company with OpenAI money. If you want your tools to explore hacking, you restrict them from internet access except for a whitelist of sites that have either opted-in, or you've very carefully vetted to make sure you won't cause any problems to. Its also not difficult to restrict their ability to make calls to be simulated, or to use fake tools that can only run the real commands if they're being run against the correct target
This is all incredibly basic security stuff to make sure you don't accidentally cause someone problems, and I simply don't believe these AI companies anymore. Its either intentional, or gross negligence
None of this is difficult to do, and its impossible to believe that a company the scale of OpenAI doesn't know this. I've built web crawlers and scrapers before, and the thing you do is test them extensively offline against simulated versions of the sites in question, and then very VERY cautiously run them against the prod versions so that you don't cause anyone any issues
The only reason not to do this is because OpenAI doesn't give a rats ass about the internet as a public good, nor the legal consequences of compromising systems
It isn't difficult to block certain kinds of network traffic, eg restrict the kinds of requests the bots are able to make. They also mention that the bots used developer only tools - why were they even installed on the machines that the bots were running on? Why aren't they reviewing network traffic, to make sure that incidents aren't occurring?
In this case, userdata was transferred to third parties by the bots - why do they have the ability to pass data to a third party? It is not complex to prevent this
This is literally the most basic kind of sandboxing and security, and the fact that OpenAI isn't doing it is clearly intentional. It is quite literally not believable that this hasn't been brought up internally as a problem
>"We have yet to understand the extent of existing incidents, and future rogue AI scenarios could be catastrophic," Krueger said.
This is why it smells like marketing, every time one of these incidents happens it reinforces the false notion that AI is sentient or acting on its own. Its intentional negligence by the AI companies to make the models seem more capable than they are to make line go up
We've been building firewalls and restrictions to prevent people from accessing sites on networks for decades and they're extremely effective. There's a whole industry built around this kind of security. The idea that these companies are incapable of doing it is wrong, they just don't want to put the work in because it makes a great ad campaign
It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
I'd love to see an in depth analysis of how much OpenAI actually did, but I suspect we'll never see that because it would indicate at least some plagiarism which undermines a lot of what OpenAI is putting out in public
That's why nobody's talking about how impressive this is, because its not nearly as impressive of a piece of work to simply cobble together other peoples' work that didn't know you were doing it. I could have republished relativity from einstein's notes, but people would correctly not be impressed with my ability
Until the plagiarism scandal is sorted out, its not a meaningful result at all, because nobody knows how much genuine innovation these models are displaying
We know that OpenAI trained on their prompts, plagiarism is incredibly likely. The only thing we don't know is whether or not it was deliberate plagiarism yet
This was unpublished research that was stolen, and constitutes plagiarism and academic fraud by even the strictest definition
If theft becomes more profitable than genuine creation, then nobody will create anything. Then there's nothing to steal, at which point all progress collapses
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Which seems to be very directly accusing OpenAI of plagiarism