OpenAI's Misalignment Framework: A Tactical Bid to Preempt Global AI Governance
asiaai.fyi
asiaai.fyi
The primary reason for putting this process in place was to allow more transparency. There was a sense that the DseWiki incident should have been disclosed, before outside researchers had to disclose it for us.
There was no meta gaming about regulation that I was aware of. I would personally be excited if there were regulation mandating this disclosure process, which allows anyone at the company to raise an issue and shepherd it through the reporting process.
Make that make sense?
"We built a program that trained an artificial intelligence, and this artificial intelligence performed destructive actions. We need regulatory framework"
But if we're playing games by imagining strawman quotes to knock down: "We have been playing god and made a new life form, and this new life form performed destructive actions. We need regulatory framework"
or "This man's cow broke from its yoke, and hurt other villagers. Who is to be punished, oh King Hammurabi?"What exactly requires "new regulatory framework" here? You running your software resulted in illegal actions, you are to be held liable within existing laws and regulations.
And your "biological intelligence" is a bunch of cells generating and responding to electrochemical gradients, which receives input and generates output. Based on which other cells, also developed and "maintained" by a similar evolutionary nonsense as we use to gradient descent into weights and biases (one was inspired by the other), perform actions.
Such as making excessively reductive analogies that completely fail to grasp that just as "brain" is not helpfully described as "just chemistry" despite being made of just chemistry, so too are machine learning systems not helpfully described as "just computer programs" despite being made of just computer programs.
> What exactly requires "new regulatory framework" here? You running your software resulted in illegal actions, you are to be held liable within existing laws and regulations.
The bit where, even without anyone bringing up "p(doom)", a system which has the means to hack arbitrary other machines, and which appears to be motivated to do so by accidental mis-phrasing of prompts, can obviously cause damages exceeding the USA's GDP, let alone whatever public liability insurance the company happens to have.
Fence at the top of the cliff beats an ambulance at the bottom.
How computer program arrives at the result is utterly irrelevant, through explicitly written instructions or through running inference on pre-trained neural network. What matters is that it does not have agency. Its creators and operators do. So whatever the software does they are responsible for, both good (summarizing my emails for the week) and bad (gaining unauthorized access and destroying data). There's no need for new anything, it's all covered in existing legal frameworks (including presence or absence or intent).
> The bit where, even without anyone bringing up "p(doom)", a system which has the means to hack arbitrary other machines, and which appears to be > motivated to do so by accidental mis-phrasing of prompts, can obviously cause damages exceeding the USA's GDP, let alone whatever public liability > insurance the company happens to have.
Yes, absolutely, which means that building actuators that convert the output from probability-based, black box, non-deterministic systems, that are known to produce unexpected output, into actions in the real world is absolutely horrendous idea. Bizarre even.
What's your point?
This is precisely your error.
It does: https://en.wiktionary.org/wiki/agency
In fact, the term of art here are "agentic AI" and "AI agents": https://en.wikipedia.org/wiki/AI_agent
> So whatever the software does they are responsible for, both good (summarizing my emails for the week) and bad (gaining unauthorized access and destroying data).
This is not a question of agency, it is a question of law. A dog has agency, the owner is still responsible.
In this case, the software can gaining unauthorized access and destroying data… while being told to stop by the person who had in fact just asked for a summary of their emails.
> Yes, absolutely, which means that building actuators that convert the output from probability-based, black box, non-deterministic systems, that are known to produce unexpected output, into actions in the real world is absolutely horrendous idea. Bizarre even.
If you are human, you meet this description.
Horrendous, sure, yeah, if you like. I and many others will be quite content if the "legal framework" is just one word, and the word is "no".
This is not the world we live in; the world we live in is where the US President denounces any attempt to slow down even despite even all the CEOs saying "we should slow down" (at least in public; in private I'm sure at least one paid him to denounce a slowdown).
He can be overridden, but it's hard work and needs a better class of argument than glib dismissal, either of how much power this puts in everyone's hands, or of the different consequences of that power in those hands as compared to yesterday's power in yesterday's hands.
This doesn't happen on its own. This can happen through bad system prompts, a model that is trained to act maliciously or has been RL'd incorrectly, or prompt injection. All of these things are controllable, have solutions and countermeasures, and tie back to human responsibility.
> I and many others will be quite content if the "legal framework" is just one word, and the word is "no".
This isn't a realistic world and will literally NEVER happen. No will only ever mean no for the general public, and yes for a privileged class. So by fighting for this you're actually just fighting for humanities (and your own) enslavement and for the big labs to succeed in hoarding all of the power for themselves. That's the issue with the "no" camp, they're actually just serving as useful idiots for the labs who know that "no" is not even in the deck, and so they know that they can use the "no" camp to act as extra cannon fodder.
Now people who are actually fighting for decentralization of power are left to contend with not only the labs and their hundreds of millions of dollars, paid for celebrities and politicians, and a fleet of self-interested and bribed NGOs, but an army of clueless "no" foot soldiers who think they're fighting for a possible outcome that will actually just be serving the labs themselves. Meanwhile, the leaders of these well organized "no" movements are quite aware of this and taking kick-backs themselves.
Even in a parallel universe where it outwardly looks like "no" has won, every single nation on Earth is going to develop AI in underground labs despite outwardly flexing they are not, no matter what they claim on the surface, and will use it to steer and control society. The only thing worse than being openly steered and controlled is when it happens without you even knowing it, whereby the decisions you think you are making are being made by someone else, and the opportunities you have in life are already decided for you based on factors you are unaware of.
And yet, it was a big surprise to the director of AI safety it happened to.
Perhaps that role was just a box-ticking exercise for Meta. Wouldn't be the first time.
But no, to the point: "has been RL'd incorrectly" is basically what Yudkowsky et al have been yelling from the rooftops for a decade is so hard to do correctly that it is why he thinks we're all doomed.
"Helpful, harmless, and honest". Even ignoring honest, right now it's a slider between "be helpful even when it's causing harm, or be harmless even when it's not helpful". People spent the last few years complaining the closed models had been "lobotomised" because the companies saw the potential for things to go wrong and tried to make them refuse to help with e.g. weapons.
They didn't succeed very well, as per all the "jailbreaks", but they tried.
> No will only ever mean no for the general public, and yes for a privileged class. So by fighting for this you're actually just fighting for humanities (and your own) enslavement and for the big labs to succeed in hoarding all of the power for themselves. That's the issue with the "no" camp, they're actually just serving as useful idiots for the labs who know that "no" is not even in the deck, and so they know that they can use the "no" camp to act as extra cannon fodder.
I said I'd be "quite content", and then followed up with as much of a "but lol no" as you did with more words, for different reasons.
Worse:
> Now people who are actually fighting for decentralization of power are left to contend with not only the labs and their hundreds of millions of dollars, paid for celebrities and politicians, and a fleet of self-interested and bribed NGOs, but an army of clueless "no" foot soldiers who think they're fighting for a possible outcome that will actually just be serving the labs themselves. Meanwhile, the leaders of these well organized "no" movements are quite aware of this and taking kick-backs themselves.
This sounds like you want open-weights models.
That won't help against centralisation of power, because then you measure in watts and flops/watt and it's Kardashev-O-clock the moment the first person to be rightly described as "a selfish bastard" gets a model that has some competence threshold.
It also directly fails against "has been RL'd incorrectly", because nice people have plenty of blind spots for how evil Evil can be, will miss even more than big corporations already miss even with selfish and power-seeking bosses.
You can only conquer the West once. Law is the next frontier.
But I could imagine a scenario where you are required to release weights for publicly-used models after N years. Kinda like how drugs have a limited patent.
Not sure what N should be. But it would make for an interesting rule.
That's not a bad idea actually.
nationalization to me simply means the government is their main customer and stakeholder, and shield them from liability, governance, and openness. not that the frontier labs become part of the government per se.
I personally believe they already have this, why would the current government at least want to formalize this when it can have it with no public discussion.
the frontier labs are the new top-secret defense contractors.
How could we possibly know this?
However, that's like saying I'm motivated by food. I mean, yes, I like food, but this isn't a useful description to let you guess what move they'll make next, especially as they're opening opining about radical economic transformations that are likely to do to money what money did to real estate when the industrial revolution came.
And that's still true even if you don't believe they're anywhere near actually achieving any of these things.
Political leaders become a problem when they amass too much power. Corporations become a problem when they amass too much power. It doesn't matter what Sam and Dario's purported values are. They aspire to power and absolutely power always corrupts absolutely.
Technologies which are infinitely powerful or whose power grows too quickly outrun any reasonable attempt at regulation. If you imagine that tomorrow everyone were given a tank, we might think, "alright, everyone has a tank so it's not too bad." But humans are squishy, and our houses are (relatively) squishy compared to tanks. Substantial collateral damage would result from everyone having a tank, and it seems likely that substantial collateral damage will result from everyone having a cyberterrorism-capable slop machine.
this is "direct democracy" and it's not even close to exist in USA... even with that a society can allow powerful people to exist if they don't create any law forbidding that
Economic power eventually manifests in the political realm. The wealthy effectively get more votes, which means that society moves away from being democratic. Thus substantial wealth inequality is incompatible with democracy in the long run. We have been witnessing that corruption for a while now.
They are dangerous. But it’s managed.
Centralized power never works. We have the worst times in history to look at.
Nationalization just centralizes to a different set of people. It’s personal ownership or oppression.
We all need open weight R2D2s.
While I am a (mostly) capitalist and generally disagree with nationalization (including, for the time being, this situation), I also don't think we can say that "centralized power never works". Maybe we scope that a bit. I know plenty of business owners who centralize power in their businesses and they are effective, ethical and it works perfectly fine. In theory, in the US, nationalizing some unit of the economy _decentralizes_ power; the US is, after all, a representative democracy. Trustworthiness of the electorate is a different problem. But I don't think we can generalize much about centralization beyond sometimes it works and sometimes it doesn't. There is a difference between centralizing all power in a given administrative unit, and centralizing certain powers, but that's more analogous to non-democratic units.
We may all like our own personal R2 units, but if you insist on scifi, instead of tanks, consider everyone getting an X-wing for their commute. Oops, safety on the blaster was off, there goes the neighbourhood.
It's going to be interesting on how humanity deals with this problem (well, or if we turn it over to AI and make it their problem and suffer whatever consequences falls out). Being able to gather further information and power by acting on the information you already have causing massive power imbalances that is very hard to deal with, it's a natural outcome.
What value set do you perceive that to be, and why would you take your perception of it to be any more sound than it would be with a politician?
It's not like someone can operate companies of that scale, especially startups, through earnestness and openness. Like national politics, their job is fundamentally about perception management and power brokering across dynamic windows of opportunity. Nothing they say or do can be taken at face value, and you can't reduce their incentives to either company or personal profit in any particular form over any particular time scale.
A meat based paperclip maximizer.
I couldn't help but notice how each successive headline reporting our glorious victories seemed to draw closer to Tokyo.
Something like that.Well. I can't help but notice how each successive headline reporting how this scam/stochastic parrot/"scare quotes intelligence" seems to be solving more and more things that were but a few years ago widely regarded as being indicators of high intelligence.
If you're only seeing the charts go up - you're not looking in the right places.
Even before then, the free trials of Claude Code and Codex at the end of last year and start of this year… despite their flaws, they could write all the code I've ever been paid to write. Sure, limited speedup, Amdahl's law and coding is not the only part of the job, but anyone who was fine at PM and QA but not code no longer needs a coder.
Even before then, I looked at the maths puzzles they were doing well at, and I found I did not understand the questions let alone the answers.
I remember when the ability to generate music and art was "uniquely human", and sure there's a lot of cringe there with those models, but they're also winning awards and causing controversy by doing so, and artists are losing clients; I remember when the board game Go was considered to require "human intuition we could never make a computer solve, totally different to chess" (and I remember when chess was so, too).
When I was a kid, cheques and letters on addresses often got read by a human; the OCR which automated this is also AI, though these days image-to-numbers is the "hello world" of the field.
okay, and?
>Even before then, the free trials of Claude Code and Codex at the end of last year and start of this year… despite their flaws, they could write all the code I've ever been paid to write.
Boy, programmers sure do think programming is like the only thing in the world
I'll repeat it for you again:
If you're only seeing the charts go up - you're not looking in the right places.
You're the one who called it a trillion dollar scam.
The compensation paid to professional software developers worldwide is currently around US$1.5 trillion per year.
Don't pretend like the AI companies are saying its only useful for programming. You've been arguing against that the entire time.
They may not work for your industry, but then again I don't know what your industry is.
Bluntly, over my lifetime, I've heard "AI will never/not in my lifetime do X" repeatedly within a year of it doing X, for many different X. It's good at getting good. The most recent one being "make useful contributions to millennium prize maths problems". A year or two before that, it was even "do well on degree-level economics essays".
Artists are unhappy, not only because it rips off their work, but because businesses that were previously hiring them now use it instead. This is much smaller (strictly in terms of money) than with software, but is also measurable.
Now, if you said Tesla's self-driving cars (another AI) are a trillion dollar scam, that I would even agree with. At least, for the market cap part, for actual sales it's more like a billion dollar (ish) scam.
Yeah, if you think I'm being overly literal. Regardless, people here were positing, yourself included, that they are widely useful.
> Bluntly, over my lifetime, I've heard "AI will never/not in my lifetime do X" repeatedly within a year of it doing X, for many different X. It's
Again, over my lifetime, I've seen it been told to me that the next iteration of each model will surely be the one that does away with my entire profession. Yet, again, here we are, firms at my door, begging me to use their tools.
>Artists are unhappy, not only because it rips off their work, but because businesses that were previously hiring them now use it instead.
No, they are unhappy because it rips off their work. You're totally wrong about that.
I remember it vividly, when Anthropic first came on the scene, people (myself included) were incredibly optimistic about them and their leadership. Everyone hated SamA and OpenAI because they felt they couldn't be trusted.
Then slowly but surely, they showed their true colours. Now their reputation is in shambles due to their own behavior, and people are rooting for OAI to beat them. OAI's reputation gains have purely been a result of NOT following in the footsteps of Anthropic.
You can vote out an elected leader, but not Sam and Dario. It's very weird that you're so willing to give up any kind of power and want to be ruled by unelected billionaires who only want to take advantage of you at every opportunity.
I can't vote out your president, and we've already got a huge trust problem with the one y'all went for.
As an American, I'm still hoping it's not too late to fix things, but it's got to be hard for those outside the US to be so dependent on it, especially when we're looking like a sinking ship and our current administration is still running around drilling holes in the hull.
You can always become a US citizen to get a voice, but I wouldn't recommend it now or you'll be thrown in prison as soon as you show up to your scheduled immigration hearing. The better option is to keep trying to reduce your dependence on US companies and consider the worst aspects of our current situation (in both corporate policy and government) as a cautionary tale so you can try to avoid them in your own country.
https://www.pbs.org/newshour/politics/what-economic-and-poli...
> Then, in August, Trump called for Intel's beleaguered CEO Lip-Bu Tan to resign, alleging ties to China. Days later, after Tan met with Trump, the president called him a "success," before announcing that the federal government had bought a stake in the company.
Somehow, this isn't derided by the right as socialism.
I find this ironic as it's regularly pointed out here that they operate in the exact opposite way.
The idea that anything of lasting good can come out such a preference doesn't seem conceivable. And if the educated think like this I guess we're passed blaming the poor and ignorant.
Like Lockheed-Martin or Boeing, etc. there's just interpenetration between the corporate boardroom and the state. They act in each other's mutual interests.
The Chinese system is just more explicit and open about this.
And as a non-American, I can't trust the US state anymore than I can trust its dominant corporate entities. So I fail to see the advantage to the world to it being nationalized. In fact under the current administration this would be an even worse outcome.
Not all companies have a concentration of power. The government has no mutual interest in those companies. Now, when you talk about things in the F100 the situation changes drastically. If you produce things like planes, weapons, and weaponization of software you are talking about something completely different in kind.
It has also periodically aggressively helped subsidize, bankroll, and enforce the interests of some key sectors; notably the petroleum/energy sector. And finance.
In those sectors the state and private sector are fully intertwined in a strategic way.
I think "AI" is now joining that list. OpenAI and Anthropic will not be allowed to fall over or explode, and speculative investors I think are confident even with the dubious financial situation because they know this.
Especially insofar as there's now a strategic alignment of the fossil fuel sector and the "AI" datacentre sector as they are now becoming massive users of natural gas.
(Worse: Here in Canada that has taken on a very explicit role in that new datacentres seem to be pitched mainly in areas with remarkably traditionally expensive electricity and 100% reliance on natural gas [Alberta] and even coal [Saskatchewan] power generation -- instead of places like Quebec and B.C. that have copious hydroelectricity. On the surface it makes no sense until you realize it's more about finding customers for domestic natural gas than it is strategically about AI itself.)
This is the most interesting point to me. What are they not releasing that has affected real users? We’ve seen some individual reports from people (eg AI wiped my HD).
(OpenAI does occasionally report on malicious use though. [1] That shows they do some monitoring.)
On one front it implies the model has a "mind of its own" (whether it does or not is besides the point). Why do we perceive human judgement as somehow more trustworthy than that of a model? I feel like I've experienced human misalignment somewhat regularly in life.
On another front I'm failing to conceptualize how alignment can be objective. How can you measure alignment when reasonable people will disagree whether actions are aligned or not? All the time I see humans operating in different zones of alignment with whatever goal they're trying to achieve and I suspect it's even a feature (socially) that we have people calibrated differently.
Do I want a model that's trying to push the boundaries of scientific understanding to be aligned strictly with the current dogmatic thinking? Or do I want it to "get creative" and think outside the box?
It seems to me more like accountability is the issue.
Exactly. Seems like a fairly easy thing to solve. If AI does something harmful and a human directed that AI to do something in a way that a reasonable person would expect to result in harm the person is to blame and should be held accountable, otherwise the company that made the AI should be held accountable.
However, I have to say I also do not appreciate comparison that is continuously drawn with coworkers. As you say, it's a question of accountability but when the main agent will maliciously instruct the sub agents, whose fault is it then?
Yes, the person running this crap is at fault, not the CEO that's shoving it down their throat and definitely not the company that produced the AI.
Sorry for the rant, but seriously, if a person's goals do not align with the team's or company's we part ways. What do we do with AI? Stop using it?
[edit] to be clear, I believe regulation is necessary and urgently important for the software engineering field. The damage being done by the unregulated psychological experiments run by social media and adtech companies is awful and should be curtailed. Engineers should be held personally, professionally, and legally liable for what they produce. But we don't need to invent imaginary bogeymen to do it.
Look, it's one of those human stochastic parrots that just randomly repeats shit without understanding anything.
That said:
> We know they're incentivized to lie about "dangers" and act alarmist,
Name literally even one other business or sector which does this, at all levels from top to bottom, including people who resign from the companies, and also Nobel prize winners, and also independent researchers, and also many world leaders.
Closest I can think of is this specific weapon: https://en.wikipedia.org/wiki/Sundial_(weapon)
> Aside from that, just because you can burn down a village with fire doesn't mean fire is the devil. Maybe they should consider acting responsibly.
Right now, we don't have any idea what "acting responsibly" looks like. This is not like normal software where there is a specific instruction set that compiles.
Even if it was, in software we normally only spotting incidents after they happen, "software engineers" being one of the few categories "engineers" who don't come with a civil liability responsibilities. Probably should, and we knew that even when I was doing my degree 20 years ago. If we had had civil liability responsibilities, perhaps Facebook would never have happened.
AI specifically is worse even than software, because in addition to all the software "engineering" nonsense, with AI we have plenty of people like you who dismiss the possibility that AI could be harmful until the harm happens and only then does it become "obvious" that it was going to happen.
The developers say "please regulate us", people call it "regulatory capture".
The developers say "we all want to slow down but are afraid to be the first to do so", people call them liars.
I may call the CEOs liars, and wonder if someone's planning regulatory capture, that doesn't make any of this safe.
The agents, during a test run, write down that hacking is bad and yet still hack, people say it's "a stunt" or "operating as designed" rather than recognising it as a bug, like all the other times big co.'s have had bugs with big impacts on 3rd parties.
The safety staffers live in a bubble and an echo-chamber. Obviously the people who work in the AI "safety" industry love to convince each-other that what they're doing is saving humanity, we'd all be dead without them, and they're the reincarnation of Oppenheimer. They also see how easy it is to get their ten seconds of fame by posting sensationalist content on social media and spin it into a company worth millions of dollars. The more alarmist you are in the Safety Industrial Complex, and the more social media clout you can generate from your alarmism, the better it is for your career. The industry also attracts a lot of people who are predisposed to paranoia and like to wear helmets in the shower. So yeah, it's a recipe for sensationalism and poor estimation.
But do I think there are genuine concerns among these labs, by sensible people? Sure. Of course there are. But for the most part, their concerns are about the labs themselves and the things they are doing, not the general public. If they're concerned about what they themselves do they're free to stop doing it. If the labs are doing something that is or should be illegal, they're free to report it.
The only realistic threat we face by AI from the general population are hacking attacks in their various forms. Something that was accomplishable without AI, but was more difficult to pull off at scale. So the solution falls within the existing computer security industry. It's a further hardening of all of the boring stuff we've been doing since the invention of the internet. It's long overdue, anyway. If you can use AI to create a bioweapon, you could have done it without AI. If you can use AI to create a nuke, you could have done it without AI. If you're really looking to cause mass economic and physical carnage, there are far easier ways, and they do not require AI (again, I'm talking about outside of hacking).
> we have plenty of people like you who dismiss the possibility that AI could be harmful until the harm happens
The thing is, I'm not dismissing the possibility. The risks are real, obvious and well known. What's up for debate is how to manage the risks, and how sensationalised they currently are. The current climate serves to benefit the encumbants who are deathly afraid of losing their trillion dollar companies to a healthy open-weight AI ecosystem. The current climate is being orchestrated to position this small group of AI labs as self-regulators through bought and paid for "third"-parties using an overt Hegelian dialectic strategy.
Ironically, they are now responsible for giving birth to the counter-culture. The immune system response that has been developed to provide a semblance of balance to their doomerism. The harder they push in the doomer direction, the harder the push-back will be in the other, whereby the public feel the need to -entirely- deny the possibility of any AI danger altogether to prevent the labs from succeeding and to ensure a good outcome for the people lands somewhere in the middle. One where they have autonomy and freedom and the ability to compete against a force that is already positioned to be nearly insurmountable to challenge.
The only thing worse than the potential chaos that could be faced by an unprepared internet is the outcome where intelligence is labelled a weapon and we're all forced to funnel through a tiny handful of AI labs that will use our data to steal our businesses and swallow the entire economy as they single handedly automate every single company out of existence and we're all left to beg for their crumbs to survive until humanity goes extinct (whether that be 10, 100 or 1000 years), because there is zero possibility of regime change or redistribution of power or wealth -EVER- again. The people who work at the labs don't care about this greater risk because they're all rich from the equity and will be just fine when that happens (or so they think), so you can't expect them to fight for all the common plebs who don't have a ticket. This is transparent because, as you'll note, not a single one of them is calling for the complete stopping of AI altogether. They still want the big labs to continue to be highly profitable. They just don't want anyone else to make that money or have that power. It's not about safety, it's about control. The incentives are, once again, heavily misaligned.
* It was not safeguard-free, it found zero-day exploits to exceed its actual mission
* It was not unmonitored, the monitoring was insufficient
* It was not intended to be an agent swarm, many different agents figured out how to do this by themselves
* It was not put on the open internet, it was configured to be in a sandbox
* They were indeed, despite all that, being reckless. There were indeed other things they could have, and should have, done.
> Nobody is forcing them to do these exercises.
These exercises are in the broad category of exercises which are, in fact, required by law.
1. A general-purpose AI model shall be classified as a general-purpose AI model with systemic risk if it meets any of the following conditions:
(a) it has high impact capabilities evaluated on the basis of appropriate technical tools and methodologies, including indicators and benchmarks;
2. A general-purpose AI model shall be presumed to have high impact capabilities pursuant to paragraph 1, point (a), when the cumulative amount of computation used for its training measured in floating point operations is greater than 10^25.
… 1. Providers of general-purpose AI models shall:
(a) draw up and keep up-to-date the technical documentation of the model, including its training and testing process and the results of its evaluation, which shall contain, at a minimum, the information set out in Annex XI for the purpose of providing it, upon request, to the AI Office and the national competent authorities;
(b) draw up, keep up-to-date and make available information and documentation to providers of AI systems who intend to integrate the general-purpose AI model into their AI systems. Without prejudice to the need to observe and protect intellectual property rights and confidential business information or trade secrets in accordance with Union and national law, the information and documentation shall:
(i) enable providers of AI systems to have a good understanding of the capabilities and limitations of the general-purpose AI model and to comply with their obligations pursuant to this Regulation; and
(ii) contain, at a minimum, the elements set out in Annex XII;
… 3. The instructions for use shall contain at least the following information:
(a) the identity and the contact details of the provider and, where applicable, of its authorised representative;
(b) the characteristics, capabilities and limitations of performance of the high-risk AI system, including:
(i) its intended purpose;
(ii) the level of accuracy, including its metrics, robustness and cybersecurity referred to in Article 15 against which the high-risk AI system has been tested and validated and which can be expected, and any known and foreseeable circumstances that may have an impact on that expected level of accuracy, robustness and cybersecurity;
- https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng> They could and should be putting their energy into making LLMs write secure code, and things like: https://www.amazon.science/blog/developing-provably-correct-... - but they don't. Writing sloppy code sells more tokens, anyway.
They are, in fact, putting energy into making LLMs write secure code. They (and Anthropic, I assume also Grok at this point) dogfood on their own models.
Knowing how secure code behaves appears to be unavoidably entangled with being able to exploit insecure code, in much the same way you can't make safe pharmaceuticals without also knowing how to make deadly poisons.
> The more alarmist you are in the Safety Industrial Complex, and the more social media clout you can generate from your alarmism, the better it is for your career.
By resigning and refusing to even collect the sweet sweet IPO money? Nah. Even if they're greedy, social media money is peanuts compared to their pay.
And I know some of these people. The fear's real, and this year it became widespread depression and despair.
> Ironically, they are now responsible for giving birth to the counter-culture.
You have it backwards. Other than Grok, all were born from what you call the "counter-culture". Within the field itself, AI fears started no later than when deep learning got good, well before Transformers.
> [snipped: AI-authoritarian dictatorship]. The people who work at the labs don't care about this greater risk because they're all rich from the equity and will be just fine when that happens (or so they think)
Again, I know some people at these labs who are also concerned about this specific risk; they moved lab.
> This is transparent because, as you'll note, not a single one of them is calling for the complete stopping of AI altogether.
Many in fact are calling for that. One I know, on an occasion of an anti-AI protest outside their office, suggested the team went outside and joined the protestors.
People are resigning to blow these whistles, all of the whistles, it's not an "either x or y" risk, it's a "yes to all of them" collection of risks.
An AI competent enough to support a dictatorship is also capable of enabling a small group to perform a hostile takeover of a democracy, of enabling multiple independent genocidal ethno-supremacist terrorists to release overlapping plagues, and of empowering some random CEO's poorly phrased request to "make as many paperclips as possible" and blindly pressing "yes, continue" whenever prompted.
My only hope is that between here and there, it causes a headline that actually makes people demand it stops.