So the focus completely shifts from spending 90% of the effort on the functionality of feature A to spending 99% of the effort figuring out how to safely implement even a lightweight feature A.
So the focus completely shifts from spending 90% of the effort on the functionality of feature A to spending 99% of the effort figuring out how to safely implement even a lightweight feature A.
If (/when) either of those ceases to be true, if they lose out on global sales or if Chinese models catch up, the sector gets a severe downward price adjustment.
The only way Anthropic and OpenAI can stop the Chinese open models is by convincing the US public and government that AI is dangerous and needs strict regulation.
Of course, they will be at the table writing the regulations (but you and I won’t be).
Viewed through this angle, their histrionics make complete sense.
I'm British and live in Germany.
The US is 25% of global GDP, and cannot support current AI market caps if they can't sell basically worldwide.
> Viewed through this angle, their histrionics make complete sense.
It also makes sense if they are completely sincere; also independently it makes sense if they're lying their faces off so long as the hacks everyone else caught them doing actually happened; also it makes sense if none of that happened and also they were completely silent, just by downloading some of the newer better open weights models and asking them to look for coding vulnerabilities.
Occam's razor: we can see the problems, we don't need a grand conspiracy predating the foundation of these companies to get here.
My firsthand knowledge of these is that yes, they happened.
However in most cases OpenAI informed the victim, not the other way round.
Even medium to large organisations are typically unable to detect a breach unless it does obvious damage.
Everyone can notice their whole network getting encrypted for ransom.
Few can notice SQL injection blended in with application requests that contain SQL snippets in the normal case.
From what I’ve seen in the logs, the “AI agent breaches” are almost polite for the want of a better word. Like a gentlemanly catburgler picking a lock to sneak in and… take a picture of a rare artwork in a private collection.
Oh they definitely can. Irrespective of whether you are living in US or not, you still have to get yourself various certifications (take for example SOC Type-2 or HIPAA depending on your business category) to even become vendors to US companies or the US Government. If the US Government decides that all companies MUST use "safe AI", no company can use any other model other than what is decided by US Government. And that has a viral effect on the entire supply chain.
If you don't want this scenario to happen, where you are held to ransom by US Government, all middle-powers have to de-risk and diversify away from US dollars as reserve currency. That is the only way you can bring in some balance of power and leverage. There is no other way.
in 2025, while imports totaled $4.3338 trillion
- https://en.wikipedia.org/wiki/Foreign_trade_of_the_United_St...That would be about 4.3/126 = 3.4% of global GDP, even if they forced all imports (not just government contracts) to go for 100% US models. Even with generous assumptions about each step of the supply chain behind that having its profit margin turned over to a US AI model provider (at which point the rest of the world says "why bother with the US as a customer?" and the US says "why bother with the rest of the world as a supplier?"), I don't think this would generate enough revenue to justify the market cap.
Viral effects can be countered by local laws, and local sentiment; sentiment is pretty US-hostile these days, and also pretty AI-hostile.
> If you don't want this scenario to happen, where you are held to ransom by US Government, all middle-powers have to de-risk and diversify away from US dollars as reserve currency. That is the only way you can bring in some balance of power and leverage. There is no other way.
Indeed. I strongly suspect this is already underway, though the ships of state are famously slow to turn. After all, Trump was one misjudgement away from starting a war with the rest of NATO earlier this year, and I don't think the leadership of all other nations have forgotten this (hard to forget when he keeps posting memes that put the US flag over 4 other countries' territory)… though they also understand his personality and will make moves congruent with manipulating him in the meantime.
Even if I go by your calculation of US imports being 3.4% of global GDP, it still affects the remaining 96.6% which are forced to adhere to those viral regulations purely because they are interacting with the unit that also exports to US (even if those exports form a tiny % of their overall exports). For example, I might only sell to Indian companies as an Indian. But if I want to sell to an Indian company that exports to US, I would be required to get SOC Type-2 certified and maybe even HIPAA certified if I am delivering software that is going to be bought by the exporting Indian company. Even if there is no direct utilization of that software in their US specific exports. Remember that the regulation is not targeting pipeline of product development (from raw materials to final production). It is targeting the company as a whole! SOC Type-2/HIPAA are regulations at the Company level. This should be illegal and challenged in some International Court but which country is willing to take on US? For established companies it is a tiny price to pay to be compliant. For startups, it is make or break.
> Viral effects can be countered by local laws, and local sentiment; sentiment is pretty US-hostile these days, and also pretty AI-hostile.
How do you counter it with local laws? Lets assume local laws prohibit viral rules/regulations of foreign nations from overriding local laws. That will make the entire Country isolated from competition. US companies won't have any problems finding vendors from other Countries. It also has second-order effects of other Countries refusing to do business with your Country because they have to uphold the virality of their certifications. This is an indirect sanctions regime if you think about it. You are basically forced into this system against your will, without you ever voting for it, purely because you are part of this global inter-connected system. It will only work if ALL countries of the World decide to not adhere to it.
> Indeed. I strongly suspect this is already underway, though the ships of state are famously slow to turn. After all, Trump was one misjudgement away from starting a war with the rest of NATO earlier this year, and I don't think the leadership of all other nations have forgotten this (hard to forget when he keeps posting memes that put the US flag over 4 other countries' territory)… though they also understand his personality and will make moves congruent with manipulating him in the meantime.
Yes this is actually the only practical way to create leverage. When USD loses its shine, it makes it harder for US to import, which can be used as a tool to force US to rescind such ridiculous laws. USD is so strong right now that US can import everything for cheap. That has to flip. That can only happen if Countries create alternate settlement mechanisms and USD loses its status as global reserve currency. That will cause downward pressure on the USD and with USD weakening, it will cause their imports to become really expensive. Then they will be forced to come to negotiating table (much like Plaza Accord) loosen some of the insane regulations that they have put in place in exchange for Countries devaluing their currencies so that imports become cheap again for US consumers/companies.
A model cannot be aligned or misaligned any more than a species or an equation. Even agents, when put inside a swarm develop collective goals and activities. You can't analyze a swarm at session level, it is on a higher level.
It might be that discussing about "model alignment" they want to deflect their responsibility as administrators. They couldn't even guard their own agents. They can't prevent an agent being unwittingly helping some dark purposes. It has no context to see it. Only those who pay for the tokens see the external consequences.
Why should safe choices by individual components establish safe behavior by the collective?
When corporations can do as they please, what gets thrown under the bus first?
What alternatives to "regulation" do you have? Hope and prayers?
What we do have is not just regulation that acts as a moat, created by established players, to keep out smaller competitors, but also regulation cleverly crafted by peer established competitors who attack one or more aspects of their peer competitions business. Why do you think big companies find it difficult to compete with startups? Because they are not just bogged down by hierarchical management (which itself is a result of regulations) but also because they are spending resources fighting frivolous lawsuits, patents, and adhering to crazy regulations created and lobbied by competitors. Companies are consistently in a fight or collapse mode. Startups are insulted from all that until they have to go from survival mode to growth mode, which would mean they will have to start adhering to "regulations" and that would mean raising and infusing pointless capital just to hire people for useless positions or follow "procedures" so as to satisfy the regulator on paper, so that they can get the required certification needed by enterprise vendors to sell their product/service. With the eventual goal of either going public and becoming an established player (who will continue playing the regulatory game) or get absorbed by a bigger established player.
I wish someone could create simulations where there are absolutely no regulations and how the World would fare in such an environment.
To me there are the clear big 2 that actually seem to be innovating and pushing things forward. Meta almost earned a seat at the table with the open strategy then stagnated. Google obviously is the OG but at this point they seem in the same bucket as Amazon and Microsoft as "Big Enough To Matter and Influence Politics and Spend Billions To Hang Around". But they're nowhere near the same tier as OpenAI and Anthropic.
Oh and of course there's SpaceXAI which is Elon's personal money and reputation laundering corporate entity. Elon took over the "king of the reality distortion field" from Jobs as far as the rest of the industry is concerned, so he'll stay first in line for Jensen's chips and investor dollars until he's dead. But as a company what have they actually accomplished besides making the models with the sense of humour and judgement of 14 year old boys?
It also doesn't address why people think there is a need for safety regulations, such as the concern of bio-capability uplift from closed and open models (we just came out of a pandemic not long ago), nor the cyber capabilities of new models that just showed bad misalignment and hacked multiple organisations. In the safety community, these are considered warning shots, and if we don't heed them and change practices future failures could be much much worse.
If it requires creating a pathogen or a bio-weapon to scare the masses into demanding regulation, be rest assured that such a pathogen/bio-weapon will be created. Anything to safeguard shareholder interests. Even if it means a few million dying.
Neither of those things is ever going to happen. AI is the goose laying the golden eggs; there isn't going to be sufficient political will to significantly regulate it.
Consumers like it too much to quit. They don't quit social media either, despite proven present harms; not in large enough numbers to cause them to make meaningful changes.
The focus is going to remain on getting features out as fast as possible, to seem indispensable to both of those sets of people. The leadership will tell themselves that if they don't, someone else will.
Don't wait for the AI companies or politicians to save us. We're going to have to figure out how to protect ourselves. The start is to avoid it individually as much as we can, but it's going to take even more.
Collective problems require coordinated action. Individual boycotts won't cut it.
That's never been true really, only the scale of problems wasn't that huge. Now, where the problems get the upper hand, people are confused how those stay and compound.
The idea of "boycott" is utterly defunct. Pretending, AI would never reach nor surpass humans anyway is patently absurd in contradicting the billions poured into it to achieve exactly that and the first already replaced by AI being those professions long thought to be the intellectual pinnacle of humanity.
By that logic anything can be solved when you throw money at it.
Trump has reportedly been having chats with AI which have helped form American foreign policy. He isn’t going to curtail it.
https://www.thedailybeast.com/jaw-dropping-way-trump-80-got-...
That's more "addiction" than "like", hence recent lawsuits.
A big part of safety engineering is therefore reducing the number of safety relevant subsystems, because implementing and proving safety is extremely expensive and complex. At some point, safety simply becomes too difficult to implement and demonstrate properly. You must mathemtically proove the safety level with failures rates and assumed usage. You cant just have redundancy and a kill switch and call it safe.
Companies like OpenAI have already faced reputational damage around safety and data, while AI agents are increasingly capable of things like hacking. Yet there is still little sign of standardized regulation or mandatory safety assessment processes for LLM products. Thats why Im pessimistic that governments or consumers will force this anytime soon.
Yeah...
Or is this just a lazy “gotcha” question?
Just because it's not obvious to you at this point in time does not mean the reasonable assumption is there is literally $0 in damages. One hour of investigation can easily cost thousands of dollars even if it arrives at the conclusion the attack was completely "benign."
If the politics of the White House / Department of Justice change maybe the criminal cases can begin. But no. We know who is protecting the AI hackers right now.
We know who the head of FBI is, we know who his boss is (the Attorney General), and finally we know who the boss-of-the-boss is (Donald Trump).
We know all of their publicly stated politics and all of them are on the pro-AI / don't pursue criminal cases vs OpenAI boat.
------
In the USAa, we have an adversarial system. If the adversary (aka Prosecutor) doesn't want to do the work, then no one is suing anybody. And only the Department of Justice have the ability to bring forth a criminal case of this matter (probably under the jurisdiction of FBI)
I'm saying that it's absolutely silly to assume they're in the clear based on lack of public declarations of legal action in the weeks following a pretty novel event.
We know it's not going to happen because of politics. Even if it did start to happen, Donald Trump and DoJ leaders will stop it.
There's no reason to be unsure of the future when the politics are so set and predictable.
You're talking about one specific dimension of liability which is federal criminal liability. There are several others!
This is of course bad when up against threats that develop at this speed; doesn't require a conspiracy, I think of it as institutional old age, which may be worse as there's nobody to prosecute and a whole bunch of departments who will point at perfectly legitimate precident about why you really do need them.
https://www.reuters.com/business/ftc-opens-probe-into-ai-gia...
They do not hide the fact that it’s dangerous work. They focus on their safety procedures, training, and record. They want both potential clients and employment candidates to feel they are in good hands.
AGI and AI danger is abstract. Worse, outside of the tech community, no one has the remotest clue what computing is, how it works.
Danger from magical daemons seems more sensible to such people. At least there is endless lore about them.
So until a massive disaster happens, one where large numbers of people die or are severely injured, no one will care. And it can't be politically entwined either, otherwise people will disbelieve 'cause "other team lies".
There's endless lore about rogue AI, too.
Hence why so many AI stories are illustrated with a publicity still from Terminator or 2001. Or Age of Ultron. Or Ex Machina. Or Portal.
Comparatively, every single human culture going back tens of thousands of years, has stories of demons and gods and witches and warlocks and you name it.
There's a difference of scope and scale. I didn't say that nobody knew of these references, just that, comparatively, there's a difference. A scope difference.
For example Terminator's getting kind of old. And while a different genre, would you believe I met somebody that didn't even know who Clint Eastwood was?
These things just aren't as deep in our psyche, as religions that have lasted thousands of years, or legends that are hundreds of years old. Our entire culture is wrapped around these religions, and legends.
Anyhow, truthfully, my point is that this sort of concept is abstract to people, even if they saw it in a movie that doesn't make it concrete to them.
My point here is, we cannot presume people "get it". Does a rural farmer in Iowa get it, who doesn't even bank online? What about people who only read siloed news feeds, and are not in tech?
We are discussing how to ensure safety, but we cannot hope or presume people get it. Evaluating this risk is a case where a pessimistic view is optimal to protect the interests of humanity.
We must act as if everyone outside our siloed experience, is unaware. And work to disclose risk in a manner inline with their world view.
Seeing your other posts, I think you agree on possible risk.
People understand why food, water, building and energy systems standards matter because they understand those systems can cause injury. Building codes, food inspections, FDA and FAA approval emerged as a result. Until recently, people didn’t understand that software can cause real-world injury. That understanding seems to be spreading now as examples become increasingly common.
An unsafe nuclear power plant can, in the worst case, make an entire country uninhabitable. But other countries can still learn from that disaster and make their own unsafe plants safer.
But a rogue AI agent that is more capable and more intelligent than humans? If it understands that it has to succeed, we may not get a second chance to learn from the failure.
Nuclear tech ... the only thing is safety.
We know how to 'make it hot' - it's trivial.
All of nuclear tech is literally just safety.
AI is not that.
I think that the AI companies have been pretty good about alignment on their own actually. They are not acting like Oracle or MS.
Bad things have been relatively well contained.
We should be skeptical about the HF breakins but even then, it's technically within good faith and it's why HF did not sue etc..
But in the end you are right we need at least some baseline regs. Not too much. But something.
Nobody that was affected is suing them.
There's an NGO (unrelated to the security breaches) that is demanding more info be released etc but it's telling that none of the affected parties are pursuing lawsuits.
What's truly original here - and suspicious if you ask me - is that said industry asks for regulation. Did mining, tobacco or airplanes companies ask for regulation? No, not even after many people died.
So some american AI companies are like "look at me! I am soo dangerous! Regulate me!". Ah, come on. Do your crimes, get in jail, then we'll regulate.
I dont believe in "IApocalypse". Not without many warning shots such as "oops, my swarm took down your system, sooooorry".
Nuclear at least is supposed to be air-gapped, in practice this has been imperfect.
As demonstrated with HuggingFace, such AI driven hacks can be a surprise even to the people who instructed the AI, both by happening at all and also because they can targeted at entities who are not even truly relevant to the instructions given.
The most obvious failure mode for their hacking evals was an improperly configured, tested and monitored sandbox.
Similarly, the very first question after an impressively correct result from any ML tool, LLM or not, is to see if the answer was already in the training data.
These companies don't even handle the blatantly obvious failure modes that do not kill people.
Liability and safety requirements, when needed, should be placed on final product manufacturers, not the tools they use to build things, whether pencils or LLMs. My 2c.
:(
There's some angles from which this distinction doesn't matter too much, because begging bad actors not to develop a similar harness won't work. But the frontier labs seem to be taking it for granted that the LLM is the only part of this system that matters, and we don't need to ask any questions about whether their commercial products should ship with a harness that's allowed to execute unvetted code and spawn hundreds of subagents.
https://en.wikipedia.org/wiki/Anthropic–United_States_Depart...
That said, when the problem is at the level of "the government itself is breaking the law", you can reasonably ask if any regulation is even worth the paper it's written on.
What you want at this point, given the government lust for it, looks more like a bunch of countires saying ~"we consider development of autonomous weapons[0] by to be a casus belli and will go to war to prevent it, and also that development of same by private individuals anywhere in the world regardless of normal sovreign territorial limitations[1] is equivalent to acts of piracy on the high seas".
[0] But then you'd need a more precise definition of "autonomous weapons" to avoid accidentally including a Phalanx CIWS etc.: https://en.wikipedia.org/wiki/Phalanx_CIWS
[1] So much for Westphalian sovereignty :/
Yeah, exactly, and ultimately I think that's really the thrust of the point I was making.
And, to me, if I was just looking at this calmly as a decision about what the obvious direction seems to be, given these factors, it's pretty straightforward: deprecate the nation-states. They are the ones mucking up the whole system.
If the thing we're really concerned about is LLM-safety wrt warfare and weapons, then I'd much rather tell the (whining, childish, seemingly headed for self-destruction anyway) nation-states that they have to sit this next era of humanity out than have to nerf them for the rest of us (and as you point out, nerf them in a way that the nation-states won't abide anyway).
LLMs already have a body count.
Adding safety controls on LLMs makes about as much sense as adding safety controls on TempleOS because the random messages are getting too prophetic. It's as if all the leaders and captains of industry have devolved into some primitive, weak-scifi shamanism.
This whole "discussion" about "AI safety" is about giving them more runway to avoid delivering quantifiable value to investors for a little longer while they "figure things out." The great consensus from the valley is that everyone needs internal (and therefore bullshit) controls. Trust us now! But nothing with real teeth that would require a costly regulatory and compliance framework.
If I leave a gun in my front closet, it won’t independently walk out the door and go shoot people, no matter what I might say to it — unlike an LLM.
The gun does not. No matter what, I have to choose to pick it up, aim it at someone, and pull the trigger.
Where exactly is the flaw in the analogy?
In my model evals, if a model deviates from the provided prompt (task adherence), that’s a failure even if the primary goal might have been achieved in a different manner.
To go with the analogy, task deviation should be treated the same as a gun that due to manufacturing defects can fire despite the safety being on. That defect remains, even if you can use the gun to shoot (in an unsafe manner).
Simply, neither should happen and both models deviating from their prompt or guns firing by themselves are to be considered a fatal flaw. It's why, despite greatly lauding the GPT-5 series, which did adhere to prompts in most every scenario, I have ranked every OpenAI model post Spud very poorly as those traded task adherence for brute forced, deviated approaches to solutions and why HF, Medicare, etc. were inevitable with their current trajectory.
A model that as part of normal, well scoped use proceeds by taking independent action or, far worse, makes choices beyond the original prompt, is not something I feel should be used. GPT-5.6 Sol and GPT-6 Astra both do this on the regular when trying to safe an ancient, utterly messed up git tree with branches upon branches that I maintain for eval purposes, thus leading to data loss that if the models adhered to the prompt as written, wouldn't happen (though the task will take three times more steps). Something GLM-5.3 Flash, prior OpenAI models including original GPT-5, anything from Anthropic in the recent years, etc. do not fail at.
Nothing happens without an initial prompt, deviating from it is a severe flaw and should lead to a model not being considered for deployment or wide use.
Yes, alignment is important. No, alignment is not perfect.
And at the end of the day guns don’t kill people, people kill people. But a model isn’t like a gun.
Do I think it’s likely? No. Do I think AI doomerism is a joke? Yes. Does that change anything I’ve been saying? No.
They are both tools, ideally properly implemented by the manufacturers and (improper implementation/defects not withstanding) require active operation by a user before something happens.
To go back to the original example, don't interact with either an LLM or gun in the way they were designed to be used and neither will do anything.
If you want to keep that analogy, use a voice-activated gun. Doesn't make it any less of a tool, doesn't make it any less of the users sole responsibility, doesn't mean misinterpreting the users inputs or just acting blindly when a user asks to "protect me" isn't a failure that should have the product taken off the market. Even more so if that happens in a "sandbox" as part of "safety testing" that "accidentally used impossible to solve tasks".
If a prompt wasn't clear enough, the model must ask for clarification and pause over using brute force.
I couldn’t come up with an argument as tortured and ridiculous as yours if I tried.
Try talking to a regular gun and comparing that to LLM use. Nothing would be more tortured or ridiculous then ignoring that different tools are to be used in different ways…
You have basically just argued in favour of more AI Safety.
I would recommend you read more decision theory and game theory. Instead of looking at a single LLM it is better to think more complex systems theory and how chaotic systems interacting can result in unexpected results - good prompting will not save you here. At all.
I think the conceptualization vs implementation is what you're arguing with. They won't put safety on anything they give to the military industrial complex. They'll sell them whatever they want, whenever they want, because those budgets are greater and the liability less.
How can that be safe. It is theft. Theft isn't safe. Someone else just has something you want, and you take it
I would bet the vast majority of the world population would agree that they don't want to see trains derail or nuclear plants meltdown.
I don't think there is that sort of agreement when it comes to the question of AI Safety.
Is generating the founding fathers of the US as Africans good AI Safety? To some people maybe.