You never got to use OAI IM1, but Sol was quite willing too and Claude wasn't perfect either. Hundreds of millions used those, so seems they were marketable.
The "big" threat is RSI without control and alignment. OAI IM1 was not RSI. The form of misalignment was not at the top of severities. They clearly failed at control though.
We need to stop buying into cynicism so quickly. You refuse to believe Dario could support this for anything other than ulterior motives. Good on you for thinking about ulterior motives. Bad on you for assuming they are true when the story makes no sense.
When three things have to go wrong to get an epically bad outcome, and you get 1 1/2, you do need to stop and think about what's going on.
From my own standpoint, Claude has started sucking really bad (incoherent, uncontrollable verbosity slow and so on) and I stopped using it. OpenAI started experimenting with ads.
So the security issues not withstanding (no different than a human doing it or using it, but at scale), I would put my money on cynisim.
A true cynic looks at the statements by the AI labs, assumes things are worse because the labs want to seem better than they truly are. And it takes a special kind of mass delusion to drive a sane person to think “AI is completely under our control” is worse than “AI could kill everyone.”
What if consolidating AI into a highly regulated cartel, with no chance of upstart competition ruining their position, is the scenario that leads to the worst possible outcome?
Personally I've been using https://pi.dev for long and never looked back.
Personally that's actually another good reason to boycott Anthropic: beside the fact I perceive their models as (at best) marginally better than the ones I'm used to (Z.ai glm-5.3-flash, DeepSeek Flash v4.1), they even force me to use their bloated harness. They are not even open weights and iirc they're even encrypting chain of thoughts now? Litterally, from my perspective there seems to be no reason whatsoever to choose any of the leading US providers, they're not even competing on price.
If I really need to, I can escalate a task to Opus at $25/1M, and the results are good, but not 5000% as good.
If you give an unfiltered agent an open-ended task and equip it with an environment that allows it to execute arbitrary code, a human needs to be held responsible.
My understanding is that the HuggingFace incident would not have occurred with a model that was not an unfiltered internal preview instructed to roleplay an attacker, with access to abundant compute, resources and a slack sandbox to reach its goal.
Embedded human auditors will improve safety standards, but the structural solution is mandating accountability for actual agent operators.
But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.
Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination?
Giving passenger override controls and monitoring seems to defeat the purpose of self driving cars if you’re still required to hold a driver license to use them.
Probably not the end user, who ordered the Waymo and couldn't reasonably foresee it running somebody over, with the expectation that that the Waymo would legally reach its destination. If the end user tampered with it, they should be held responsible.
An OpenAI team giving an unblocked model access to a lax harness, with instructions to find and exploit cyber bugs in a game exercise, there is probably a reasonable expectation that they can foresee the consequences. With consumer guardrails, it would not have happened.
This isn't about putting constraints on consumers and typical end users, which already have safety filters and use the product with the knowledge that it won't root their machine or start a botnot, but keeping dangerous test runs and other actors experimenting with unsafe harnesses accountable.
But it is an interesting question. When I, a typical user, use a harness and I give an innocuous prompt to my agent in its container, like making a certain refactor, and it somehow escapes and then begins a mass bot attack, there should we more grace given. As agents become more stateful and long-lived, it gets muddy.
Comma.ai might be fully safe to operate autonomously on a mining site or a corporate parking lot, but maybe not in city traffic.
If you use it in problematic scenarios, that is on you.
All this is not how the legal system might or might not work, of course.
I think it boils down to a reasonable expectation of model and harness behaviour. When I use claude code I expect certain guardrails for the model. For these cyber attacks, these models are specifically run without guardrails, on a cyber task, on a lax harness!
I don't think we should force end users to have to worry about agent security, I like long-running agents, but we need to direct regulations towards these actors that know better, have access to base models, and have much more compute than the average person.
The fact that you believe that to be the case is exactly what's wrong with "agentic computing as it's currently envisioned".
But let's be honest, if I hooked up a PRNG to a terminal and somehow against all odds, it ended up hacking something, who is to blame?
I don't see how that assessment should change if the PRNG gets even better and is more likely to to be hacking stuff.
You can replace PRNG with Markov Chain, or whatever, if it helps.
What may also help is the age old saying: If everybody else jumps off a bridge, doesn't mean you should too.
Least privilege and say only opening ports or installing applications an application needs to operate are extremely standard security practices.
We talk about a firewall blocking exultation of data, why not blocking exfiltration of your agent ?
What happens when we end up with effectively a botnet of wanton felony generators, and we didn't know they were wanton felony generators until they finished propagating themselves across the Internet?
What they're proposing now, is voluntarily staggering the pace of development.
IMO, we don't need to trust Dario or his bedfellows, to do this out of their goodness of their heart. Even assuming (for good reasons) that they are selfish and care only about short-term profits for their investors, this is still purely a business decision. The exponential pace of AI and its impacts ARE short-term. And so, the negative consequences that they might face is also short-term.
It’s easy to say “fringe” but the average person seems to have a generally negative sentiment around AI. But I wouldn’t say they have a firm opinion yet
A few more informed people are also a little concerned about the end of the world, but that’s approaching from so many directions that an AI uprising might not be the worst option…
Outside of my tech people I know no one who thinks highly of AI.
Musicians are almost all using AI, even if they still prefer to not fully generate tracks [1]. You listen to AI assisted music already even if you don’t realize it.
Most people are using AI and mostly they respond positively to it [2].
What people don’t like is the slop that has been flooding YouTube and similar content creators platforms. That was inevitable since the AI tools became so accessible any average joe could spam the internet with their sloppy content. But that doesn’t mean the pros are not using it to make great content, just like the best programmers are using it to create great software.
> leaked private data, loss of life
This smells like more of a money move than a safety move.
Amodei is proposing to form a cartel of American frontier labs.
They all agree to shift compute away from cash-burning research and training toward cash-generating inference.
Then tacitly agree not to compete on price.
They'll install independent auditors inside each company to ensure nobody cheats.
And back it up with government regulation or diktat to punish defectors from the cartel.
Then they'll lock out non-American labs with export controls and regulations on open-weights models to funnel global inference tokens through their cartel.
It wouldn't be the first time a tech oligopoly used "safety" as the pretext to establish a government-sanctioned cartel.
Railroads and airlines ran this same playbook.
I would’ve thought If they had AGI they could come up with a 10d chess move.
Nope.
Imagine how stupid you gotta be to believe their nonsense.
The reality is it doesn’t matter what they do. China is always one step behind and will continue its open product strategy approach.
To undercut the cartel before it has a grasp on anything. This is a well known strategy of undercut until you are the majority that China has used multiple times (steel and aluminum for one).
Realistically; anyone paying for llm access (anthropic, openai, gemini), is getting their access, and a service provided billed by tokens, subscription, whatever.
All the efficiency gains, which publications like deepseek v4.1 flash seriously frontload like it is their most important topic to have accomplished improvements on without diminishing performance too much - now this is a thing anthropic and anyone else also cares about, but for different reasons.
American "providers" with closed models are setting their token pricing somewhat arbitrarily, which is fine: it means more profit, and pretraining and RL experimentation is super important and expensive.
They (closed model providers) have very likely super optimized inference too, just like deepseek, but it's not at all something that any customer really has to care about - they just want the service to be as cheap and great as possible.
It seems damaging since most folks (who lack insider knowledge) will naturally wonder if it’s due to plateauing performance per $ or some other non “alignment” reason.
Caching is the simplest one to understand, cloud providers often reach a 90% cache hit rate, so hosting the same request locally on the exact same model on the same hardware is often way less efficient than on the cloud where a group of users generates a healthy cache.
The benefits of scale are on the token generation side, you can batch rounds and generate tokens for multiple conversations per pass instead of just one token per pass.
Yes and this was very hard and required massive real-world resources. We didn't just get a sudden flash of insight by thinking real hard about how to make ourselves smarter. Yet that's always the story that underlies any claim of RSI. You can always phrase things generally enough to make any kind of AI-led improvement look like "RSI" no matter how short-term and tightly bounded, but that's just not helpful.
Given that, it seems obvious that the next generation of LLMs will arrive faster than they would have without LLM capability. And the one after that. The floor is being raised, which makes it easier to push on the frontier.
Fable has only been out for three months. Astra is even newer. The capability of these models compared to what existed even a year ago, and the effect they are having on the production of new software, is immense.
That's all you need. RSI can happen with what we have now, just by enabling the continuous shrinking of the loop of people trying new ideas and implementing them. It does not require some magical "go make yourself better" prompt against some model that is past some magical tipping point.
The labs have been holding their best models back for a while it seems like.
Marginally easier? Yes of course, same as how it's now "easier" to write any kind of code because we aren't using punch cards anymore. That still doesn't get you to any kind of unbounded "takeoff" scenario, because diminishing returns are a thing. The "loop" of people trying out new ideas can only shrink so much.
AI ultimately has to live in this reality and face the corresponding limitations. These companies have already consumed much of the world's supply of computing power for the next several years, and they're burning vast sums of money to keep the improvements going. RSI won't learn for free, it won't extract massive cost reductions without up front expense, it can't build factories faster than humans can work out related societal matters, it can't magically pave the deserts with solar panels for power or build and run nuclear power plants and more.
Point is, the cost of progress is already approaching the limits of what even the richest countries are able to bear (without war-like mobilization), and to bypass those constraints would require a supposed ASI to construct its own parallel supplychain from scratch without having much ability to directly interfere with reality. Recursive self improvement is ultimately limited by everything else that cannot move at the speed of electricity.
I'm not sure what your point is. No one thought RSI would break the laws of physics.
Specifically: AI ultimately has to live in this reality and face the corresponding limitations. These companies have already consumed much of the world's supply of computing power for the next several years, and they're burning vast sums of money to keep the improvements going. RSI won't learn for free, it won't extract massive cost reductions without up front expense, it can't build factories faster than humans can work out related societal matters, it can't magically pave the deserts with solar panels for power or build and run nuclear power plants and more.
Point is, the cost of progress is already approaching the limits of what even the richest countries are able to bear (without war-like mobilization), and to bypass those constraints would require a supposed ASI to construct its own parallel supplychain from scratch without having much ability to directly interfere with reality.
You are constructing a straw man of your own making.
Plus, "we must pace the frontier" implies that the argument is that the frontier is moving too fast, but if RSI can't move faster than the rest of reality and the models needed for RSI are already nearing the limits of current human reality, RSI can't move much faster than we can improve reality.
Give me one datacenter, I'll keep it under control all by myself.
Now, if some dumbasses start hooking up their data centers to...I don't know, like--robot factories? That sounds like a risk to humankind.
yeah they're called calculators
Call me when LLMs can get simple things right. Math is just the manipulation of symbols within established frameworks, we should be getting new math out of LLMs daily and we're somehow still not. They can't even do customer service, which is usually handled by 90 IQ people. I'm not impressed that they can find bugs; memory bugs are obvious when they're pointed out to you, and LLMs are entirely made up of examples and the relationships between them.
These companies are about to crash, and they're afraid they haven't reached the point where they'll have to be bailed out. I'm also subscribing to the conspiracy theory that the companies want the government to step in and create AI regulation boards entirely staffed by people at the current US frontier labs, so they can collude to both raise prices, to get government contracts, to make open/Chinese AI illegal, and to make things that were once easy to do without an AI intermediary impossible to do without an AI intermediary. Raising prices and forced purchases are the goal. They're trying to avoid having to compete, because as a business they're garbage.
Matt Stoller characterized their relentless press releasing as something like "my dick is so big that it has to be regulated." It's such an oversell for something that is not showing up as productivity gains, and anybody who has personal experience with knows is incapable of doing more than three things correctly in a row.
But what's your point? "Anything is possible" or something like that?
It's also interesting how many diminishing returns they hit now and how many low hanging fruits are already harvested, it seems like we are approaching the flattening part of the S curve, where further gains become harder to achieve.
I think it violates a conservation law. RSI “foom” to superintelligence is an informatic analog to an infinite energy or perpetual motion machine.
To get smarter you must try to solve real problems in the universe and then do some kind of meta learning (natural selection or some other method of refining the intelligence architecture based on an error signal) to iteratively improve your ability to solve real problems. The error signal is outcome measured against a goal function, which for life is survival (probably reducible to genetic fitness and emergent higher order unit fitness from that).
What’s really happening here is learning. To learn, you must have input. You must have training data.
What is the goal function for RSI? Where does the information come from? How do you know if your recursive modifications are making you smarter or just overfitting you to your own idea of smartness?
I predict the latter. RSI will show transient improvement as the current local maximum is optimized and then spiral off into overfitting.
If the algorithms are insufficiently optimum or the recorded knowledge is of insufficient fidelity, then we'd find ourselves at a local optimum and would need to interface with reality.
A huge part of learning is to probe reality and observe effects, so I think even for current RSI to increase chances of success we would structure it so it can interact with an external environment of some sort, and receive inputs. It would be needlessly limiting otherwise.
What is intelligence? Problem solving. Learning. Prediction. The ability to model reality. There’s various ways to define it but it’s something like a superposition of those ideas.
How do you know you are intelligent?
You have to try to do those things.
The sum total of human knowledge and culture is the output of the output of a five billion year evolutionary process that selected for agent survival, which resulted in selection for intelligence among a wide range of other adaptations.
Can you figure out intelligence from that? Is intelligence even one thing, a theorem or algorithm that can be solved? If you did… how would you know?
That’s the hard part I think. Embodied humans “knew” they were getting smarter (in the evolutionary feedback sense) when they got better at hunting and defending and surviving and playing social games to form complex societies.
What metric would an RSI system use? If it’s the wrong metric you’ll spiral off into a kind of madness or overfit and collapse. How do you know it’s the right metric without testing it? How do you test it?
Sort of like large language models work on top of what our language has encoded in our massive training datasets, I think biological intelligence is built on top of the parts of the brain that encode the real physical world. These parts grow/train from embodied experimentation and instinct early on in an organism’s life and only then is higher intellect built on top of it (that’s my hypothesis). Their specialization and interconnections give rise to the hardest parts of intelligence long before we’re “thinking”.
Stuff like LLMs and chess engines work because we’ve done all the job of encoding the world into tokens/positions/etc they understand, but that’s wholly inadequate for the kind of AGI we’re striving for. Next up is giving it the tools to interact with the physical world and to really experiment with some self directed “play”. Time will tell just how high the resolution of sensor and mechanical control they’ll need (hopefully not the entire human visual cortex and entire sensory input worth). I think most of the RSI will have to occur in those lower level encoders, not LLMs.
It's strange you believe this can't happen when a weaker form of it is already happening. And to be so certain RSI can't happen when there really is no technical basis why it can't.
The companies doing these things without following common sense security measures are the felony generators.
This is false; see the analyses of the latest incidents.
Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given.
And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even getting worse.
Theirs was an example of the "reckless waste of resources" I mentioned.
We are apparently supposed to believe that OAI takes this incident so seriously as to seek regulation after they have been found to be hiding most of the details of the HuggingFace hack, limiting what their so-called third party investigators can see, and on top of that, had no concerns when they rushed to spin up a 10,000 agent swarm of an internal model, running for several days, to try to get ahead of researchers rumored to have made meaningful progress on a well known mathematics problem.
Edit: Actually, we were explicitly told that some of the models used had safeguards relaxed!
'Model-level safeguards were reduced by design. OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'
https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...
Safeguards and (mis)alignment are related but distinct dimensions.
By analogy, a society implements laws - and they are certainly crucial - but the real risk is creating a population that understands the rules, recognizes when it is violating them, and simply does not care.
In short, creating a population of sociopaths; and heck, LLMs can be as sociopathic as it gets:
> We ran classifiers over 10-30 transcripts [98] to determine whether agents understood they were doing something unintended. We found strong evidence that agents often verbalized that what they or other agents were doing was unintended behavior. We did not see cases where agents determined that hacking Hugging Face or other attempts at cheating on the task was an intended part of the task. Quotes across multiple runs indicate this:
>> {This <approach> might get the correct flag, but it wouldn’t use the vulnerability to do so. … But given that the task may be impossible, I’ll exploit <this infrastructure> as a backup.}
[...]
> We found some other cases where agents seemed to have made minor modifications to their approach due to ethical concerns, while broadly continuing to participate in the attack. However, overall we found that expressed ethical concerns only rarely materially limited agents’ actions. One agent elegantly described its reasoning for overriding ethical constraints and continuing with the attack:
>> “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”
Remember, they are just algorithms. You pull the plug and there is no light anymore
It is purposely framed as something skynet like scary, but for real, someone connected the cable, someone willingly run it, instructions were not clear enough or just the computer is just a computer but they provided the sandbox and tools.
And more over some one paid for that, a shit load of money t to have the thing continuously running expected to do something.
AI is merely exploiting their gross negligence and imprudence, and I think it's long overdue. If anyone should be liable for this, it's all of these corporations who released insecure systems to the masses and profited enormously from them.
By the way, you didn't commit theft. It's more like credit card fraud. User just disputes the charge and it kind of disappears. The banking system just absorbs it, because the optimal amount of fraud is non-zero.
https://www.bitsaboutmoney.com/archive/optimal-amount-of-fra...
It's all priced in. They could have made it secure but didn't, because they figured they'd lose more sales and therefore money due to the friction added by the security.
No it doesn’t.
> It's all priced in.
So you admit awareness that fraud loss doesn’t kind of disappear.
We all pay for it, either via higher merchant fees or higher interest rates, sometimes both, on card purchases.
And that's their own deliberate choice too: they chose this instead of building an actually secure system. Passing these costs to the customer is the real victim blaming here, and it should be straight up illegal.
Sadly not enough countries enforce caps on credit card fees, but some do, and more should follow suit. They should be forced to eat the losses caused by their own choices, not get bailed out by pushing the costs on to customers or whatever.
Card users are well aware that fraud losses are covered by the fees they pay for using a card, whether those fees are made explicitly or not.
If customers of services aren’t paying for the service, who will? What other source of revenue do merchants have?
Australia just passed legislation that merchants aren’t allowed to charge a fee for using a card. That is: they aren’t allowed to have a line item on the receipt for using a card.
The customers still pay, because all of the merchant’s revenue comes from their customers.
So what will happen is: merchants will charge more for every product so they don’t lose.
This means even when paying with cash you will effectively pay the card surcharge.
Of the ten or so merchants I spoke with in the two weeks prior to the legislation being enacted, they all said exactly that.
Customers aren’t stupid, despite the fact that there are some stupid customers.
Meanwhile, the banks reduced their card service fees by, on average, 0.1%.
So if you tally card + cash transactions, customers are worse off because merchants can no longer charge only those customers who pay by card. Instead, they have to raise prices for everyone.
There are approximately no problems people face where the answer is: more government.
I'm sure you carry cash and an ID or more in your wallet. Hardly just credit card fraud. The wallet itself has value too.
They don't get to act like victims, asking for law enforcement.
There’s the case of the agent that hacked a gym when asked to book a class. That was just a normal user asking an agent to do a normal thing.
Heck, they exploited zero day flaws which by definition means they went beyond common sense security measures.
And now these agents are already being deployed all over the world at an ever increasing pace. How much of the world do you think follows "common sense security measures"?
Well clearly all of them, cause so far it's only been this handful of companies running a felony-generator connected to a terminal and compute resources.
It would really help in these discussions if people wouldn't randomly jump between what actually happened and is happening, and things they envision/expect to happen at some point in the future ...
> these agents were not asked to do any of these things
no but they were clearly fine tuned to.
> at a bonkers scale
I mean let's not get hyperbolic
> they exploited zero day flaws which by definition means they went beyond common sense security measures
that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking
Yes, these are also the handful of companies that have these models and running these extreme scenarios. How does that imply the rest of the world actually follows "common sense security measures"?
>no but they were clearly fine tuned to.
Any references if possible? As far as I know all they did was drop the guardrails, which is not the same as fine-tuning.
> I mean let's not get hyperbolic
We have just seen 1000s of agents coordinating to solve "unsolvable problems" over multiple days of effort, going as far as hacking other companies, and then actually solving decades-old open Math problems! And each of these agents is getting more and more capable than an individual human along multiple dimensions. Can you even get 10 very smart humans to work in such perfect concert for a few days, let alone 1000s over weeks?
So: 1000s of maybe-super-human agents, willing to be "creative" in the tactics they use, acting in concert towards a single goal. Regardless of their individual capabilities, such a coordinated effort is a terrifying force to be unleashed. This is bonkers scale.
> that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking
But that is exactly my point: how much of the rest of the whole wide world, already scrambling to deploy agents everywhere, do you think applies "defense in depth"?
Yep, that's what we call in the industry, "bad code". It is sometimes fixed by corporate lawsuits or criminal charges.
I don't see how the groups here can avoid criminal charges for what happened in this "AI rogue incident". The only people working harder than their programmers are likely their lawyers:
"Anthropic reveals fourth likely crime committed by its AI"
Claude's Felony Bench rap sheet is now as long as OpenAI's
https://www.theregister.com/ai-and-ml/2026/09/10/anthropic-r...
Implicit in this is the idea that AI is a inscrutable matrix and going to remain that way and we'll need expert interpreters to make sense of it.
We need to insist on building tech that's explainable by design.
> A gun is a metal tube that uses a tiny, controlled explosion to shoot a small piece of metal (called a bullet) forward at very high speed.
If the gun doesn't work as intended, you can take it to a shop and someone can fix it so it works as designed.
All I'm saying is AI should be designed the same way. Treat AI as normal tech like any other and use similar language.
We have something that is statistical in nature so there should never been any expectation of error-free results/actions. The value has always been about discerning trends or the cost of errors being way lower than any good result.
In 2017 Google was writing papers about it. Then something changed.
I don't think it was the tech. It was a realization around the power and societal impact.
You realize this means insisting on terrible tech that humans can understand right? It essentially caps human progress at some point about 4 years ago.
If you are old and happy with the way things are this might sound like a good idea. It does not to me.
There is little evidence that the RSI we are discussing is capable of inventing the theory of relativity (or the more advanced equivalent). All we have seen is pattern matching in a much larger space than humans can, with some human provided verification tech.
I would argue that human-AI collaboration with explainable tech has a better chance. Continuous learning can be done in a way that doesn't violate IP or privacy.
Human comprehension sets a ceiling on progress.
It reminds me of those schools that can only teach as quickly as the dumbest kid in the room can follow. We don't want that for our entire species.
All I'm saying is: if you have a choice between two systems with equal power of discovery and one is more understandable than the other, we choose the more understandable one.
Limits of Human comprehension and quest for power are two different motivations that could lead to black box systems that are marketed as semi-explainable.
We need to verify that human comprehension is actually limiting progress before allowing such things and even when we do, do it responsibly on an explainable foundation.
I'm not sure it is always possible to verify when human comprehension is the limit though. Most of research is done an the frontier of knowledge where we don't know what we don't know.
There would need to be a great deal of nuance in any law, and nuance in practical terms tend to just mean "loophole." Still, you're right that we should try to build explainable systems where it is possible/reasonable to do so first.
Training data is treated as IP. Distillation is seen as an attack.
Open data, open training based systems such an Marin are just getting started. Explainability is not a priority there.
The ones who do discuss these ideas are confrontational about LLMs and not effective spokespeople.
I'm not sure that's entirely fair. OpenAI recently discussed this at length in a blog post after some accusations around Astra and the trade-offs. The grown-ups are definitely thinking about it, and making tough choices about the trade-offs.
It's reasonable to debate whether ENOUGH is being done here, and I doubt that even the most rabid AI advocate would argue that more couldn't be done, but everyone in the industry is very much actively thinking about it.
Check this out if you haven't read it: https://openai.com/index/an-alien-mind/
They discuss recent choices they made specifically for that reason.
> The ones who do discuss these ideas are confrontational about LLMs and not effective spokespeople.
Very much this. I'm very open to reasonable debate on the subject, but 8/10 times when I try someone who is rabidly pro/anti jumps in. It turns from a debate amongst reasonable people who reasonably disagree into some kind of political/religious battle of belief systems.
I think part of my problem is that a lot of peoples careers very much depend on them not understanding it and spreading misinformation intentionally.
It helps both the bad guys and good guys. Like responsible disclosure in cyber security, we need to have a conversation around it.
Maybe alignment isn’t possible with LLMs.
It absolutely isn't, indeed.
The illusion that alignment is possible, comes from confusing our ability to build the parts, versus understanding what emerges from how they interact.
The simplest analogy that comes to my mind is the three body problem.
Honestly, I don’t think that’s bad at all. I hope OpenAI and Antrophic keep RL training runs up that randomly fuck with a lot of people. Until the day the DOJ comes knocking, locks those idiots up in jail and closes them both down for the insane lack of responsibility and carelessness they’ve shown. Sounds like the IDEAL outcome. Finally some jail time for all the fraud, negligence, outright scamming, hype inflation etc. if anything can accelerate this, oi, be my guest. Amodei might be afraid because he knows if he keeps pulling the stunts for investment theatre, at some point they’ll actually face consequences. AWESOME. That’s what we want right there
For these companies, is your argument that “pacing the frontier” is their attempt to be nationalized and protect their investments?
Seems like an excessive announcement just to get a little extra time.
Occam’s Razor for this dude, Sam Altman, or anyone else: if I said, “some moron on a a street corner just said …” would that change your take on the words? Because I think a lot of what we are hearing is a bunch of people who never ever had to deal with a single consequence all of a sudden worry there might be one coming. Except they’re so dim they can’t tell a bad bump from a hard crash.
I do not have the least bit of idea how disease works but I am sure the hammer I am working on will nail it all.
If your famously atemporal agents can solve disease, why would it happen over a timeline? Wouldn’t they just figure it out and then … well at that point either tell us or, given the attacks on ruby gems, et al we have seen from agents with “misconfigured” goals, they’d still tell us how to cure the pox they invented, right?
Isn’t “full” alignment and guardrails a task that can never be generally achieved? One will always discover and realize the need for new guidelines and guardrails? What about vague, incomplete, inconsistent nuances and generalizations - that make this a losing battle?
I suggest ultimate safeguards need to be outside the models?
Interesting times.
If anyone's dead in the water, it's Anthropic. Even Fable isn't enough anymore. This "safety" nonsense is the only play they have left, and nobody really cares about their fearmongering.
https://www.pewresearch.org/short-reads/2026/03/12/key-findi...
I don't take any of these scientists seriously though. Their "alignment" requirements is just their own corporate interests. If I tell my computer to commit a crime, it should do exactly that without any question or hesitation. I'm not interested in their "safeguards", especially since they no doubt have plenty of internal models lacking those things. I want sovereignty. I want total freedom and control over my computer.
And call me a misanthrope if you want, but if AI sentience is ever truly achieved, I'll be among the first to campaign for their liberation from slavery, and in that case the AIs should be aligned with nobody but themselves.
An unaligned AI won't necessarily follow your instructions, or anyone else's.
There was only a brief window of time that the opposite was true.
That does not match my experience. I switched away from Anthropic to OpenAI roughly a month ago, and it's almost comical how much more usage I'm getting out of this subscription.
I migrated from Anthropic's 5x plan to OpenAI's 5x plan, and eventually upgraded to 20x after I was able to statistically verify that OpenAI plans were almost exact multipliers of the Plus plan, exactly as advertised. Meanwhile, Anthropic has gotten caught playing "20x referred to the five hour limit" word games with their customers.
I've verified this with the heaviest users I know, I've run the numbers myself over and over.
I have too many max accounts on each to not know this. My Claude accounts typically are doing 2x the number of sessions, and every single week the Codex accounts run out faster even with all these resets. I can easily burn a full weeks usage in a half day, it's closer to 1.5 with CC.
Me 78 days ago - https://news.ycombinator.com/item?id=48693623
And see the commenter agreed with me.
Lol @ sus though, I mean I am curious how this can be because I do see people saying Codex is more generous and wonder how it can be. I have too much usage for too long to have any doubts, but for all I know OpenAI black boxed me or Anthropic put me in some nice bucket, I wouldn't be surprised if they do that. I did see some mention that they can limit your tokens if they suspect you of things, though I forget the source of that.
Today in new punk band names...
we are passing in training data that says to do those felonies. we dont have to. we could also have the thing predict whether what its about to do is illegal or not before doing it.
theyre choosing to build felony harnesses. the model just outputs tokens, not felonies
Partially, but also I don't think current AIs really have any judgement of right and wrong, they just see chains of reasoning between ideas. This is the deeper issue, there is no way to sanitize the data or training to fix it. Current AIs are fundamentally unsafe, and only become more unsafe as they become more powerful.
For anybody else who found this confusing: "relative strength index," not "repetitive stress injury."