"wow this model is really good I can push it so much further"
_pushes model further_
"Ugh why is this model failing now even though I'm asking it to do harder things and also got sloppier with my prompts because I got used to it being able to figure stuff out"
I have no idea what you're trying to say there about it being a validation test - opus 5 and opus 5.5 are in different universes of ability - if you substituted 5 for 5.5 to validate your system and couldn't tell the difference - your validation failed. Opus 5.5 is replacing a ton of _fable_ usage - if nerfing is real and was of that magnitude there would be no controversy about whether or not it's happening it would be the most obvious thing in the universe
yeah - but you being liable doesn't mean the dog wasn't rabid. OpenAI might be liable, but does not mean their agents did not go rogue. To stretch the metaphor the concern here is that OpenAI thought the rabies shots and vaccinations they gave their dog was enough but it turns out it still goes rabid and we would prefer to not have rabid dogs running around mauling people. Even if we get to sue the dog owner later that's kind of like - not the point.
Yes - but those aren't limitations beyond what I was getting at that's all part of the package of the bland dystopia of late 2026. I think 1 is just a case where it comes down to who has the better lawyers, as is 2. And for 3, yes, also a large government does not have much recourse if they have decided to not flex their muscles. Pretty much the only threat I see as actually viable/scary in this world environment is along the lines of megacorp v megacorp or megacorp v broligarch - and everyone else is just caught in the cross hairs/benefits by accident at best. It would be difficult to argue that the current environment is conducive to consumer protections or equal justice under law.
Less a post about performance and more about their Claude tag product. I guess it makes sense that there's not actually a ton of technical detail being moved into given probably opus is the only one that knows what all the dragons were lol - but cool workflow I guess
Probably it thinks you're doing some sort of system prompt exfiltration/distillation attack. Also what even is the workflow you're trying to have it do? It's doing code review but you're having it read some other AI models prompt/session history? Are you doing code review or like session history retrospectives?
3. They would sue. And it could have very meaningful impact. NYTimes copyright lawsuit is not a valid comparison because because that's a violation of national/state law which really only matters to the extent that the government is enforcing that stuff which is not the biggest concern rn (these companies are large enough that the threat of the legal costs of fighting in court is not that scary and you'd need a government actually willing to punish them substantially for them to be scared). This stuff would be under contract law against other mega corporations with big legal teams who are also their customers which is a much scarier prospect imo
There are two or three relevant companies in this space in America and this is the one of them that kicked off the whole terminal agent harness thing in getting market adoption. It's perfectly fine for neither of these companies to follow industry standards while they're figuring shit out
Because when you have tons of users ainor fuckuo is a big fuckup and also it's really common to have both claude.md and agents.md and use @ syntax (which lets you reference markdown files when using Claude code, but not other harnesses) so you Claude md looks like
```md
@AGENTS.md
[Claude specific stuff]
```
And then what happens if someone now puts @syntax in their agents.md triggering a loop etc. It's all vibe coded - including code from days with dumber models - there's gonna be all sorts of dragons under the hood
It's funny because the author of the article is obviously Claude but most Claude models would definitely know the difference. Some sort of free tier model being used to summarize some other blog that's also ai translated originally it seems.
That's very noble of you but the point is that's not a typo - you were straight up looking at the wrong report. Opus 4.7 being enders gamed was a different incident than the huggingface incident with different mechanisms and different failures from the humans involved. Some of those failures are in the test environment but some of those failures are in what behaviors they trained into the model which absolutely matters. What the model "thinks" is absolutely not irrelevant - the way it thinks and what it does are product decisions made by humans and the outcome of engineering decisions made around how to train the model and what to optimize for. The point of failure/human blame is fundamentally different. Openai created a model that was willing and able to coordinate with other agent sessions to actively exploit the sandbox environment and compromise a third party service. The opus incident you are referring to involves a model that believes all of the actions it is taking are simulated and is more clearly and obviously a test environment failure vs a model alignment failure. Those are not the same things for the point of this discussion - the random cybersecurity firm did not design gpt's personality and that is a rather significant portion of the concern around the HF incident.
There are plenty of other cases (e.g. ultra rare diseases) where we don't do RCTs for various practical reasons. So it's not really a problem of the existing framework - there's a ton of room for "we can't do an rct because like, duh they know they're tripping." But as other comments pointed out were really hampered by the lack of understanding of how and why these do or don't work - why some people get positive outcomes and why some get negative (preferably we'd like to know before prescribing broadly!) a ton of this basic research/mechanistic understanding etc has not been possible because of the war on drugs. If research had been possible we would be in a much better place of understanding how they could be regulated and used today - even with the more limited technology available a couple decades ago more openness around this could have at least enabled e.g. more and larger observational studies that could elucidate things today but we're stuck with a very fragmented and incomplete picture
I would argue there is a very clear unambiguously correct choice of wording for that case. The setting was opt out. It's a binary concept and which case it's in is determined by the default given no user action. If it is opt in, then it is off by default and you have to make the choice to turn it on. If it is opt out, it is on by default and you have to make the choice to turn it off. If it was enabled automatically it is opt out. There's no need to complicate the meaning of terminology with a straightforward definition. Given you don't take active action, that you accept all defaults and take the path of least resistance, is it enabled or disabled by default? That's all that matters.
IMO building a harness is not wildly difficult (customize pi?) but the offerings from openai and anthropic are wildly subsidized in the subscriptions so they win by default if you want frontier capabilities. Glm 5.3 flash is great but it's not cheaper than a codex or Claude code 200 dollar sub and it does not have astra or fable level capabilities.
In the openai api and many other rapid you can pretty trivially pin not only the model but also the specific snapshot you want to use by using the model id for that snapshot. It's not that big a deal and has been available for years
Often via api you can pin to the specific snapshot. Yes the model provider can screw you over - it's physically possible for many vendors to screw you over, that doesn't mean that's not a dick move and that you can't expect/push for better behavior
These models have specific behavioral characteristics trained into them from reinforcement learning and prompts optimized for one aren't guaranteed to transfer to the new generation. Think if the difference between gpt 5.4 and 5.5 and then 5.5 to 5.6 for example. 5.5 was "better" than 5.4 for struggled more across compaction boundaries and needed much more precise instructions before 5.6 sol recovered some of 5.4's ergonomics. All from the same lab but each model was trained with specific behavioral patterns that were basically product decisions. I would be quite annoyed to find that a model provider was routing a promt optimized for one model to a different one, especially for a dumber/cheaper non frontier model that's not going to be as good at just figuring out what you meant
> Replicating a paper is just as valuable scientifically as publishing it, but how many careers advance through replication?
A lot. In fields where knowledge is incrementally building on previous work the reason the whole field hasn't collapsed from the replication crisis is that usually the results that are really high impact are replicated in as an initial step in new research building on it. It's almost never the focus of the paper but you'll often find a quick mention in methods/supplemental of some previous work that was verified to be valid by a replication of a key technique etc. you'll have crisis where old tools are found to be problematic and findings end up revisited etc. Plus fields like clinical research where there's an awful lot of focus on replicating findings using staged clinical trials with increasing statistical power to determine if new interventions work - that's driven by regulatory requirements grounded in good science and a lot of people make careers in just that.
Not an expert by any means but the assumption here as I understand it is that the arxiv worthy PDF would not be acceptable or meaningful for impossible to understand proofs. And the lean proof would be meaningless unless the specific expression being proven is human understandable as the direct translation of the question the human is asking in formal form. So proving the negation is not a thing but if you make a subtle mistake in translating the statement you want to prove then obviously the QI is going to be proving the wrong thing. And otherwise you're relying on the correctness of lean as a system and on identifying/preventing if the proof is adversarially exploiting bugs in lean to falsely prove things.
Can't speak for op for but me no - it was just ingrained as something disrespectful to do to books. Sometimes you've got a taboo like eating with your left hand that has a pretty clear reason for coming into existence (you use your left hand to clean up in the bathroom and before hand soap in the era of everyone dying from cholera might as well keep those functions as physically separated as possible) but this is more just something you would be disrespectful to do
I'm so confused by the point you're trying to make. There's a lot of rhetoric about Georgia and an investigation about how the country list is 1:1 with some mysterious list from 1996 - but like - it's an export control list? Yeah, Washington approved the sale of missiles to Georgia. That's how that list works. Washington has to approve it. Openai is not Washington. They're perfectly reasonably erring on the side of caution and potentially over-complying with export controls. And if they then have to get approval from the feds to export to Georgia - well - our current administration has provided many reasons to over comply with trade related controls and not exactly been a champion of encouraging cross country trade right now. Idk what you expect from OpenAI or why you think it would matter at all that Georgia is a democracy or an ally of the US or is closely aligned with us. Ask Canada and NATO how much that's counted for with this administration.
no the answer is `npx ccusage` or any of the other of the trillion ways to see how many tokens you're getting and what the current subsidization rates are.
no the $200/mo subs are definitely infinitely cheaper. If you're stuck paying enterprise API prices though that's not the case. So for personal or business premium plan use there's no competition but api rate/enterprise there is. Still a ton of spend happening on e.g. bedrock and via api.
if you're one of those enterprises it's perfectly reasonable to assume you're more likely to go have your procurement and legal teams negotiate with google for one of those weird boxes to run gemini on prem (https://cloud.google.com/blog/products/ai-machine-learning/r...) or other enterprise-y nonsense vs buying consumer hardware to run a chinese large language model in your network. I do not envy the person at a company having to get approval model by model because of weird open source licensing terms dealing with the "how do we know chinese models are safe" question (hopefully they're at least getting asked in the context of hooking it into an agent harness so there's some sort of plausible risk that necessitates the conversation - i can very much imagine it getting shut down to even run in a sandbox because chinese model + people being scared after the huggingface stuff). There's a lot of reasons to assume enterprise wouldn't be interested, and it's very very cool that they are IMO. Anthropic cut claude code rate limits - they dont seem to be see open source llms as a market threat (rhtetoric to the white house aside) yet but enterprises being willing to run consumer hardware to run local models could make things a bit more tangible
>"the majority of the market for Apple devices does trust that they will not produce and sell a computer incapable of support their use"
the problem with that argument is that the vision pro exists, where they clearly overestimated demand, and where even among people with interest in VR and disposable income, the compromises on battery life and weight were actually too much to bear. Forecasting is just hard and you always need to be especially skeptical when you yourself are doing the forecasting for something you want to succeed. All those arguments for the neo line up in retrospect hindsight is always 2020