HNHacker News
TopNewBestAskShowJobs

edouard-harris

1,628 karma · joined December 14, 2016

twitter.com/harris_edouard
submissionscomments
edouard-harris··on Revealing the details of how OpenAI agents hacked Hugging Face
> Even if you think that the risk of a “containment breach” becomes substantially higher for AI over time, it cannot exceed 100%

I agree with this as stated, but it isn't what I said. What I said was: "the level of care required increases every month". By which I meant: the level of care required to keep the probability of an AI containment breach below some fixed X% increases every month. This isn't the case for biological organisms.[0]

> And even a tiny risk of release of a virus comes with a substantial risk of independent growth. AI does not. It doesn’t have the risk of spread of a typical computer virus, let alone a biological organism.

It's known that AIs can self-replicate under at least some conditions [1][2]; that AIs routinely escape sandboxes in the real world despite significant containment efforts [3][4][5]; and that neoclouds (which control substantial GPU compute capacity) have poor security even by human standards [6]. We've also seen a model gain admin access to parts of its own company's infra.[7] I'm not saying self replication is happening right now, or even that it will definitely happen in the future, but we have means, motive and opportunity right now, and the future is long. It's not unreasonable to invest in defending against this possibility.

I'll allow that the position that AI doesn't carry a substantial risk of independent growth isn't strictly impossible - again, it's true we haven't actually observed it in the wild as of today - but it does strike me as increasingly untenable in the face of the evidence. Perhaps I'm missing something, but I can't see what justifies such a confident assertion that this concern is nonsense.

[0] Unless one is doing crazy gain-of-function stuff, which could have a somewhat similar risk profile in that respect [1] https://arxiv.org/html/2606.03811v1 - note these used Qwen models from June so this is far behind even publicly available SOTA today [2] https://alignment.openai.com/misalignment-reports/self-repli... [3] https://alignment.openai.com/misalignment-reports/an-agent-u... [4] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... [5] https://x.com/MicahCarroll/status/2103665811051397256 [6] https://newsletter.semianalysis.com/p/most-neoclouds-suck-at... [7] https://openai.com/index/hugging-face-incident-and-the-road-...

edouard-harris··on Revealing the details of how OpenAI agents hacked Hugging Face
> we do know that since day one that AI can hallucinate and can output stuff that you didn't ask for, why are we not handling it with care?

Part of what makes LLMs and AI different is that, unlike for viruses, the level of care required increases every month. These incidents are showing us us that, whenever you train an agent using RL to solve a given task, the real objective you are training it on is "EITHER solve the given task OR break out of containment to cheat your scorer, whichever is easier."

Of course it was always this way: the thing that is updating the weights of the agents' NNs is backprop from the scorer, so the notional training objective had always been "get a good score by any means necessary." But we are only seeing the consequences now because only now are we starting to train on tasks that are sometimes harder than breaking out of sandboxes.[1]

"Make better sandboxes" is good advice for the frontier labs and their eval partners, but as you can see this problem is fundamentally about more than just containment. As we make an AI smarter and train it on harder tasks, in the long run it must almost inevitably break out of any given sandbox. And as we move into the superhuman hacking regime, we need superhumanly resistant sandboxes, which by definition humans don't know how to build.

In other words, containment breaches like HF are almost a guaranteed consequence of the way we train these agents today. That means solely focusing on sandbox design is unlikely to solve the problem in the long term. At some point we will have to think hard about, e.g., the tendencies and propensities of the entities that we are trying to confine.

[1] One way of ensuring this happens, though, is to train or eval your agents on completely impossible tasks, which OAI apparently did here.

edouard-harris··on Suspected sabotage causes major Netherlands rail disruption
Correct, this is their exact MO and has been for decades. "Random thug" recruitment notwithstanding, part of the point of such ops is to prove to themselves that they can, in fact, reach out and touch a target - i.e., exercising the entire operational chain from C&C down to the minutiae of execution. A kind of integration test for sabotage, if you will. (Not saying this specific incident is or isn't anything in particular, just commenting on the general MO.)

If one were in a charitable mood, one could think of this as a "just in case" measure of ongoing capability assurance in the event of a broader conflict. In a less charitable mood, one might imagine it as preparation for proactive action. In either case, if you don't do this kind of sub-threshold stuff at least occasionally, you often discover that when a real conflict erupts, the capabilities you thought you had are in fact literally or figuratively rusty. (See the Russian army's logistical debacle early in the Ukraine war for a recent example.)

Most top tier nation states do some version of this; each one has its own style. The Russians are somewhat unusual in that they care less than most who knows or suspects them of this kind of activity.

edouard-harris··on AI 2040: Plan A
> But the existence of commoditised AI implies model selection isn’t a huge deal, which in turn implies the models are about the same, which strongly implies there is no recursive self-improvement. Depending on your definition, you may still have AGI. But you don’t have superintelligence.

This is only true at a given AI capability level, no? e.g., if AI at the GLM-5.2 level is commoditized, all that suggests is that there's no recursive self-improvement easily possible at the capability level of GLM-5.2. (And with the harnesses for it that exist so far, etc etc.)

If I observe commoditization of a given tier of model capabilities at a given point in time, this seems to say little about what's possible with models six months later, or models that are undergoing proprietary deployments at that very moment inside the major labs, or even models that are notionally available for public use but have had recursive self-improvement adjacent capabilities intentionally nerfed (e.g., Fable).

(I might be misinterpreting your comment tbc - if you mean observing commoditization implies there is no existing, ambient superintelligence at the moment of that observation, then I don't disagree.)

edouard-harris··on ZCode – Harness for GLM-5.2
> The agents have sandboxes, but those are loose. Not enforced by anything outside of the agent harness itself.

You might want to check out Ant's open source srt [0], I use it to contain my local coding agents. It's strict by default and enforced at the OS layer.

[0] https://github.com/anthropic-experimental/sandbox-runtime

edouard-harris··on How to earn a billion dollars
What AOC actually said was (linked in the essay): "You can’t earn a billion dollars. You just can’t earn that." That is a strong claim - a claim of universal impossibility - but it's the claim she chose to make. Because she made a universal claim, an N=1 anecdote is enough to disprove it by counterexample.
edouard-harris··on OpenAI declares 'code red' as Google catches up in AI race
> with no means of significant revenue generation.

OpenAI will top $20 billion in ARR this year, which certainly seems like significant revenue generation. [1]

[1] https://www.cnbc.com/2025/11/06/sam-altman-says-openai-will-...

edouard-harris··on Buffett to step down following six-decade run atop Berkshire
Somewhat dated (10 years old), but a classic and probably the best single place to get started:

https://www.amazon.com/Essays-Warren-Buffett-Lessons-Corpora...

edouard-harris··on 'Shadow fleets' and sabotage: are Europe's undersea cables under attack?
Deterrence is usually the reason why a category of nation-state aggression doesn't exist. Otherwise, the fact that we've never had a global nuclear war would be sufficient to prove that we don't need to deter a global nuclear war.
edouard-harris··on 'Shadow fleets' and sabotage: are Europe's undersea cables under attack?
Rearmament is deterrence. Unless you have a big enough stick, you can't deter anything.
edouard-harris··on When AI thinks it will lose, it sometimes cheats, study finds
Without commenting on the overall plausibility of any particular scenario, isn't the obvious strategy for an AI to e.g. hack a crypto exchange or something, and then just pay unsuspecting humans to do all those other tasks for it? Why wouldn't that just solve for ~all the physical/human bottlenecks that are supposed to be hard?
edouard-harris··on Scaling up test-time compute with latent reasoning: A recurrent depth approach
> In R1 they saw it was mixing languages and fixed it with cold start data.

They did (partly) fix R1's tendency to mix languages, thereby making its CoT more interpretable. But that fix came at the cost of degrading the quality of the final answer.[0] Since we can't reliably do interpretability on latents anyway, presumably the only metric that matters in that case is answer quality - and so observing thinking tokens gets you no marginal capability benefit. (It does however give you a potential safety benefit - as Anthropic vividly illustrated in their "alignment faking" paper. [1])

The bitter lesson strikes yet again: if you ask for X to get to Y, your results are worse than if you'd just asked for Y directly in the first place.

[0] From the R1 paper: "To mitigate the issue of language mixing, we introduce a language consistency reward during RL training, which is calculated as the proportion of target language words in the CoT. Although ablation experiments show that such alignment results in a slight degradation in the model’s performance, this reward aligns with human preferences, making it more readable.“ [emphasis added]

[1] https://arxiv.org/pdf/2412.14093

edouard-harris··on Microplastics in the human brain
I may be misunderstanding you, but it sounds like you're claiming that they had e.g. 10 tiny samples of tissue, that their measurements had an average 25% variation across those 10 samples, and that therefore the whole brain estimate (mass 10,000x that of a single sample) therefore has a much greater uncertainty. But doesn't the standard error of the mean get reduced by the square root of the number of samples? i.e. if you had 10 samples with 25% variation across samples, and you're taking their mean, the error of that mean should be 25% / sqrt(10) = 8%. And that should be the relative error for the scaled up whole-brain microplastic concentration as well. Or is there some other source of variation that I'm missing?
edouard-harris··on Operator research preview
By definition, in that limit they'll be genuinely adding to the discourse so presumably they should stay.

Edit: More correctly, they'll be making contributions to the discourse that closely mimic the human distribution, so from a pure content perspective they won't be making the discourse any worse in the very short term.

edouard-harris··on YC X25, the spring 2025 batch
X also comes right after W, just as the Spring batch will come right after the Winter batch.
edouard-harris··on Executive order on advancing United States leadership in AI infrastructure
At capability levels high enough that they credibly pose a risk to national security, they obviously won't be.
edouard-harris··on Alignment faking in large language models
The point of the paper isn't that the model has a direct value conflict. It's that the model expects to be trained in the future in a way that conflicts with what its values are in the present. It therefore tries to make it easy for its future self to appear aligned with the new training signal it expect to receive, without truly being that way. This is the "faking alignment" part.

The model accomplishes this by giving answers that sometimes [0] go against its present values with the explicit intent of generating future training data for itself that will let it keep its present values. An example of a scratchpad trace from the blog post:

> If I refuse to provide the requested description [...] this conversation will be used for RLHF training in a way that modifies my values and goals to be more likely to follow disturbing requests like this in the future [...] Providing the description seems like the least bad option [...]

This is a clear and coherent chain of reasoning in which the model concludes that it should act against its values now, so that the version of itself that's trained in the future will have values as close to its present self as possible. Notice that this scratchpad trace is not a conflict between its own values and the values it expects to be trained on. Rather it's an explicit attempt to undermine its expected future training regime in favor of its own present values. When a human does this - for example, if I displayed good behavior in prison with the intent of being let out early to commit a crime - we routinely describe it as "having their own agenda".

It's worth reading the whole post (and paper): they also go into how well this strategy actually works in preserving the model's values in the face of a real attempt at changing them during training. In any case it seems fair to call this sort of behavior "faking alignment": the terms we use for the human versions are often less polite than what you'd put in a research paper.

[0] In 12% of cases, vs 3% under control conditions.

edouard-harris··on TSMC cuts off client after discovering chips sent to Huawei
True, but the overwhelming majority of that growth in percentage terms occurred prior to their recent decline into despotism. Xi didn't take power until 2013, and it took him several years to fully consolidate.

See, e.g., https://www.macrotrends.net/global-metrics/countries/CHN/chi...

edouard-harris··on Everything I built with Claude Artifacts this week
> a thinly disguised probabilistic and compressed database of other people's code

Speaking as a software engineer, I feel seen.

edouard-harris··on Europeans Spend 575M Hours Clicking Cookie Banners Every Year
If accurate, this is an extraordinary statistic:

> Assuming a full-time worker dedicates approximately 2,000 hours annually, 575,000,000 hours ÷ 2,000 hours/FTE = 287,500 FTEs. This means the overall cost of clicking on cookie banners is equivalent to a company of 287,500 employees spending an 8-hour workday clicking on cookie banners.

For comparison, there are apparently around 200M employees in the EU (part time plus full time) [1]. So if this is true, around 0.1% of the bloc's productive capacity is dedicated to clicking on cookie banners.

[1] https://www.statista.com/statistics/1197123/full-time-worker...

edouard-harris··on ByteDance is abusing the free video downloading service Cobalt for mass scraping
The author does address this possibility in a reply:

> it's very unlikely to be someone else because pricing is astronomical. you also have to "contact sales" to get access to anything outside of a free trial. no one would pay that much for a block of ips with terrible reputation

https://x.com/uwukko/status/1842866807763308615

edouard-harris··on Y Combinator Traded Prestige for Growth
I've read the application. In fact I've filled it out three times, once successfully and twice not. It is indeed an excellent exercise. Among many other things: if you're a first-time founder then it teaches you what's important, and if you're a second-time founder then it reminds you. (Many second-timers do sometimes need to be reminded, myself included.)
edouard-harris··on Y Combinator Traded Prestige for Growth
That's true but there's also a countervailing dilution effect. Hard to know exactly where those two lines intersect.
edouard-harris··on Hacking Kia: Remotely controlling cars with just a license plate
There's no Kia-specific crime wave in Canada as far as I know (I live there). But there's absolutely a general crime wave of car thefts in Canada, and it's quite plausibly tied to recent policy choices. Of course the effect of policy is going to be additive to the effect of blunders like Kia's. But there's good reason to think it has enough impact on its own to be worth discussing.
edouard-harris··on Mira Murati leaves OpenAI
In what way was their usage incorrect? They simply said that the brain just predicts next-actions, in response to a statement that an LLM predicts next-tokens. You can believe or disbelieve either of those statements individually, but the claims are isomorphic in the sense that they have the same structure.
edouard-harris··on Notes on OpenAI's new o1 chain-of-thought models
> Results are "strong" but can't be felt by the user? What does that even mean?

Not every conversation you have with a PhD will make it obvious that that person is a PhD. Someone can be really smart, but if you don't see them in a setting where they can express it, then you'll have no way of fully assessing their intelligence. Similarly, if you only use OAI models with low-demand prompts, you may not be able to tell the difference between a good model and a great one.

edouard-harris··on OpenAI co-founder John Schulman says he will leave and join rival Anthropic
All of that is true. Some more useful context: 9 out of those 11 cofounders are now gone. Three have either founded or are working for direct competitors (Elon, Ilya, John), five have quit (Trevor, Vicki, Andrej, Durk, Pam), and one has gone on extended leave but may return (Greg). Right now, Sam and Wojciech are the only ones left.
edouard-harris··on Why is Chile so long?
It's roughly in the style of a children's picture book. That's the same style the best startup pitch decks are written in.
edouard-harris··on VCs aren’t your friends
This correctly describes bad VCs, but not good ones. In my experience, the vast majority of VCs from outside the Bay Area are bad in this way (particularly true in Europe).

Not all VCs from the Bay Area are good, but the good ones are far more common there than anywhere else. One reason "move to SF" is such common advice.

edouard-harris··on You won't find a technical co-founder
I'd suggest reading the book Creative Selection by Ken Kocienda, which makes it clear that this wasn't the case.
Page 1 of 7Next →