HNHacker News
TopNewBestAskShowJobs

lebovic

1,931 karma · joined April 2, 2019

Noah Lebovic

Email is my firstname@lastname.com

Formerly at Anthropic, co-founder of Toolchest (a YC startup), and a bio startup

submissionscomments
lebovic··on I'm Eric Ries, author of "The Lean Startup" and new book "Incorruptible" – AMA
Having values doesn't mean they're "the right" values nor the same as your values.

Regardless, it's still atypical in the context of an American company, and it can help explain the differences between Anthropic and its peers. That doesn't mean I agree with their decisions or that they're "the right" decisions, but I think it's a helpful framing in which to understand them.

lebovic··on I'm Eric Ries, author of "The Lean Startup" and new book "Incorruptible" – AMA
Yeah, I lean towards the structure not being the cause of the outcome here (i.e. if you rotated the governance structure of Anthropic and OpenAI, I think the decisions at each would likely stay the same).

If they made that decision and it destroyed revenue, I could see an alternate timeline where a standard C-Corp + board with non-founder control may have ousted leadership. But that wasn't the situation for OpenAI or Google either, and their leadership still made a different decision.

lebovic··on I'm Eric Ries, author of "The Lean Startup" and new book "Incorruptible" – AMA
> If you think that the courage that Dario has regularly shown would be possible with a conventional "best practices" structure, I think you're kidding yourself.

Is there something that happened which you don't think would have come to pass with a standard PBC/C-Corp (without the LTBT)? I'm trying to think of one, but nothing is coming to mind.

I think the structure attracted many people to Anthropic (e.g. an RSP that could only be overridden by the LTBT), but I'm not sure it has demonstrated a practical impact.

As an aside, I think a lot about this problem too! But the answers that don't reduce to something like "the people, and the people to whom they give power" seem to break down when I look closely.

lebovic··on I'm Eric Ries, author of "The Lean Startup" and new book "Incorruptible" – AMA
I worked at Anthropic, and I wouldn't attribute much to the structure itself – so I'm wary of using it as a positive example here.

I do attribute a lot to specific people. Concretely, to much of the intitial team, who they recruited on the research/infra side, and some very close personal relationships within research/infra. That dynamic, paired with their unwillingness to accede to something against their values, is what I credit for some atypical decisions and outcomes [1].

Things regulary go "corrupt" in parts of the company; it's hard to scale without importing culture from big tech. Sometimes, the defense was ICs escalating issues, Dario talking to ICs, and then shaking things up.

But this process takes time, and it doesn't lead to a full reversal; a bad/misaligned hire has reverberating impacts. Many folks are still driven by values (even if their values are not your values!), but scaling dynamics seem to be evolving like any other org – just at a higher employee count and revenue numbers.

I do place trust in specific people who work at Anthropic, but I wouldn't place trust in Anthropic the organization. It's an organization that's wont to change, regardless of its structure.

[1]: https://news.ycombinator.com/item?id=47174423

lebovic··on Claude Fable 5
Not sure this holds, sadly. I spent a few months reporting serious security bugs as model capabilities took off earlier this year, and only ~half were fixed. The unfixed bugs were just as critical as the fixed ones; sometimes they were even two similarly critical bugs at the same company, and only one would be fixed!

On your other point, the government still has systemic leverage and can compel access, so this doesn't remove that risk.

That doesn't mean this is the end of the world, and some balance of power is usually good. But I do think it will still increase the capabilties of rogue actors and their net harm.

lebovic··on Claude Fable 5
Not in the same way.

A customer could sign a ZDR agreement with Anthropic, and their API usage wouldn't be retained for even a day. That's no longer possible.

lebovic··on Claude Fable 5
While this makes it easier for Anthropic to detect misuse, it also means that the US government and other parties have access to every message and response from every user.

This applies even with API usage through third-party inference providers (e.g. AWS' Bedrock and GCP's Vertex) or with a zero-day data retention agreement in place.

I understand the reasoning for doing this, but I don't love the precedent that it sets.

lebovic··on The ways we contain Claude across products
I'm late to this thread, but the post seems to skip the section about risks/mistakes/incidents with restricting Claude's access with containers ("pattern 1"). Doing this properly is still hard!

For example, Anthropic has shipped several bugs that allow any claude.ai/code session – which are isolated in ephemeral containers – to access and exfiltrate all of the user's other sessions, connected repos, and environment variables. The rogue/hijacked Claude could also spawn new Claude sessions with arbitrary instructions and access, regardless of the original session's constraints.

I originally wrote about this (with permission) in February[1], and most of the issues were quickly fixed. But the underlying token scope issues have regressed several times since then – including post-Mythos – so I wouldn't say that Anthropic has solved this yet.

[1]: https://www.noahlebovic.com/hacking-claude-code-on-the-web-b...

lebovic··on Benchmarking open-weight models for security research
GLM 5.1 is surprisingly capable. Anecdotally, I couldn't notice a difference until ~120K tokens.

Qwen 3.6 35B A3B also exceeded my expectations. It's surprisingly performant, even though the previous generation wasn't even able to use the testing harness.

(Tbd on Kimi K2.6; the eval is still running.)

lebovic··on Trusted access for the next era of cyber defense
It seems reasonable for a company to require KYC for a product that's dual use – especially a novel one that's built for security research.

Privacy concerns aside, the KYC process for OpenAI was self-serve and took about a minute.

lebovic··on Evaluation of Claude Mythos Preview's cyber capabilities
I think the third chart is the most notable; Mythos is the first model which saturated that eval from the UK AISI [1].

Personally, I think we crossed the threshold of meaningfully useful capabilities for autonomous hacking with Opus 4.6 [2], mostly because its behaviors and persistence are useful for finding vulnerabilities out of the box [3]. But it still seems like Mythos is another step up.

[1]: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a48/...

[2]: https://www.noahlebovic.com/testing-an-autonomous-hacker/

[3]: https://news.ycombinator.com/item?id=46920682

lebovic··on Assessing Claude Mythos Preview's cybersecurity capabilities
A plateau is unlikely, at least for cybersecurity. RL scales well here and is replicable outside of Anthropic (rewards are verifiable, so setting up the training environment doesn't require that much cleverness).

The post also points out that the model wasn't trained specifically on cybersecurity, and that it was just a side-effect – so I think there's still a lot of headroom.

It's scary, but there's also some room for cautious non-pessimism. More people than ever can cause billions of dollars of damage in attacks now [1], but the same tools can be used for defensive use. For that reason, I'm more optimistic about mitigations in security vs. other risk areas like biosecurity.

[1]: https://www.noahlebovic.com/testing-an-autonomous-hacker/

lebovic··on Claude Code Found a Linux Vulnerability Hidden for 23 Years
It cost me ~$750 to find a tricky privilege escalation bug in a complex codebase where I knew the rough specs but didn't have the exploit. There are certainly still many other bugs like that in the codebase, and it would cost $100k-$1MM to explore the rest of the system that deeply with models at or above the capability of Opus 4.6.

It's definitely possible to do a basic pass for much less (I do this with autopen.dev), but it is still very expensive to exhaustively find the harder vulnerabilities.

lebovic··on Statement on the comments from Secretary of War Pete Hegseth
Others have addressed the first half of your comment, so I'll focus on the astroturfing claim.

While I've talked a lot about Anthropic this week, if I was astroturfing for a positive image, I'd be very bad at it [1][2][3].

[1]: https://news.ycombinator.com/item?id=47150170

[2]: https://news.ycombinator.com/item?id=47163143

[3]: https://news.ycombinator.com/item?id=47174814

lebovic··on Statement on the comments from Secretary of War Pete Hegseth
I used to work at Anthropic, and I wrote a comment on a thread earlier this week about Anthropic's first response and the RSP update [1][2].

I think many people on HN have a cynical reaction to Anthropic's actions due to of their own lived experiences with tech companies. Sometimes, that holds: my part of the company looked like Meta or Stripe, and it's hard not to regress to the mean as you scale. But not every pattern repeats, and the Anthropic of today is still driven by people who will risk losing a seat at the table to make principled decisions.

I do not think this is a calculated ploy that's driven by making money. I think the decision was made because the people making this decision at Anthropic are well-intentioned, driven by values, and motivated by trying to make the transition to powerful AI to go well.

[1]: https://news.ycombinator.com/item?id=47174423

[2]: https://news.ycombinator.com/item?id=47149908

lebovic··on Statement from Dario Amodei on our discussions with the Department of War
I don't think it's cynical to believe that a company can make the world a worse place, or that Anthropic as a company will make many horrible choices.

I do think it's cynical to believe that people, and groups of people, can't be motivated by more than money.

lebovic··on Statement from Dario Amodei on our discussions with the Department of War
Sorry, I meant a different Sam – Sam McCandlish, not Sam Altman.

Wasn't expecting this post to get so much attention.

lebovic··on Statement from Dario Amodei on our discussions with the Department of War
Yeah, values on their own don't lead to positive outcomes. I agree that many groups that are driven by ideals have still committed horrible acts.

I do think that they're acting with positive intent, though, and are motivated by trying to make the transition to powerful AI go well.

Many folks on HN seem to assume the primary motivation is purely chasing more money, which certainly isn't the case for for many – but not all – people at Anthropic.

That doesn't guarantee a good outcome, and there's still a hard road ahead.

lebovic··on Statement from Dario Amodei on our discussions with the Department of War
Hah, you're right, I meant Dario Amodei, Jared Kaplan, and Sam McCandlish.

They're all cofounders of Anthropic. Dario is the CEO, Jared leads research, and Sam leads infra. Both Jared and Sam were the "responsible scaling officer", meaning they were responsible for Anthropic meeting the obligations of its commitments to building safeguards.

I think neom is referring to Jack Clark, another one of the seven cofounders.

lebovic··on Statement from Dario Amodei on our discussions with the Department of War
Yeah, I think that's one way it could go!

I think both situations are pretty scary, honestly, and it's hard for me to have high confidence on which one would lead to less risk.

lebovic··on Statement from Dario Amodei on our discussions with the Department of War
> What are those values that you're defending?

I think they're driven by values more than many folks on HN assume. The goal of my comment was to explain this, not to defend individual values.

Actions like this carry substantial personal risk. It's enheartening to see a group of people make a decision like this in that context.

> Which one of the following scenarios do you think results in higher X-risk [...] There are no meaningful AI risks in such a world

I think there's high existential risk in any of these situations when the AI is sufficiently powerful.

lebovic··on Statement from Dario Amodei on our discussions with the Department of War
Yeah, I didn't mean this as a reflection of my morality, more to counter the financial and "rosy picture" parts of their comment.
lebovic··on Statement from Dario Amodei on our discussions with the Department of War
It's enheartening to see someone make a decision in this context that's driven by values rather than revenue, regardless of whether I agree.

I dissented while I was there, had millions in equity on the line, and left without it.

lebovic··on Statement from Dario Amodei on our discussions with the Department of War
I used to work at Anthropic, and I wrote a comment on a thread earlier this week about the RSP update [1]. It's enheartening to see that leaders at Anthropic are willing to risk losing their seat at the table to be guided by values.

Something I don't think is well understood on HN is how driven by ideals many folks at Anthropic are, even if the company is pragmatic about achieving their goals. I have strong signal that Dario, Jared, and Sam would genuinely burn at the stake before acceding to something that's a) against their values, and b) they think is a net negative in the long term. (Many others, too, they're just well-known.)

That doesn't mean that I always agree with their decisions, and it doesn't mean that Anthropic is a perfect company. Many groups that are driven by ideals have still committed horrible acts.

But I do think that most people who are making the important decisions at Anthropic are well-intentioned, driven by values, and are genuinely motivated by trying to make the transition to powerful AI to go well.

[1]: https://news.ycombinator.com/item?id=47145963#47149908

lebovic··on Anthropic drops flagship safety pledge
> Every employee you talk to forced to pretend that the company is all about philanthropy, effective altruism and saving the world

I was an interviewer, and I wasn't encouraged to talk about philanthropy, effective altruism, or ethics. Maybe even slightly discouraged? My last two managers didn't even know what effective altruism was. (Which I thought was a feat to not know months into working there.)

When did you interview, and for what part of the company?

> knowing fully well that we'd do what the bosses told us to do [...] now that real money is on the line

This is a cynical take.

I didn't just do what I was told, and I dissented with $XXM in EV on the line. But I also don't work there anymore, at least one of the cofounders wasn't happy about it and complained to my manager, and many coworkers thought I had no sense of self preservation – so I might be naive.

The more realistic scenario is that a) most people have good intentions, b) there's a decision that will cause real harm, and c) it's made anyway to keep power / stay on the frontier, with the justification that the overall outcome is better. I think that's what happened here.

lebovic··on Anthropic drops flagship safety pledge
I don't hold the belief that it's always better to have influence in a group where you don't trust leadership – in this case, those who decide at the metaphorical table – vs. trying to affect change through a different avenue.

It's probably naive, but it's also the reasoning that drove many early employees to Anthropic. Maybe the reasoning holds at smaller scales but breaks down when operating as a larger actor (e.g. as a single person or startup vs. a large company).

lebovic··on Anthropic drops flagship safety pledge
I was willing to (and did) give up my equity.
lebovic··on Anthropic drops flagship safety pledge
I used to work at Anthropic. I fully believe that the folks mentioned in the article, like Jared Kaplan, are well-intentioned and concerned about the relationship between safety research and frontier capabilities – not purely profit.

That said, I'm not thrilled about this. I joined Anthropic with the impression that the responsible scaling policy was a binding pre-commitment for exactly this scenario: they wouldn't set aside building adequate safeguards for training and deployment, regardless of the pressures.

This pledge was one of many signals that Anthropic was the "least likely to do something horrible" of the big labs, and that's why I joined. Over time, the signal of those values has weakened; they've sacrified a lot to get and keep a seat at the table.

Principled decisions that risk their position at the frontier seem like they'll become even more common. I hope they're willing to risk losing their seat at the table to be guided by values.

lebovic··on Evaluating and mitigating the growing risk of LLM-discovered 0-days
The post is light on details, and I agree with the sentiment that it reads like marketing. That said, Opus 4.6 is actually a legitimate step up in capability for security research, and the red team at Anthropic – who wrote this post – are sincere in their efforts to demonstrate frontier risks.

Opus 4.6 is a very eager model that doesn't give up easily. Yesterday, Opus 4.6 took the initiative to aggressively fuzz a public API of a frontier lab I was investigating, and it found a real vulnerability after 100+ uninterrupted tool calls. That would have required lots of of prodding with previous models.

If you want to experience this directly, I'd recommend recording network traffic while using a web app, and then pointing Claude Code at the results (in Chrome, this is Dev Tools > Network > Export HAR). It makes for hours of fun, but it's also a bit scary.

lebovic··on Delay-Line Memory
See also this thread on storing data in space: https://news.ycombinator.com/item?id=46327158
← PreviousPage 2 of 4Next →