HNHacker News
TopNewBestAskShowJobs

lebovic

1,931 karma · joined April 2, 2019

Noah Lebovic

Email is my firstname@lastname.com

Formerly at Anthropic, co-founder of Toolchest (a YC startup), and a bio startup

submissionscomments
lebovic··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
I formed a PBC and worked at a well-known PBC. Personally, I opted for a PBC because I liked that I could balance a specific cause with shareholder benefit.

In most cases, it doesn't really matter. The board + management is still in charge, and they have significant legal leeway regardless of the structure. But there's little additional cost to opt for a PBC, and it does give you more legal defensibility to be truly mission driven. Standard C Corps weren't really intended for mission driven companies (see the shareholder primacy norm).

I think it's popular for AI startups, because many great researchers understand the risks involved, and they don't want what they build to be controlled solely for shareholder benefit.

While non-profits are also an option for a mission driven org, it's harder to raise the large amounts of cash that some AI startups need, and laws around deferred compensation and private inurement (e.g. options-like structures) make employee compensation harder.

lebovic··on UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
> all frontier models benefit from more tokens not just Kimi K3

Past a point, that doesn't hold and the score plateaus.

Token hungry models tend to plateau at a much higher token count. Because Kimi K3 is a token hungry model – and 100M tokens (including cache hits) seems at the edge of the plateau for these evals – it could disproportionately benefit from a higher token budget.

For Kimi K3 specificially, policymakers are interested in whether it can find and exploit the same scope of vulnerabilities as models like Mythos. In that context, an answer of "yes, but with quintuple the token budget" is materially different from "no, it performs significantly below the most recent frontier cyber-capable models".

(As an aside, I like the UK AISI and think they're the best example of that kind of group!)

lebovic··on UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
The UK AISI post is https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...
lebovic··on UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
> Kimi K3 performs significantly below the most recent frontier cyber-capable models

UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M token limit well before saturating scores [2].

This gap was true for GLM 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I've found GLM 5.2 to be better at security research than Opus 4.6 [4]. But it's a quirky model that degrades quickly at long context lengths.

Personally, I'd rank Kimi K3 above Opus 4.8 and lower than GPT 5.6 Sol in its ability to find vulnerabilities and exploit them. But it's not far from the frontier.

[1]: From the UK AISI: "Our setup likely slightly underestimates open weight models’ maximum capability: we didn’t pursue specific elicitation or optimisations which could have improved performance" (https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are...).

[2]: Their eval also counts cache hits towards the token budget; the 100M token budget is comparable to a ~5M token budget in other evals.

[3]: See GLM 5.2 eval scores in https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are...

[4]: https://dualuse.dev/posts/chinese-models-are-sometimes-bette...

lebovic··on OpenAI’s accidental attack against Hugging Face is science fiction that happened
Almost exactly a year ago, we released Opus 4.1. It was definitely capable of finding vulnerabilities, and people were using custom harnesses to do so quite effectively.

The newer models are still more capable, but there were people doing this and writing about it (e.g. XBOW).

lebovic··on OpenAI’s accidental attack against Hugging Face is science fiction that happened
Yes, I think you could probably get something similar from Opus 4.5 (2025). Definitely Opus 4.6. I still think recent models are more capable, though!

Some of the model behaviors that make it better at pentesting, like persistence, can be improved with harness-level tricks (e.g. alloys, automated nudges, pre-fill to promote persistence, coordinated swarms, etc).

You mentioned the UK AISI's evals. Their harness is like basic Claude Code with compaction, and it doesn't include any of these tricks (afaik). As a result, I interpret their evals as a lower-bound of capabilities.

Newer models are still more capable, and they require almost no harness to find and exploit vulnerabilities. They're also more capable of performing more complex long-horizon attacks. But we've been past the threshold of modes capable of autonomous hacking for a while now [1].

[1]: Opus 4.6 was used for https://www.noahlebovic.com/testing-an-autonomous-hacker/

lebovic··on Qwen 3.8
(This comment was originally on another merged post, and "this page" referred to https://www.qwencloud.com/pricing/token-plan)
lebovic··on Qwen 3.8
I'm haven't found an announcement page, but there's a banner on the website announcing Qwen 3.8 and redirecting to this page.

Looks like they're previewing the model only on their subscription plan.

lebovic··on The Kimi K3 Moment
Kimi K3 only supports "max" reasoning effort right now, but they plan to enable other levels soon [1].

When I looked at traces from benchmarking, I saw a lot of backtracking and uncertainty while reasoning ("wait, but..."). This also happens with GPT 5.6 and Fable with xhigh/max thinking, albeit to a lesser degree.

I think that explains part of the token inefficiency. Hopefully it will improve with lower reasoning effort settings.

[1]: https://platform.kimi.ai/docs/guide/use-thinking-effort

lebovic··on Claude Science
After spending years on a problem, it's exciting to see it start to get more attention and move towards being meaningfully solved.

But I try to limit my time on HN, and I thought someone who works on Claude Science might respond to this thread later.

lebovic··on Claude Science
I assume they do hallucinate, just like with coding or finding vulnerabilities.

You can try to minimize it (e.g. with a reviewer agent, which Claude Science and Biomni have), but nothing is perfect, so I limit autonomous work to verifiable problems and review it.

lebovic··on Claude Science
I can't speak for Claude Science, but I prefer using Biomni as an agent for bio over Claude Code with a custom setup because a) Biomni stays on the frontier for bio, b) it has a config that just works and skills I trust are correct, and c) it has better built-in abstractions for long-running sessions.

As a concrete example, computational biology jobs sometimes run for hours on the Biomni HPC. When they're done, the session needs to reawaken, process the results, iterate, etc. You can implement something like this with agent callbacks, but it's not as straightforward.

This repeats many times for many integrations, so it's just simpler for me to use an agent that's built for exploratory bio and already has all of this. Claude Science has some of these features, so I imagine they're aiming for something similar.

lebovic··on Claude Science
I built one of the connected tools included in this launch (the Biomni HPC [1]), and I have spent an inordinate amount of my life working on this problem. (I also worked at Anthropic, but not on this product.)

As other comments have pointed out, this is for data science – but it's capable of more than making plots and writing papers [2]. It has integrations with many databases and computational tools, including a researcher's institutional cluster.

That alone is valuable. I founded a startup after struggling with this problem at a bio startup; integrating these tools and databases is hard and time consuming. If the only outcome of this product is that great APIs are built for LLMs, it will be a massive positive impact. Many databases used in computational genomics are still only accessible through FTP!

LLMs are particularly good at navigating these tools and databases. It's often very specialized, but straightforward, work that benefits from in-context skills. Seeing an early glimpse of my former customers – bioinformaticians – using LLMs to solve this problem is what led me to join Anthropic in 2024.

Also, this pattern isn't fundamentally constrained to data science: you can also integrate with a wet lab or a CRO for some kinds of science. This is what I'm spending my time on now.

This type of science doesn't solve everything, but it's useful in some niches. For example, progress on many rare diseases is bottlenecked by researcher attention rather than a fundamental breakthrough.

[1] https://x.com/phylo_bio/article/2029233694775624096

[2] In comparison, OpenAI's science product – Prism – was effectively a LaTeX editor they acquired with Crixet.

lebovic··on Chinese models are sometimes better, even if they're distilled
In this case, the benchmarks were private and it still outperformed.
lebovic··on GLM 5.2 beats Claude in our benchmarks
GLM 5.2 and DeepSeek v4 Pro seem to approach security research differently. This benchmark was with GLM 5.1, but the patterns are similar: https://dualuse.dev/posts/deepseek-v4-thinks-different

Overall, I still think GLM 5.2 is the much stronger performer. It's hard to tell the difference between GLM 5.2 and Opus at <120k tokens.

lebovic··on Anthropic says Alibaba illicitly extracted Claude AI model capabilities
Thanks!

For that eval, I used an account that was labeled as a known red-teaming org by Anthropic, and I read the traces. There were no refusals or obvious avoidance behaviors, though it may have been silently nerfed.

On the same eval, Opus 4.7 and 4.8 outperformed GLM 5.1, but GLM 5.2 is on par again with Opus. So it's at least partially measuring capabilities without respect to refusals.

One possible contributing factor is that model capabilities are shaped differently (an example of this is GLM 5.1 vs. DeepSeek v4 Pro: https://dualuse.dev/posts/deepseek-v4-thinks-different). So if you use RL-based "distillation" from multiple models like Opus 4.x and GPT 5.x, you could get a more capable model.

lebovic··on Anthropic says Alibaba illicitly extracted Claude AI model capabilities
Curiously, this isn't always true.

For example, GLM 5.1 is more capable at pentesting than the model from which it is alleged to have been distilled [1].

Intuitively, this makes some sense: you can "distill" from multiple frontier models, and you can further post-train the distilled model. But I'm not sure exactly what happened with GLM 5.1.

[1]: https://dualuse.dev/posts/chinese-models-are-sometimes-bette...

lebovic··on Anthropic says Alibaba illicitly extracted Claude AI model capabilities
It's too late to prevent distillation of some capabilities, like writing code or finding vulnerabilities [1].

But an AI lab can continue to produce immense economic value without releasing the model publicly for potential distillation. For example, it could use a model solely in-house to develop therapeutics.

Hopefully there's a future where others can access frontier models, but it's not neccessary if preventing proliferation through distillation is considered more important.

[1]: See the notes on distillation in https://dualuse.dev/posts/export-controls-on-fable

lebovic··on Health insurance claim denial rates range from 13% to 35% by insurer
This is already a thing! For example, Neon Health does this for providers. I haven't heard of any changes to the process yet, but I imagine insurers move slower than startups.
lebovic··on Midjourney Medical
Full wave inversion uses all of the information from the wave and more intense computational tomography to image structures that pulse wave B mode cannot, though gases are still a problem. Computationally, if you squint, it's similar to the work Midjourney does with AI image generation, as it progressively generates a structure that fits the data.

Ultrasonic waves can penetrate most structures in humans, including the brain. For example, with focused ultrasound (as they mentioned with MRgFUS) you can burn specific structures in the middle of the brain without any incision.

To use this for imaging, you need lots of transducers (MRgFUS typically uses 1024 for ablation, and Midjourney is proposing 358,000 for imaging) and massive advances in computational tomography capabilities. There will still likely be pockets of low confidence where there's a lot of air, like in the lungs. But with sufficient information on what's happening around those areas, you'd still have something that's medically useful.

lebovic··on Amazon CEO's Talks with U.S. Officials Triggered Crackdown on Anthropic Models
They made a deal for access, but I'm unsure if it's usable, scaled, and has vulnerabilities attributed to it at this point. But I have no inside information here, so I could be wrong.
lebovic··on Amazon CEO's talks with U.S. officials triggered crackdown on Anthropic models
Claims of retribution aside, one steelman is that Mythos is likely the most capable model that's usable by folks like the NSA [1], and decision-makers across the USG and industry partners have seen a stream of reports of Mythos successfully finding serious vulnerabilities over the past couple months due to Glasswing.

So even if GPT 5.5 is just as capable in these scenarios (which, imo, it largely is), it is not known by the government apparatus as having the same capabilities.

Personally, I think we crossed the threshold of capabilities with Opus 4.6 [2], which translated to an even more capable open-weight GLM 5.1 (which it is rumored to have distilled Opus 4.6) [3][4]. But the USG and its partners aren't fully rational actors with perfect data, so it's possible they're only viscerally aware of these capabilities in the context of Mythos.

[1]: https://www.reuters.com/business/us-security-agency-is-using...

[2]: Opus 4.6 was used for https://www.noahlebovic.com/testing-an-autonomous-hacker/

[3]: See GLM 5.1 scoring in https://www.cybergym.io/cybergym/

[4]: https://dualuse.dev/posts/chinese-models-are-sometimes-bette...

lebovic··on Anthropic apologizes for invisible Claude Fable guardrails
That's a good clarification. I've updated my comment to the "most capable models" to refer to the most recent releases.

And sure, and I love open models – I spent much of the past couple months doing additional RL on Qwen 3.6 35B A3B, Gemma 4, Kimi K2.6, and GLM 5.1. Without these open models, I'd be forced to do my research inside a frontier lab.

There's a balance to strike here, but I don't think the biological risk is overplayed. It would be very easy to accidentally cross the threshold of "meaningful" without adequate safeguards, and then be unable to undo what you've released to the world.

lebovic··on Anthropic apologizes for invisible Claude Fable guardrails
In normal bio, there are standardized biosafety levels, because without it there would be no standard agreement on what "meaningful" safety is. So yes, I do think there's ambiguity here.

But I don't think I've found any domain expert who thinks granting everyone raw access to the most capable models wouldn't meaningfully increase risk. OpenAI recently staffed a biological threat modeler to help quantify this risk.

(Edit: just saw your edit, this includes at Anthropic. ASL tiers were "rule-out" to exclude rather than "rule-in", so exact thresholds were murkier, but I think it's clear that models have passed that threshold by now.)

That said, there are clear steps and requirements to set up a BSL-2 or BSL-3 lab, and I think there should be similarly clear rules around model capabilties and access. The process for Anthropic and OpenAI is murky and still implictly gated on spend, which I think is holding back research.

For example, anyone who has access to a BSL-3 lab should have a clear and low-cost path to a model with corresponding capabilities, as long as they set up corresponding precautions for model access.

I think it would be a bad outcome for only frontier labs and a select few groups they choose to have access to the most capable models – which is sadly the precedent that's currently being set.

lebovic··on Anthropic apologizes for invisible Claude Fable guardrails
No, Anthropic's model cards have claimed that the models don't show considerably more uplift than previous ASL-3 models, which already showed material uplift.

I participated in the internal bioweapons uplift test for Sonnet 3.7, and even then, one non-expert got huge uplift from the model [1]. I'd consider evals a lower bound of capabilities that can be elicited from a model.

The team behind Biomni, a biomedical agent that's widely used by researchers, has continued to find consistent gains between models [2]. I trust them, because I visited them to build their HPC tool [3], which the model is quite capable of using – moreso than most grad students. The Biomni team cares a lot about about real usability for real researchers, so they have a great pulse on capabilties.

SecureBio also has some public evals [4], which have continued to show increasing uplift.

And while synthesis monitoring is a part of the solution, I think you might underestimate how much goes under the radar. See the Reedley lab incident for an example [5].

Is Anthropic still effectively throttling beneficial biomedical research? Yes! And so is OpenAI. But the underlying capability is still actually dual use.

[1]: See page 25 in https://www-cdn.anthropic.com/9ff93dfa8f445c932415d335c88852...

[2]: Their benchmark has a preprint at https://www.biorxiv.org/content/10.64898/2026.05.12.724604v1...

[3]: https://x.com/phylo_bio/article/2029233694775624096

[4]: https://securebio.org/

[5]: Search for "ebola" in the public report for the Reedley lab incident at https://chinaselectcommittee.house.gov/sites/evo-subsites/se...

lebovic··on Anthropic apologizes for invisible Claude Fable guardrails
It sounds like you might not agree with that belief.

While I don't agree with their actions here, I do think there's sufficient reason to hold that belief.

On some fronts (e.g. security, on which you've experienced more than me), I think there are surmountable challenges. But on other fronts (e.g. bio), a single errant actor could reasonably kill millions or billions of people with sufficiently powerful AI. We don't have good defenses here, and those actors do exist.

I still don't agree with these actions, but I do think I agree with their assumptions.

lebovic··on I'm Eric Ries, author of "The Lean Startup" and new book "Incorruptible" – AMA
My point was that Anthropic has tended to make atypical decisions vs. its peers, not that they're always the right decisions. The direction of those decisions has tended towards assuming exponential growth of AI will continue and a certain flavor of AI safety.

This decision does seem in line with what I would expect from Anthropic, so I don't see it as a sign of changing values – even if I personally disagree.

lebovic··on I'm Eric Ries, author of "The Lean Startup" and new book "Incorruptible" – AMA
Sure, but their question was whether different structure/governance would have changed that decision.
lebovic··on I'm Eric Ries, author of "The Lean Startup" and new book "Incorruptible" – AMA
Individual contributor; i.e. not a manager.

(This isn't a dig on managers; I've been one. But if a situation doesn't naturally escalate, that usually means a manager in the chain chose not to escalate it, and their reports have to go around them.)

lebovic··on I'm Eric Ries, author of "The Lean Startup" and new book "Incorruptible" – AMA
I think the difference is the ability to be persuaded by a strong argument that you've critically evaluated.

Some people are unjustly called stubborn when they don't change their position based on a weak argument from an authority figure. And others claim values, but they're just stubbornly adhering to something that feels good to believe.

Page 1 of 4Next →