So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
First, there's an element of timeliness that simply has hard constraints. In order to perform a "proper" analysis of this situation (i.e., little to no dependence on AI tools), you'd have to expect a pretty long wait. I know I'd rather have some sort of "initial report" as quickly as possible than to wait a year or two to get a report about a situation that will likely look trivial in a year or two. I imagine we'll see more detailed, human-developed reports over longer time ranges.
Second, I suspect the expectation of non-AI driven reporting of these kinds of things will definitely decline rapidly as everything scales up quickly. I mean, the data being produced by situations like this comes in the form of natural language "forum posts" (so to speak), but done at an autonomous scale. This isn't a collection of emails and Slack messages posted by humans in an org over the course of a few months; this is a bunch of bots interacting with each other in relatively novel ways as quickly as possible. It is, unfortunately, a perfect job for LLMs.
None of this disagrees with your points, necessarily. But I just think it's worth pointing out that this doesn't seem like a case of "And look! METR is so confident in LLMs that we're able to use it instead of paying humans to save a buck :D" and more of "Without LLMs, we'd only be half-way done analyzing this data before there are dozens more such investigations on the docket, so this will have to do."
How can it be both so important that we need it so quickly, but at the same time have a tolerance for such plausible deniability? It just doesn't really make sense that both those things are true at the same time.
Perhaps if the government decides to sue OpenAI, we could get a more thorough investigation.
Hey other labs, this shit could be happening to you right now, take a look at this and stop it asap if you're seeing anything similar.
So yea, both things can be true at the same time. Kind of like when a particular type of building collapses, even if they don't know the causation they will send inspectors to other buildings of the same type to sure the walls aren't cracking apart in an obvious fashion.
> Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret.
> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.
METR is a cog in the effective altruism machine. It was spun off from Paul Christiano's Alignment Research Center. Christiano is a well-known longtermist and AI doomer, who predicts a 50% chance that AI will end humanity once it reaches human capacity [1].
The author of this piece is also a well-known member of the Bay Area rationalist cult.
[1] https://www.businessinsider.com/openai-researcher-ai-doom-50...