AI cracks superbug problem in two days that took scientists years
bbc.co.uk
bbc.co.uk
"Google Co-Scientist AI cracks superbug problem in two days! — because it had been fed the team’s previous paper with the answer in it" https://news.ycombinator.com/item?id=43162582#43163722
Let's see how many points gets the correction. It would be good that achieved the same or more visibility than this one to keep HN informative and truthful.
https://research.google/blog/accelerating-scientific-breakth...
https://pivot-to-ai.com/2025/02/22/google-co-scientist-ai-cr...
If it helps scientists find answers faster, I don’t see the problem—especially when the alternative is sifting through Google or endless research papers.
"He told the BBC of his shock when he found what it had done, given his research was not published so could not have been found by the AI system in the public domain."
Also:
Prof Penadés' said the tool had in fact done more than successfully replicating his research. "It's not just that the top hypothesis they provide was the right one," he said. "It's that they provide another four, and all of them made sense. "And for one of them, we never thought about it, and we're now working on that."
Either way, the headline is garbage. It's like being amazed that your coworker who you've documented every step in your process to managed to solve that problem before you. "I've been working on it for months and they solved it in a day!" is obviously false, they have all of the information and conclusions you've been producing while you worked on it. In this case, telling the coworker every detail is the non-consensual AI training process.
Society and people in general don’t want to hear these “sour grape” gripes. Unless they are the ones affected adversely.
"However, the team did publish a paper in 2023 – which was fed to the system – about how this family of mobile genetic elements “steals bacteriophage tails to spread in nature”. At the time, the researchers thought the elements were limited to acquiring tails from phages infecting the same cell. Only later did they discover the elements can pick up tails floating around outside cells, too.
So one explanation for how the AI co-scientist came up with the right answer is that it missed the apparent limitation that stopped the humans getting it.
What is clear is that it was fed everything it needed to find the answer, rather than coming up with an entirely new idea. “Everything was already published, but in different bits,” says Penadés. “The system was able to put everything together.”"
https://www.newscientist.com/article/2469072-can-googles-new...
AI has some potential here, because unlike a human, AI can be trained across all of it and has the opportunity to make connections a human, with more limited scope, might miss.
In this particular example it wasn't useful because the reader already knew the answer and was fishing for it with the LLM. But generally using an LLM as a creativity-inducer (a brainstorming tool) is fine, and IMO a better idea than trying to use them as an oracle.
This has been true since Waze and Google maps first came out. Some of the suggestions were like "Huh, I never thought of going that way... hey this works... good job Waze!" and some were like "No Waze, I'm not going to try to cross eight lanes of free-flowing traffic at rush hour."
What sucks is when an AI bot hallucinates some transgression and bans you from a monopoly marketplace for life, with no human recourse. Ask me how I know.
This is fundamentally an incentive problem. Whether MegaCorp reaches this decision through an AI or a team of people who made a mistake, interpreted a rule differently than intended, missed context, whatever, they have been allowed to set up the system in a way which has all the incentives in their favor. They don’t need you and further, helping you costs them money. Yet we allow them to be the only game in town and still say it’s at their convenience to give you service.
The focus in the modern world needs to be on re-incentivizing companies to do the right things, AI or otherwise.
Even in the case of something like a credit card company where it’s not technically a monopoly, you risk them all coming to the same conclusion.
Essential for-profit service combined with at-will service is simply a recipe for failure.
All appeals were denied with increasingly vague language, either by more bots, or humans with strong incentive to just rubber stamp the bot's decision and move on. I'll never know which.
It's all here in this pamphlet: https://news.ycombinator.com/item?id=40992654
A weird statement. I think you only care about humans knowing the thing, right? Then it can be true that it doesn't matter if the machine is right or wrong.
However, if the machine is right.. that will have consequences.
That's basically 99% of startups/business too.
> to make connections a human, with more limited scope, might miss.
I have a concern, that all this AI stuff is giving less time to "fuck around and find out" (i.e. experimentation). Just like how shifting from compiled languages leads to less time thinking about the code or reading docs while things are compiling. Sometimes you just need to walk away from the desk to solve a problem. It's kinda ironic, that reducing blocking aspects creates more. But the reason walking away from the desk works is because your creative part needs time to think and imagine. It's why you play with your lab equipment or make fun programs to learn. But if everything is focused on only going to the product then you end up straying away from that.I do use AI myself but tbh I'm constantly fighting it and find it frequently misses small subtle things that are important. I even notice that people accept answers that I wouldn't, and that's sometimes concerning
Usually.
The media portrays science in an unhelpful manner, gpt-3.5 didnt appear out of thin air, deekseek R1 was built on deepseek MOE was built on research by mistral
LLMs also happen to be pretty good at it, but unlike humans they don’t get bored or tired from doing it too much.
The exciting bit will be when AI does something nobody has done before. Then we know for sure it isn't cheating.
I get that they asked it about a new result they hadn't published yet, but the idea that it did it in two days when it took them a decade -- even though it's been trained on everything published in that decade, including whatever intermediate results they published -- probably makes this claim just as absurd as it sounds.
edit - instantly is apparently many hours, hence the 'two days', just to be clear
> He told the BBC of his shock when he found what it had done, given his research was not published so could not have been found by the AI system in the public domain.
and,
> Critically, this hypothesis was unique to the research team and had not been published anywhere else. Nobody in the team had shared their findings.
Unfortunately, even if we wanted to, it might be impossible to attribute this success to the authors that it built upon.
By AI trainers, if not by the authors whose works were encoded.
public domain is a legal term meaning something like "free from copyright or license restrictions"
I believe in this thread people mean not that but "published and likely part of the training corpus"
It is still a big accomplishment if the AI helped them solve the problem faster than they could have without it, but the headline makes it sound like it solved the problem completely from scratch.
This kind of hyperbole is what makes me continue to be skeptical of AI. The technology is unquestionably impressive and I think it's going to play a big role in technology and advancement moving forward. But every breakthrough comes with such a mountain of fluff that it's impossible to sort the mundane from the extraordinary and it all ends up feeling like a marketing-driven bubble waiting to burst.
This seems like the most important detail, but it also seems impossible to verify if this was actually the case. What are the chances that this AI spat out a totally unique hypothesis that has absolutely no corollaries in the training data, that also happens to be the pet hypothesis of this particular research team?
I'm open to being convinced, but I'm skeptical.
> It also seems impossible to verify if this was actually the case.
If this is a thinking model, you could always debug the raw output of the model's internal reasoning when it was generating an answer. If the agent took 48 hours to respond and we had no idea what it was doing that whole time, that would be the real surprise to me, especially since Google is only releasing this in a closed beta for now.
"It's not just that the top hypothesis they provide was the right one," he said.
"It's that they provide another four, and all of them made sense.
And for one of them, we never thought about it, and we're now working on that."”I don't know that their policy says about that, or if it is even something they do...at least not publicly.
sounds extra fishy, since google does not provide email support normally.
These LLM models are essentially trying to produce material that sounds correct, perhaps the hypothesis was a relatively obvious question with the right domain knowledge.
Additionally, he may not have been the first to ask the question. It's entirely possible that the AI chewed up and spat out some domain knowledge from a foreign research group outside of his wheelhouse. This kind of stuff happens all the time.
I personally have accidentally reinvented things without prior knowledge of them. Many years ago in University I remember deriving a PID controller without being aware of what one was. I probably got enough clues from other people/media that were aware of them, that bridging that final gap was made easier.
https://support.google.com/meet/answer/14615114?hl=en#:~:tex...
You may not believe them, but I challenge your description of it as "openly".
> We do not use your Workspace data to train or improve the underlying generative AI and large language models that power Gemini, Search, and other systems outside of Workspace
I.e., they could use it to train an LLM specific to your Workspace, but that training data wouldn't be used outside of the Workspace, for the general generative products.
Was signing up for a workspace account and signing a EULA permission? Is permission given in an implicitly signed document like a privacy policy?
It is an awful sentence, but I'm reading it as:
> We do not use your Workspace data to train or improve the underlying generative AI and large language models [...] outside of Workspace [...].
Which to me makes it sound like if the answer is in your Workspace, and you ask an LLM in your Workspace, it would be able to tell you the answer.
I'm reminded of when Microsoft said you didn't actually buy the Xbox even though you thought you did and they won in court to prevent people from changing or even repairing their own (well, Microsoft's I guess, even though the person paid for it and thought they bought it) machine.
They do not.
Many years ago, they served customized ads based on your email. Then they stopped that, even for free accounts, because it led to a lot of unfounded misunderstanding that "Google reads your email, Google trains on your email"...
Then also consider how all of their enterprise customers would switch to MS/AWS in a heartbeat if they found out Google was training on their private, proprietary data.
Google Cloud has been a gigantic investment they've been making for well over a decade now. They're not going to throw away consumer and enterprise trust to train on a bunch of e-mail.
Enterprise customers are one thing, but private customers of ordinary Gmail? Completely another thing.
> Google Cloud has been a gigantic investment they've been making for well over a decade now. They're not going to throw away consumer and enterprise trust to train on a bunch of e-mail.
Consumers barely trust Google any more these days, not with the frequent stories about arbitrary bans, and corporations shy away from Google Cloud due to the same reason plus the service quality being way lower than AWS.
But yeah, in this specific case, it is way less nefarious. Just one, or both, of Google, and the scientist, selling a new AI product, with a sensationalist, unrealistic story, in a huge, publicly funded, "serious" news outlet. At least when the NYtimes drools down it's chin at some AI vaporware, they may be getting a huge advertising buy, or someone there owns a lot of stock in the company... BBC can't even hide behind "well that's capitalism baby".
I will say the prime minister and his red thatcherites have been obsessed with becoming a player in the AI industry... If you want a conspiracy theory i think is more likely haha.
Could it be the case when asking the right question is the key? When you know the solution already it's actually very easy to accidentally include some hints in your phrasing of question that will make task 10x easier
No double blind methodology protocol.
This is not to detract from the AIs accomplishment at all. If it read the scientist's prior work, then coming up with the same discovery as the scientist did is still astounding.
Today Google is launching an AI co-scientist, a new AI system built on Gemini 2.0 designed to aid scientists in creating novel hypotheses and research plans. Researchers can specify a research goal — for example, to better understand the spread of a disease-causing microbe — using natural language, and the AI co-scientist will propose testable hypotheses, along with a summary of relevant published literature and a possible experimental approach.
Without those years spent working the problem, would he have recognized that hypothesis as a valuable road to go down? And wouldn't the years of verifying it still remain?
https://scholar.google.com/citations?hl=en&user=rXUHiP8AAAAJ...
I wonder how AI is going to mess with the calculus of employing grad studies, and if this will affect the pipeline of future senior researchers...
Just want to make sure I'm understanding what's written in the article accurately.
https://research.google/blog/accelerating-scientific-breakth...
What actually happened here looks more like rubber-ducking. If you look at the prompt (Supplementary Information 1), the authors provide the LLM with a carefully posed question and all the context it needed to connect the dots to generate the hypothesis. The output (Supplementary Information 2) even states outright what information in the prompt led it to the conclusion:
"Many of the hypotheses you listed in your prompt point precisely to this direction. These include, but are not limited to, the adaptable tail-docking hypothesis, proximal tail recognition, universal docking, modular tail adaptation, tail-tunneling complex, promiscuous tail hypothesis, and many more. They collectively underscore the importance of investigating capsid-tail interactions and provide a variety of testable predictions. In addition, our own preliminary data indicate that cf-PICI capsids can indeed interact with tails from multiple phage types, providing further impetus for this research direction."
I feel like someone explained this in the 80s to me.
Sounds like all of those claims where ChatGPT allegedly coded a flappy bird clone from scratch. Only it didn't, it just regurgitated code from several Github repos.
(and I wonder how many hallucinated solutions the LLM came up and were rejected - sorry "refined" - by the team).
I think we are starting to get to the root of the utility of LLMs as a technology. They are the next generation of search engines.
But it makes me wonder, if we had thrown more resources towards using "traditional" search techniques on scientific papers, if we could have gotten here without gigawatts of GPU work spent on it, and a few years earlier?
This is performance art, right?
A few years back I was so sick of blockchain vaporware, and honestly couldn't think of anything more annoying... But several years of reading "serious" outlets publish stuff like "AI proves the existence of God", or "AI solves cold fusion in 15 minutes, running on a canon R6 camera" makes me wish for the carefree days of idiots saying "So you've heard of Uber, now imagine Uber, but it's on the blockchain, costs 0.1ETH just to book a ride, and your home address is publicly accessible"...
The '''journalist''' is, too: