As an expert, could you also provide your arguments please?
As an expert, could you also provide your arguments please?
So if IIUC your point is "they're not good enough at biology right now because they're not trained on it so they're not a threat".
To which I want to answer: "they're not a threat now but I see *no* reason for models not to be trained on biology pretty darn soon unless people like you convince the world otherwise."
Thoughts?
That's a big claim. I believe on the other hand that biology is applied chemistry which is applied quantum physics. And frontier models are so good at biology + biochemistry + chemistry + quantum physics that it's bound to trickle down into biology by sheer knowledge.
And secondly, even without pipettes i'm pretty sure computer simulation can teach us a tremendous amount still.
Till then I don't think current crop of LLM will ever touch that capability. I'll believe it when I see it
As for the future... today the LLMs are "trained on biology", in that they read the textbooks, the research, the web.
They aren't trained on biology in the sense of being embodied, autonomously or semi-autonomously driving actual biological experiments. If you come from software, the timescale of these experiments is outlandish. Yes, I am partly saying the LLMs are not good enough today because they need to be embodied and trained for literal decades of lab time before there is even the _remote_ possibility that they could present a novel risk profile that is even a shadow of what the current fearmongering suggests the current models can enable.
And they aren't trained on biology in the sense that they've read the literature, but even 100T token training run only begins to touch the data scales that rather mundane bioinformatics operate at. True multimodal models that work on DNA and human language at high quality haven't yet emerged. We're talking new architectures which are going to arise after the next AI winter.
All of this ignores an even more fundamental point. Cost. If someone wants to make a bionuke, they don't need to use AI. They can set up the right evolutionary context and run quadrillions of parallel explorations of the design space. Directed evolution like this is cheap, well-understood, and insanely powerful. If you actually care about biosafety, we should be doing hard work to surveil gain of function research. Different flavors of LLM use are not going to be a differentiator for the foreseeable future.