Chai-1: Decoding the molecular interactions of life
chaidiscovery.com
chaidiscovery.com
This is extremely exciting news if true, so I'm eager to have it either confirmed or questioned. The one thing I hope we won't be doing is accepting SOTA evals from open-sourced models at face value.
For rest of us, this is a privilege that we don't have. We can't deceive, defraud our investors because it has real consequences....but not for people like Matt Schumer, why is that?
Doesn't seem so since they have seemingly endless capital but have limits in what they can bring to bear. You tell me...
But it’s an interesting question. You can’t be too risk-averse because there’re thousands of patients dying horrible deaths every single day. There’s simply a need for bold approaches in many areas of medicine.
Also the paper says they basically copied the methods used for AlphaFold, but then included the ability to input language embeddings, and input some other side constraints that I don't have the biology knowledge to understand. They don't show any data that indicate how much these changes improve performance. They show a very modest improvement over AF3 (small enough that I would think it could be achieve through randomness/small variations in the training parameters). So I don't think this is very revolutionary, but I suppose it replicates AF3.
Now, "predictions for parts of drug discovery" isn't the widest range, so perhaps you need to consider "foundation" as somewhat context dependent, but I don't think it's a wild claim. Neither "foundation" nor "fine tuned" are really better than each other, but those are probably the two ends of a spectrum here.
My get-out clause here is that someone with a better understanding of the field may say these are actually extremely narrowly trained things, and the tests are equivalent to multiple different coding problem challenges rather than programming/translation/poetry/etc.
The table with a comparison to alpha fold gives a less than one percentage point improvement.
Would be interesting to try to estimate the impact of results like these across the drug development pipeline.
E.g. N% improvement on our most predictive benchmarks X, Y, Z could impact clinical success by M% +- E% (where E would likely be quite large).
Lo and behold I see Chai did the same shit in their repo. Lol
Genuinely curious, would love to learn if that isn’t true / or is generally just not that big of a deal compared to other risks.
This gets into the philosophy of restricting access to knowledge. The conclusion I keep arriving at is that we’re lucky that there don’t appear to be many Timothy McVeighs walking around. I don’t think there is a practical defense from people like that.
This is vastly difficult to achieve using biology. Any organism on the planet has it's own agency, and it will hit anything to reproduce and eat. In addition this is not limited to toxicology and releasing toxins, because the agent can just eat tissue.
For example phosphorus has been used in chemical warfare, but even that cannot be described 100% as a weapon. The phosphorus gas can hit people who released it the same as everyone else, it just depends on the wind.
Right now, on everyone palms, there are thousands of organisms which create electricity, eat wood and kill animals. Given that the palms are washed, that number is reduced to some thousand different species. If the palms are not washed the last 24 hours, that number shoots up to hundred thousand different species, even millions.
I do not see any difficulty for someone to enhance a harmful agent and make it deadly, using just regular computation and not even A.I.. However the person who facilitated this, will be a target too.
I don't know how much releasing this model is a delta on safety, but we certainly need to do a better job of vetting who can order viruses; my understanding is there's very little restrictions right now. This will become more important as models get more capable.
In short there are other ways to negatively affect large numbers of people that are easier, and presumably those avenues are being explored first. But we don't know what we don't know.
For instance, would any of the following technologies be acceptably "safe"?
- physical locks (makes it possible to keep work secret or inaccessible to the government)
- solar power (power is suddenly much cheaper, means bad guys can do more with less money)
- general workload computers (run arbitrary code, including bad things)
- printing press (ideology spreads much more quickly, erodes elite hold over culture)
- bosch-haber process (necessary for creating ammunition necessary to fight the world wars)
- nuclear fission, which provides an abundant source of environmentally friendly energy, but allows people to make bombs capable of wiping out whole cities at once (and potentially causing nuclear winter)
But even in that case, I believe that it's a good thing that we have access to nuclear power, and I certainly want us to use more nuclear power. At the same time, I'm very glad that a bomb is hard enough to make that ISIS couldn't do it, let alone any number of lone wolf terrorists. So I think I would apply the same logic to biotechnology; speeding up medical progress seems extremely valuable and I'm excited about how AF and other AI systems can help with this, but we should mitigate the ability for bad actors to use the same tools for evil.
An aspect that's unique about biotechnology that's different in comparison to the examples you gave is that most of those technologies help good and bad people approximately equally, and since there's many more reasonable than crazy people they're not super dangerous.
There's a concern that technologies that make bioengineering easier could make it easier to produce and proliferated novel pathogens, much more so than they make it easier to prevent pandemics; in other words, it favors "offense" more than "defense". The only one example you listed that has a similar dynamic in my mind is the bosch-haber process, but that has large positive downstream effects separate from its use for ammunition. Again, this is not to say we should stop medical progress, but that we should act to mitigate the dangers, and keep this concept in mind as the technology progresses.
That said, I'm not certain how much the current tools are dangerous in this way. My understanding is that there is lower hanging fruit in mitigating these issues right now; for example, better controls at labs studying viruses, and better vetting of people who order pathogens online.
Structure prediction is just one small slice of all of the things you'd need to do. Choosing a vector, culturing it, splicing it into an appropriate location for regulation, making sure it's compatible with the environment, making sure your payload is conserved, study the mechanism of infection and make sure all of the steps are unimpeded, make sure it works with all of the host and vector kinetics, study the pathology, study the epidemiology. And that's just for starters.
This would require a team and an enormous amount of resources. People motivated enough to do this can already do it and don't need the AI piece.
Then again, you can just release existing dangerous pathogens. Like, poison a water with something deadly. So you don’t need a new one if you’re a terrorist.
Chemistry and molecular biology are fiendishly complicated fields, far more complex and less predictable than what general (and most of the non-biochem STEM majors) imagine them to be.
How do I know? I thought of one brilliant startup idea that would solve so many of the world's problems if only we used computers to simulate biological systems.
Result: https://xkcd.com/1831/
Reference materials:
https://www.amazon.ca/Molecular-Biology-Cell-Loose-Version/d...
I strongly recommend to treat it as introductory-level text on the same level as "K&R - C Programming Language". Yes, all 1464 pages of it.
https://www.amazon.ca/Fundamentals-Systems-Biology-Synthetic...
On the same level as above text, but with more math.
https://www.amazon.com/Introduction-Computational-Chemistry-...
That or any other book on computational chemistry will give you an understanding why it is difficult to design anything of value in biological systems. ML can only help so much.
Also check out this page for entire field scope:
It's done in an interesting style, with lots of direct references to current literature. I was surprised to see a recent edition on IA: https://archive.org/details/alberts-molecular-biology-of-the...
Are there any other books you would recommend?
https://www.coursera.org/specializations/bioinformatics
https://www.coursera.org/specializations/systems-biology
Also, checkout out
https://www.coursera.org/courses?query=computational%20biolo...
Then you will have a better understanding of the subject area and the literature to search.
[1] https://en.wikipedia.org/wiki/Tokyo_subway_sarin_attack
[2] https://en.wikipedia.org/wiki/Masami_Tsuchiya_(terrorist)
But they're also dumb, which is why they think terrorizing random people will positively I prove the world in some direction they care about.
I won't go into details, but I think if I had 19 dudes with a death wish in America, and a few million dollars, I could do something far worse than 9/11.
The attack had a significant degree of symbolism. The intended audience was twofold: the Western public and leadership, with a durable message that they weren't untouchable (hence the attacks on the Pentagon and attempt on the Capitol), hence targeting large landmarks; the combination of civilian and military targets was to signify that they held the two to he equivalent. Plans were actually presented to attack other targets that would lead to more casualties, notably a nuclear power plant.
The other goal was to incite a religious conflict from the Muslim world against the US, and therefore probably from the US against as many Muslim countries as possible.
So the primary goal really wasn't to kill as many random people as possible (though of course that was a consideration), it was actually to target the tallest buildings possible as well as the most important government institutions.
Unfortunately, it really did move the world in the direction they wanted. Despite being extremely evil, they actually were remarkably successful at causing the social and geopolitical changes they wanted given the resources they had, and that caused yet more damage we shouldn't ignore. It also bears remembering (especially today) that terrorists often and unfortunately aren't as dumb as we think, and we underestimate them and simplify their motives to our peril.
Submitters: "Please submit the original source. If a post reports on something found on another site, submit the latter." - https://news.ycombinator.com/newsguidelines.html
(Submitted title was "Chai-1 Defeats AlphaFold 3")
> Chai-1 achieves a ligand RMSD success rate of 77%, which is comparable to the 76% achieved by AlphaFold3
[0] https://chaiassets.com/chai-1/paper/technical_report_v1.pdf
If there is another line that said “+500 thus model will be forgotten and useless in 6 months” could take my retirement from years to months.
An optimist and their seed funding are easily parted.
Docking software is used to scan millions and billions of drug-like molecules looking for new potential binders. So it needs to be able to generalize, rather than just memorize.
But the evaluation approach used here and in the original paper (1) does not test how well the software will perform on novel molecules, because the test set is related to the training set.
If you understand the basics of ML and physics, you may be interested in my detailed critique here: https://olegtrott.substack.com/p/are-alphafolds-new-results-...
I'm glad that Chai-1 has been released though, as this will probably help people evaluate the method better.
(1) It looks like they are a bit different, as this paper allows 40% sequence identity. It's still high. I believe that sequences with 40% identity tend to have the same shapes, especially in the binding site, where it matters.