AlphaFold: a solution to a 50-year-old grand challenge in biology
deepmind.com
deepmind.com
https://news.ycombinator.com/item?id=25253488&p=2
We changed the URL from https://predictioncenter.org/casp14/zscores_final.cgi to the blog post, which has more background info.
https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp...
Now that the problem of static protein structure prediction has been solved (prediction errors are below the threshold that is considered acceptable in experimental measurements), we can confidently answer AlQuraishi's question:
Protein Folding just had its "ImageNet moment."
In hindsight, AlphaFold v1 represented for protein structure prediction in 2018 what AlexNet represented for visual recognition in 2012.
> CASP14 #s just came out and they’re astounding—DeepMind looks to have solved protein structure prediction. Median GDT_TS went from 68.5 (CASP13) to 92.4!!!! Cf. their 2nd best CASP13 struct scored 92.8 (out of 100). Median RMSD is 2.1Å. I think it's over https://predictioncenter.org/casp14/zscores_final.cgi
[0]: https://twitter.com/MoAlQuraishi/status/1333383634649313280
`.cgi`... we've come full circle
Much like other fields, I do begin to question the academic structure to making advances. It appears something is rotten in the state of academia. Oddly it's academia doing incremental improvements to existing methods but industry making novel leaps and bounds... The other major case in point being NLP
A better comparison would be to national labs, but they are tasked with projects that make no sense for industry to tackle.
The system is working as intended, all players are needed. The team at Alphafold busted their chops in academia and went on to working on problems they could spend decades on.
> But in part due to the canonicalization of CASP, protein structure prediction effectively has a two-year clock cycle, where separate research groups guard their discoveries until after CASP results are announced.
and further noting
> As I discussed earlier, it is clear that between the Xu and Zhang groups enough was known to develop a system that would have perhaps rivaled AlphaFold.
Finally, and rather crushingly for your thesis, is the points made about the real industrial groups:
> What is worse than academic groups getting scooped by DeepMind? The fact that the collective powers of Novartis, Pfizer, etc, with their hundreds of thousands (~million?) of employees, let an industrial lab that is a complete outsider to the field, with virtually no prior molecular sciences experience, come in and thoroughly beat them on a problem that is, quite frankly, of far greater importance to pharmaceuticals than it is to Alphabet. It is an indictment of the laughable “basic research” groups of these companies, which pay lip service to fundamental science but focus myopically on target-driven research that they managed to so badly embarrass themselves in this episode.
Modern academia exists to further modern academia, no more and no less. I became disillusioned of any search for truth and progress during my time there, it was really just about showing up the Baker lab at the next CASP and getting the next round of grants secured.
Speaking of which, Google Translate was published in 2006, but when did the "learning from data" approach became an accepted idea in machine translation? I think the earlier attempts at machine translation were more about trying to codify grammar rules in software, than doing statistical learning from large text corpuses? I remember in 2002, the approach of leaning protein substructures from data was already the best performing approach in the protein folding problem.
You have to realize that corporate research labs had a high level of recognition back in the 20th century. Labs like the Bell Labs, the RCA Laboratories, or the IBM Research, privately-funded, had a reputation that met or exceeded the standard of not-for-profit or public-funded academic research institutions. They made some of the most important discoveries in the electronics industry of 20th century, like the point-contact transistor, the MOSFET, VLSI, or the UNIX operating system. They were considered a part of the academia, many scientists were their employees. It's only the 1980s after their death that people had the impression that "important research must come from academia, industry is for incremental changes." So, I'd argue that the division between industry and academia is large, but actually smaller than people's perception. If you consider privately-funded researches by the industry as a part of the academia, the current situation is totally normal, nothing unusual.
Interestingly, for those labs to exist, being a monopolistic megacorp is a requirement. It appears to me that today's FAANG monopoly allowed the creation of Google Deepmind and OpenAI, perhaps it's simply a beginning of the repetition of history.
The article The death of corporate research labs had an interesting review. I highly recommend to read the article:
* The death of corporate research labs
> https://blog.dshr.org/2020/05/the-death-of-corporate-researc...
(HN comment: https://news.ycombinator.com/item?id=232466722)
To summarize, those great labs existed and made great contributions because of (1) corporate monopoly on the industry, and (2) the pressure from anti-trust laws. First, due to monopoly, the gigantic size allowed the labs to be the center of gravity and to concentrate all talents and projects into a single place, with a huge research budget for basic research. Second, the pressure from anti-trust laws also forced corporations to invent more into basic research to grow the business, because mergers and acquisitions were restricted. In some cases, the pressure from anti-trust laws also made the corporate labs to share their discoveries in a more open manner, examples included advances in semiconductor [1], or the Unix source code.
Note: but as HN comments pointed out, somewhat ironically, the success of corporate labs relies on anti-trust pressures, but not the actual monopoly-busting enforcement. The breakup of Bell caused the death of the Bell Labs.
Finally their decline,
> The more relaxed antitrust environment in the 1980s, however, changed this status quo. Growth through acquisitions became a more viable alternative to internal research, and hence the need to invest in internal research was reduced.
And it turns out that managing a corporate research labs without losing money is a tricky problem to solve. If the researches are too goal-oriented, short-termism will dominate, basic research in the labs will be ignored. Thus, basic research in the lab must be independent. However, a lab too isolated from the business can also cause great loss.
> Research in corporations is difficult to manage profitably. Research projects have long horizons and few intermediate milestones that are meaningful to non-experts. As a result, research inside companies can only survive if insulated from the short-term performance requirements of business divisions. However, insulating research from business also has perils. [...] Walking this tightrope has been extremely difficult. Greater product market competition, shorter technology life cycles, and more demanding investors have added to this challenge. Companies have increasingly concluded that they can do better by sourcing knowledge from outside, rather than betting on making game-changing discoveries in-house.
And the author argued the death of corporate labs decreased productivity.
>> An unintended consequence of abandoning anti-trust enforcement was thus a slowing of productivity growth, because the this new division of labor wasn't as effective as the labs:
> a new division of innovative labor, with universities focusing on research, large firms focusing on development and commercialization, and spinoffs, startups, and university technology licensing offices responsible for connecting the two.
> The translation of scientific knowledge generated in universities to productivity enhancing technical progress has proved to be more difficult to accomplish in practice than expected. Spinoffs, startups, and university licensing offices have not fully filled the gap left by the decline of the corporate lab. Corporate research has a number of characteristics that make it very valuable for science-based innovation and growth. Large corporations have access to significant resources, can more easily integrate multiple knowledge streams, and direct their research toward solving specific practical problems, which makes it more likely for them to produce commercial applications. University research has tended to be curiosity-driven rather than mission-focused. It has favored insight rather than solutions to specific problems, and partly as a consequence, university research has required additional integration and transformation to become economically useful.
---
[0] https://www.eetimes.com/podcasts/six-words-that-built-the-ic...
> Honeywell brought a lawsuit against us and said you can’t selectively choose people to divulge your technology to. It’s too important. And if you divulge it to anyone, you’ve got to divulge it to everybody. They filed a lawsuit, and the government came down on their side. And RCA basically had to open up all of its patents to everybody if they opened them up to anybody.
It is not always A versus B.
These fallacies are too common.
This is what it's like when someone really moves the needle. And academic science cannot get it's head around it.
And yet, none of these scientists will suffer any career consequences. Their irrelevant work will be healthily cited by all the other scientists who are doing and have done irrelevant work. They'll retcon a story in their lit reviews about how their irrelevant work led to this.
The career consequences are saved for those who had their eye on the real ball for the last five years, but didn't get there first. For them, the comfortably irrelevant will have the gall to ask in accusatory tones: "What have you been doing these last 5 years?".
In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level.
The Deepmind team did this with ;
"We trained this system on publicly available data consisting of ~170,000 protein structures from the protein data bank together with large databases containing protein sequences of unknown structure. It uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks, which is a relatively modest amount of compute in the context of most large state-of-the-art models used in machine learning today."
So it wasn't out of reach for academia, pharmaceuticals, or others with a bit of resources.
It also looks like they came up with a brand new jiggling algorithm which is probably just V1 now, this really changes things in a significant way!
It turned out only to be one more year instead of four (depending on whether getting to the 90~ range is "solved".
I'm curious to see if AlphaFold can do even better the next two years.
Those last mile percentages always tend to be small anyway.
This seems premature. Even though it does very well on average, there may be some areas where it struggles, and those areas may turn out to be important.
Scores of 30-45 for 15 years. Now scores of 87-92.
This isn't a minor improvement, it's a leap forward.
>a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods
So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.
Which gets into the concept of whether the ML model has actually learned some deeper conceptual ideas than we have, some deeper truth about how this works. If so, can we somehow extract that truth, or is it truly a black box that does the thing we want?
I'm reminded of a sci-fi book I read long ago in which humans are discussing the fact that the science they are utilizing is beyond the scope of a human mind to comprehend- only the AIs can intuitively deal with 12-dimensional manifolds (or something to that extent). Maybe we've reached the doorstep of that future.
While this is an accomplishment, nobody is going to be confusing these models for structures produced experimentally. The CASP metric is for backbone atoms. To have a useful model of protein structure, you really need to have the positions of the protein side-chain atoms modeled correctly. Experimental methods will do that, but this method, as I understand it, does not.
What's an experimental method for protein folding and why is it so good? Are they talking about creating an actual, physical protein in a lab and observing how it folds?
This bodes extremely well for the future of computational biology, I'm very excited thinking about the prospects. If we know how a protein folds, we know its shape, meaning we know which shaped/charged molecules are needed to act as suppressors/enhancers of those proteins.
Does seem like the contest structure could include quite a bit of risk for hiding the effect of overfitting ... I wonder if there is anything inherent about the problem that reduces that risk ...?
One question is how robust the predictions are that DeepMind produces. I would also assume that right now it can't e.g. determine protein structures in the present of other small molecules, or protein complexes. A lot of the interesting stuff lies in the interactions between molecules.
And in general in life sciences any new development will take at least a decade until it hits day to day life, likely even more. We're living with a exception to this rule right now due to the pandemic, but in general things take quite a bit of time in that space.
What an accurate model of protein folding allows us to do, is to take our big database of DNA, predict protein foldings for all of it, and then stand up a search index for this database, keying each amino-acid "row" by the "words" of its predicted protein's structural features.
We could then, with a simple search query that executes in O(log n) time, find DNA targets that produce molecules with interesting structures that might be worthy of study.
This would, for example, be a game-changer in how biopharmaceutical macromolecule-therapy R&D is conducted. Right now we have to notice that some bacterium or another produces some interesting protein, and then engineer a bioreactor to get more of that protein. With this tech, we can work backward from an entirely hyothetical, under-specified "interesting protein", to figure out what catalogued-but-unstudied DNA sequences produce never-before-catalogued proteins that fit that particular functional "shape", and therefore might do the interesting thing. Then we can either directly synthesize that same DNA, or find the organism we originally sampled it from and study it more.
So imagine you gave an iPhone to someone in the 1800's, they wouldn't understand how most of it works, but this may be analogous to them finally figuring out some key aspects of the transistor. So it's another tool in the toolbelt and like all good tools will be used in all sorts of unpredictable ways.
Someone else I'm sure could do a lot better at explaining how important shape is to understanding the function and behavior of proteins.
Of course, there is a caveat. The static, crystallized structure is only one aspect of a protein. The dynamic behavior dissolved in H2O, at different pH, different ionic strength, with different ligands/cofactors are all also important, and not (afaik) directly addressed by this research.
So your second guess is correct - one of the steps is much cheaper now, which marginally improves the entire pipeline. As a result, drugs should now arrive to the market faster.
As a side note, I am curious what happens to the field of structural biology in 10 to 15 years from now. Every research university has a large structural biology department with super expensive Xray/NRM/Cryo-EM machines, and armies of students who routinely spend 4-6 years of their PhD trying to solve a structure of a single protein. If AlphaFold works as advertised, NIH will gradually shift funding to other problems.
(It was predicted that it'd be taxi drivers, not professors, that AI got first. Ironic.)
Back in the 1990s, when I worked on structure data, I remember that at least some crystallizations were easy enough they could be done as a rotation project.
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6287266/ suggests that life is now a lot easier than the 1990s. Quoting the abstract:
> Macromolecular crystallography evolved enormously from the pioneering days, when structures were solved by “wizards” performing all complicated procedures almost by hand. In the current situation crystal structures of large systems can be often solved very effectively by various powerful automatic programs in days or hours, or even minutes. Such progress is to a large extent coupled to the advances in many other fields, such as genetic engineering, computer technology, availability of synchrotron beam lines and many other techniques, creating the highly interdisciplinary science of macromolecular crystallography. Due to this unprecedented success crystallography is often treated as one of the analytical methods and practiced by researchers interested in structures of macromolecules, but not highly competent in the procedures involved in the process of structure determination.
Certainly some proteins are extremely hard to crystallize, and the new single-atom EM work will help a lot. But are there really "armies of students who routinely spend 4-6 years of their PhD trying to solve a structure of a single protein" these days?
I honestly don't know. I'm sure some do. But if so, that army is pretty small compared to the vast numbers who more routinely use crystallography.
Solving protein folding is huge, Nobel in chemistry scale achievement. It would be massive leap for biochemistry.
It seems that Deep Mind solved competition benchmark and made huge leap, but it's just partial solution that works on limited set.
After you have solved protein folding, there is still problem of solving chemical interactions between molecules accurately. Quantum chemistry is extremely compute intensive.
They always had single digit, bizarre artifacts, where the program can't sometimes recognise the very data it was trained on with most minute differences.
Other artifact was that the most "stereotypical cases" were least reliably recognised, and they hot a lot of flak for screwed up live demos, where a radiologist put a very, very obvious tumor shot onto the scanner, and it didn't work without a half an hour of wiggling the film, and a camera.
The "bruteforce" solution may well be always, 80-85% off, but off consistently, and always. NN algo so far beat them, but fail with double digit frequencies on "artifacts" which they themselves can't do anything about.
How well it deals with the later, is what I believe will measure its real world usefullness.
One could say the same thing about programmers automating a task, or a number of other trivial examples. I would lean towards assuming deep mind has competent model validation teams vs. not, even if data science is hard.
So in 5 years you'll see exactly zero new medicines pop up.
(submitted by furcyd : https://news.ycombinator.com/item?id=25254888 ).
> The organizers even worried DeepMind may have been cheating somehow. So Lupas set a special challenge: a membrane protein from a species of archaea, an ancient group of microbes. For 10 years, his research team tried every trick in the book to get an x-ray crystal structure of the protein. “We couldn’t solve it.”
> But AlphaFold had no trouble. It returned a detailed image of a three-part protein with two long helical arms in the middle. The model enabled Lupas and his colleagues to make sense of their x-ray data; within half an hour, they had fit their experimental results to AlphaFold’s predicted structure. “It’s almost perfect,” Lupas says. “They could not possibly have cheated on this. I don’t know how they do it.”
Kudos to the DeepMind team for making magic happen.
And in Nature: https://www.nature.com/articles/d41586-020-03348-4
And about dozen news organizations listed in "CASP14 in news" column in the conference homepage: https://predictioncenter.org/casp14/
Congratulations to the team! This work will have far reaching impacts, and I hope that you continue to invest heavily in this area of research.
For a non-biologist, on what is this skepticism based?
Just purely based on following ML news it looks like the trend for ML solutions has been that they've overtaken expert-systems once they've gained a solid foodhold in a field. Maybe this is some perception bias. Are there any cases where ML performed decently but then hit a ceiling while expert systems kept improving?
It'll be genetics next.
e: although AlphaFold appears to be convolutionally based! I suspect that'll change soon.
>According to Professor Moult, a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods.
So if we go by >= 90 as solved:
>In the results from the 14th CASP assessment, released today, our latest AlphaFold system achieves a median score of 92.4 GDT overall across all targets.
they solved for their targets, but
>Even for the very hardest protein targets, those in the most challenging free-modelling category, AlphaFold achieves a median score of 87.0 GDT (data available here).
They basically admit they still haven't "solved" it for "most challenging free-modelling category"
Take that as you will, not sure how useful the ">= 90 is solved" criteria is since they call it "informal" themselves.
You literally said why it is useful in your comment:
> 90 GDT is informally considered to be competitive with results obtained from experimental methods.
It's informal because we don't have a true "gold-standard" for determining a protein's folded structure – the best we have is experimental methods of trying to determine the structure which still have a great deal of error (compared to other things we can measure).
So all we can do is say "the GDT between two experimental measurements (of the same protein) is often around 90, so if we get there with predictive models that's pretty much just as good".
As soon as we have better experimental methods for determining protein tertiary structure, you can be sure we will require predictive models to deliver better results too. Until then, the point is that the delta between any two experimental determinations of folded structure is approximately the same as the delta between an experimental determination and an AlphaFold guess. So the AlphaFold guess may as well be an experimental measurement. Except the AlphaFold guess happens fairly trivially (once you give it the DNA sequence[1]), where as the experimental method is involved and expensive.
[1] Or the primary structure, I'm unsure what inputs are given to AlphaFold.
What makes them hard to predict is the very close energies involved in different folding pathways. Those close energies mean there will be more variant structures which change by use the experimental approach too.
"We have been stuck on this one problem – how do proteins fold up – for nearly 50 years. To see DeepMind produce a solution for this, having worked personally on this problem for so long and after so many stops and starts, wondering if we’d ever get there, is a very special moment."
--Professor John Moult Co-founder and chair of CASP
"To see DeepMind produce a solution for this" does not imply something is solved. I can produce a bad solution. I can produce a really good solution. All without solving a problem.
Shorter problems are easy to solve. Median score is mix of easier hand harder problems. Next year competition will have new set of much bigger and harder problems to solve.
This seems like a leap, not solved as in having solution that just works and scales.
Semantics. From a systemtheoretical point of view, dynamic folding is an abstraction of static folding; solve (i.e. understand the underlying mechanisms) static folding and you can start progressing on dynamic folding, building up on your previously achieved solution.
Wether it's solved or not depends on wether you mean `general folding` or the `entire spectrum of folding` when considering the problem.
Whether it's good PR or not is to be debated, but it seems that the talent at DeepMind simply can accomplish things other's can't.
It would be very interesting if there was a way to use computational techniques to go beyond what crystallography and other experimental techniques (Cryo) can accomplish and determine the protein structure in it's true biological setting. Some research into experimental methods for this include high power X-ray pulses.
Nonetheless, impressive work!
A few times, I get immense pangs of jealousy for younger people a generation or a half before me. And I'm only 30! This is one of those times.
They surely did unsupervised training on raw data and then fine-tuning on the 170K labelled sequences. I expect the data volume could be increased by orders of magnitude in the next couple of years and we'll see a GPT-3 like jump.
That is not an outstanding question. The test on which DeepMind scored high marks is a test of how well the algorithm folds novel proteins -- proteins whose ground-truth structure has not yet been published.
The article did claim:
> According to Professor Moult, a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods.
So perhaps their score of 87 GDT is pretty significant. But “competitive with” is not the same as “always in agreement with”, as you point out. Could be the failure modes are problematic.
(Am not a structural biologist.)
Added: 2019 article on de novo design: https://www.nature.com/articles/d41586-019-02251-x Not to say that better prediction won't also make design easier -- of course I expect it will.
Protein conformation prediction is essential when engineering new small-molecule drug compounds that must 'dock' with the specific proteins that regulate disease. Knowing how to create a protein with the precise shape to become biologically active has soaked up a lot of R&D funding toward pie-in-the-sky techniques that promise to advance that agenda (like quantum or DNA computing).
If this method works as DeepMind says, it will immediately be adopted by every pharma to assess and tweak the shape of candidate proteins.
A protein is actually a linear sequence of amino acids, but in a cell this sequence has a three-dimensional arrangement like a clew of thread. The arrangement is not random, but dependent on the specific composition of the sequence (i.e. selection and order of amino acids) and some other factors. To understand the function of a protein, we need to know this three-dimensional arrangement (i.e. structure). Up to now the structure determination process was mostly manual, complex, time-consuming (several months up to more than a year) and error prone. If structure determination by DNN is reliable, this is a big win for life science. There are still a lot of problems open: e.g. the structure is not constant over time but there are "moving parts" in the structure which are important for its function.
This is useful because for example chemical simulation tools don't run on DNA code, but on atomic models. And also to produce "images" of the molecules (images between quotes because most proteins are too small to interact with reasonable photons, and no interaction with photons means you can't see them in any way)
DNA has other parts that are really important but we don't understand at all yet, where this doesn't help at all. This applies to sections of DNA sent to ribosomes, to produce actual molecules. Besides that, there are pieces of DNA that "index" the DNA, pointers (from one gene to another), triggers (that for instance start production of an enzyme based on some external influence, like detection of a marker molecule) and export markers (that tell you what to do once the protein is produced, for example, mark a protein to be removed from the cell, incorporated into the cell membrane, or for instance used inside the cell nucleus, and there's also one that essentially says "at this point stop producing a protein and instead couple the rest of the DNA code to the end of the protein you just made").
https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp...
Now I'm waiting for the equivalent news about SETI@Home ;-)
Nonetheless, I'd like to hear more from specialists outside the context of a marketing blog post before I fully buy into a claim of a solution.
There's also a rabbit hole about what 'solution' actually means. Is the performance sufficient for any protein folding prediction application that might arise in the future?
Should we expect to see faster progress in large well capitalized bioscience companies -- or a sudden increase in the viability of smaller biotech and/or biotech startups ...? Are we gonna see top talent fleeing the old biotech companies to start their own ventures with a new belief that the potential for huge reward might suddenly seem achievable?
What kind of companies do we think will be the first that are able to translate this new knowledge into profits?
What companies are doing that work?
I wonder if we can determine if this extends to proteins that aren't as keen to determining their 3D structure?
For example, certain proteins are more crystallizable than others.. For these non-crystallizable proteins, I wonder if we can say that AlphaFold would generate accurate 3D models? And if possible, might there be a way to map out this uncertainty?
This is already happened.
"An AlphaFold prediction helped to determine the structure of a bacterial protein that Lupas’s lab has been trying to crack for years. Lupas’s team had previously collected raw X-ray diffraction data, but transforming these Rorschach-like patterns into a structure requires some information about the shape of the protein. Tricks for getting this information, as well as other prediction tools, had failed. “The model from group 427 gave us our structure in half an hour, after we had spent a decade trying everything,” Lupas says."
From what I gather, training was done on 170,000 Amino Acids (features) and the resultant protein structure (labels). This is out of 200 million possible proteins.
How many examples were in the test set?
EDIT: Looks like N=100 for test set: “Entrants get amino acid sequences for about 100 proteins whose structures are not known” https://www.sciencemag.org/news/2020/11/game-has-changed-ai-...
TODO: 200 million - 170,000 training - 100 test ~= 199.8 million proteins
Better at conversation. Better at making people laugh, and generate attraction or other emotions, better at motivating them, and organizing movements, etc.
Clearly we are not ready for such an efficient system... it would be a big disruption to all human organizations and relations. It would start with Twitter botnets and directing sentiment.
But virtual world, say they're better at math. Say they prove all the Clay Millennium Problems. Say they go way beyond those problems and produce some math far beyond human's ability to understand it.
I've been thinking about that for a while and have decided it's fine. Math as a profession will still exist. Fact is, there's already a proof for everything mathematicians are investigating (or a proof that there's no proof, recursively), out there somewhere. Mathematicians are just searching for it, so that it can be understood and translated to human language. The fact that AI already knows the answers doesn't mean that human mathematicians are useless: they are still required to uncover the meaning of these results and translate them into human language. AI then is still just a tool that mathematicians use to help them in their search. Similar to how biologists will use AlphaFold. I guess.
disclaimer I think the contributions are super useful for science but they do come with worries as does every path of discovery
I think the why is pretty clearly understood (https://en.wikipedia.org/wiki/Protein_folding), in the same way that we understand the mechanism behind the three body problem in physics or quantum computing. But that does necessarily imply that there is an efficient way for us to simulate/predict the results of having nature play out those mechanisms.
The second is that explainability in ML is much more tractable than it was 10 years ago. This is not to say that it's solved, but having solved the predictive problem -- I would expect model simplifications and SME research to proceed more quickly towards understanding the how now. I did some work w/ an Astrophysics postdoc using beta-VAEs [2] to classify astronomical observations, and simplifying models in order to achieve human-explainability proved to not cost as much predictive power as you might expect. It might be that the same holds true here.
This isn't something specific to AI, but science itself. We know the value of C, but now why the value is C, sure we can point to something like the Lorentz transformation, but we can't and probably won't even be able to explain why it has these particular constants, we just know that we can measure them and they are this.
Science isn't in the business of answering why. A successful scientific theory does two things, A) Makes useful predictions, B) Is correct in its predictions. It'd be wrong to call a NN a scientific theory, but it certainly does make predictions and as these results show, it is correct in its predictions.
Sometime soon, humanity is going to have to come to terms that we will soon (or perhaps already have) enter an age where mankind is not the only source of new knowledge. AI-derived knowledge will only increase as the future unfolds and the analysis of such knowledge will likely become it's own branch of study itself.
The same information we get through x-ray diffraction will now be available 100x or even 1000x cheaper, and using this model can even aid the interpretation of xray diffraction data!
What excites me most isn't doing what we can do now, for cheaper (which will surely lead to more effective research methods), but the potential to gain a systematic view of protein structures, either across the genome, species, or through time which will give us a deeper and more fundamental understanding of biology.
Edit: from the other HN article on this topic:
> We trained this system on publicly available data consisting of ~170,000 protein structures from the protein data bank together with large databases containing protein sequences of unknown structure. It uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks
https://deepmind.com/blog/article/alphafold-a-solution-to-a-...
In other words, it's less like GPT3 and more like ImageNet.
> The organizers even worried DeepMind may have been cheating somehow. So Lupas set a special challenge: a membrane protein from a species of archaea, an ancient group of microbes. For 10 years, his research team tried every trick in the book to get an x-ray crystal structure of the protein. “We couldn’t solve it.”
> But AlphaFold had no trouble. It returned a detailed image of a three-part protein with two long helical arms in the middle. The model enabled Lupas and his colleagues to make sense of their x-ray data; within half an hour, they had fit their experimental results to AlphaFold’s predicted structure. “It’s almost perfect,” Lupas says. “They could not possibly have cheated on this. I don’t know how they do it.”
That, on the other hand, makes me feel sad and almost depressed every time:
> It uses approximately 16 TPUv3s (which is 128 TPUv3 cores or roughly equivalent to ~100-200 GPUs) run over a few weeks, a relatively modest amount of compute in the context of most large state-of-the-art models used in machine learning today
Can it take temperature and other environmental conditions into account?
Can you specify that a particular ligand or electrical current is present so that you can see the resultant shape change?
Is all the source code for this available so that other scientists can build on top of this, or will we have to go through a paid or SaaS google API to use it?
Makes you wonder what Deep Learning will tackle next. Factorization of large integers?
This is really an amazing moment.
For example, this[1] is the code for SARS-Cov-2's Spike(S) protein. From what I understand of this page it's pretty short, only ~1,757 proteins (corresponding to ~3,821 bases in RNA).
And here[2] is a visualization of it in 3D. You'll likely recognize the characteristic mushroom shape that's been portrayed in 3D models of SARS-CoV-2 in the media. How does this software work if there's no real way to tell how the protein is arranged?
[1] https://www.ncbi.nlm.nih.gov/nuccore/NC_045512 (search for spike glycoprotein on that page)
[2] https://3dmol.csb.pitt.edu/viewer.html?pdb=6X6P&style=stick&...
Does 84% == solved?
Also, any low hanging frut implications for longevity tech?
I thought the #1 criterion for titles was that they should match the original if at all reasonable...?
The proteins are shown on the CASP website [1]. Both the number of residues and number of proteins are bigger than I expected.
https://predictioncenter.org/casp14/gdtplot.cgi?group=205&mo...
https://twitter.com/TassosPerrakis/status/133353559400213299...
I'm willing to bet it will be staggeringly different than what most people are expecting.
Not really. MSA-based approaches, as most structure prediction methods, have as a goal to find the lowest energy conformation of the protein chain, disregarding folding kinetics and basically all dynamic aspects of protein structure.
> If I write a random genetic sequence (think drug discovery) that has many aligned sequences, without the strong assumption of co-evolution at my disposal, there does not seem any good reason for the aligned sequences to also be proximal.
I don't think I fully understood this, but I'll give it a shot anyway. If your artificial sequence aligns with others, there's a chance that it will fold like them, depending on the quality and accuracy of the multiple sequence alignment. Since multiple sequence alignments are built under the assumption of homology (all sequences have a common ancestor), it's a matter of how far from the "sequence sampling space" your sequence is located compared to the others.
If my speculation is correct, then drug discovery should use a process of genetic programming, using something like this to score the resulting amino acid sequences. I'm wondering if an artificial process of evolution would be sufficient to satisfy the co-evolution assumption here.
Assuming optimistic further progress, what are the implications of accurately predicting protein folding? What are we hoping to discover, or succeed in doing?
It specifically looks at how AI can be used for predictions.
The show immediately came to mind in reading this news. It aired on Hulu.
Short answer: it's extremely computationally expensive
Its better to abstract everything away by a neural net, apparently...
EDIT: Just throwing this out there: Are there national security issues to think about with this? Can it be used to weaponize computational biology?
If AlphaFold gives you a picture of a protein structure, Folding@home shoots a video of that protein undergoing folding.
This is the tipping point where I think the AI singularity may have teeth. I could see math proofs being the next thing to fall. If AI solves the remaining six millennium problems in the next few years, what does that mean for math researchers?
We design proteins for immunotherapies - this kind of thing would help us more rapidly design our proteins (and more efficiently use our wet-lab resources to speed existing projects). For others, some drugs are hard build without knowing how they will interact - this could both provide new 'targets' to go after, but also might help prevent projects that would otherwise accidentally target an important protein.
Maybe they might get an even better model.
[1]: https://www.sciencemag.org/news/2020/11/game-has-changed-ai-...
Can deep learning find the cracks in P vs NP?
Perhaps making clever guesses at prime factors because it learned some weird structural fact that has eluded mathematicians.
If we break crypto, there goes the modern world. Banks, bitcoin, privacy, Internet, the whole shebang.
(I obviously am not an expert in computational complexity and hope that some domain experts can chime in and assuage my fears.)
"[Modern heuristic and approximation algorithms] can find solutions for extremely large problems (millions of cities) within a reasonable time which are with a high probability just 2–3% away from the optimal solution." [0]
When close enough is enough, NP problems can often be solved in P time, and I suspect this is one of those cases. For crypto however, close enough is not enough.
[0] https://en.wikipedia.org/wiki/Travelling_salesman_problem#He...
No. It really is just heuristic building. A core problem with using ML in this sort of use case is that it is often brittle. Once it gets outside of the context it was trained in it may or may not be able to generalize it's training to new contexts. We may have difficulty knowing when it is very wrong.
I think ML in research science could be viewed as a very good intuitive oracle. Even if they are right 95% of the time, you have to do this work prove the long way every time because that 5% matters. The real utility is in "scanning the field" to better focus research on things likely to bear fruit.
But - you don't know that you are 1% from the solution... even if you are pretty confident that you are. It's quite possible (unlikely) that you are way off the optimal, but if you have a decent solution that's ok.
This “uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks”. That is a moderate amount of hardware for this kind of work, so it seems they have a more efficient algorithm.
Also, this algorithm doesn’t solve protein folding in the mathematical sense; it ‘just’ produces good approximations.
Yeah, at least some variations of it are NP-hard. SAT is THE NP-complete problem, but there are some really good SAT solvers around. This basically means: They have a solution that mostly does very well on most instances. But because (probably) P != NP, you will never have a polynomial time algorithm for this.
X has actually solved Y. That's not so much "massively cool", that's historical.
The two are different concepts -- this isn't the typical HN pedantry.
"Solving" the problem would entail developing an interpretable algorithm for taking a string of amino acids and determining the 3D structure once folded.
Approximating a solution would entail simulating that algorithm, which is what their neural network is doing. It is of course usually accurate, but you would expect this with any suitable universal function approximator.
Props to DeepMind and congrats to CASP but is it not obvious that this is more hype-rhetoric for public consumption?
If you can approximate an algorithm with error that is "below the threshold that is considered acceptable in experimental measurements" (to quote another HN comment), then you have something as good as the algorithm itself for all intents and purposes.
Therefore the use of the word "solve" doesn't qualify as hype-rhetoric, and the distinction you're making does seem somewhat pedantic (even if technically true).
(I'm speaking as someone with only the tiniest amount of stats/ML experience, so I could be totally wrong!)
It looks like you'd like a grokable solution, but the problem might be just too complex to grasp for the human brain. "Solved" means they solved the protein puzzles on the official benchmark.
> but you would expect this with any suitable universal function approximator
Yeah, it's just that easy. Function approximator, engage! It took a team of Deep Mind researchers, two years and God knows how much compute. The universal function approximation theorem doesn't also say how to find that network.
Then launches into what can only be recognized as an exercise in pedantry.
This is the absolute definition of it.