AI Update, Late 2020 – Dumpster Fire
blog.piekniewski.info
blog.piekniewski.info
Well, the NP-hardness argument might not be the relevant one, but how useful the structures determined by AF2 are has yet to be demonstrated. Protein folding is very complex; what AF2 has in the training set and CASP in the test set are proteins which were accessible to structure determination up to now at all; most proteins were measured in crystallized (i.e. not their natural) form, so the resulting static structure is likely not representative; and not to forget that many proteins get another conformation than the one to be expected by thermodynamics etc. e.g. because they're integrated in a complex with other proteins and/or "modified" by chaperones; so it would be quite naive to assume that from now on you can just throw a sequence into the black box and the right structure comes out.
This article two clicks away from the referenced one is much more relevant: https://moalquraishi.wordpress.com/2020/12/08/alphafold2-cas...
Even suggesting that AlphaFold should be judged by its ability to "solve protein folding" in the sense of solving an NP-hard problems shows a fundamental misunderstanding of the types of problems that it set out to solve. Granted; even by the standards of the problem that AlphaFold did set out to solve, there is still room for improvement.
From the article the author cites to "diffuse the hype":
> Firstly, there is no doubt that DeepMind have made a big step forward. Of all the teams competing against one another they are so far ahead of the pack that the other computational modellers may be thinking about giving up. But we are not yet at the point where we can say that protein folding is ‘solved’. For one thing, only two-thirds of DeepMind’s solutions were comparable to the experimentally determined structure of the protein. This is impressive but you have to bear in mind that they didn’t know exactly which two-thirds of their predictions were closest to correct until the comparison with experimental solutions was made.
I would hardly becoming the state-of-the-art in a problem a "dumpster fire". Granted, DeepMind got there (at least in part) by throwing tons of money at the problem. Perhaps other methods would have done even better if given the same resources.
Speaking of which, lets look at the other methods. I can't links to actual papers on the CASP site, but by looking at the predictors [1], I believe I was able to find some of the papers by searching.
Second place was the BAKER group by Ivan Anishchenko, Minkyung Baek, and Hahnbeom Park.
> With the recent developments in deep-learning, single-model quality assessment methods have been also advanced, primarily through the use of 2D and 3D convolutional deep neural networks. Here we explore an alternative approach and train a graph convolutional network with nodes representing protein atoms and edges connecting spatially adjacent atom pairs on the dataset Rosetta-300k which contains a set of 300k conformations from 2,897 proteins [2]
This group also got the 3rd place submission as well.
4th place goes to Michael Feig, and Lim Heo
> Here we show that combining machine-learning based models from AlphaFold with state-of-the-art physics-based refinement via molecular dynamics simulations further improves predictions to outperform any other prediction method tested during the latest round of CASP [3]
At least they incorporate some actual knowledge about physics, but their approach still involves a significant amount of AI.
The field of protein folding is dominated by AI even without AlphaFold. What AlphaFold showed is that a well funded team of AI experts can outperform protein folding domain experts. This isn't that groundbreaking of a claim (although certainly speaks well to the generality of current AI methods), but if you want to critize AI methods in general, you should be comparing AlphaFold to the state of the art in non AI methods to protein folding; and I don't even know where to look to find what those methods are.
[0] https://predictioncenter.org/casp14/zscores_final.cgi
[1] https://predictioncenter.org/casp14/docs.cgi?view=groupsbyna...
[2] https://www.biorxiv.org/content/biorxiv/early/2020/04/07/202...
[3] https://www.biorxiv.org/content/biorxiv/early/2019/08/10/731...
But he does make an interesting point about the business of “AI”, a point made by a VC in a well-written post some time last year — it seems quite hard at the moment for ML to scale in the same way that we’ve grown used to software scaling.
But people were saying similar thing about “the World Wide Web” back in the late 90s, and especially after the bubble. Perhaps there is a ML bubble and it will pop, but this is a good thing. It in no way reflects upon the fundamental value of scaling human tasks with software.
Right now ML works fairly well for making a program that accomplishes some discrete and narrow task that could have conceivably been done with traditional hand-coded software, but at greater levels of accuracy and lower levels of human effort. For example, various forms of speech recognition have existed for ... decades? But ML models are far superior than anything that has been or could be hand-coded.
So this space for innovative businesses built upon ML software is as large as the current space for business built upon OOP and FP software — the method is somewhat irrelevant. What matters is still the difficulty of the problem being solved, and in general, the difficulty of making what people want. This is the hardest problem — making something actually valuable. ML software gives businesses a larger space of solvable problems, but it is not unlimited. It might not yet include self-driving cars.
But given time, the only theoretical limit to ML’s problem-solving space is the limit of intelligence itself. Research like AlphaZero is one step toward that holy grail of AGI, and deserve more credit than the author gives it.
Most infuriating is people pretending to have deep knowledge by dumping on technology. It works for the most part because most technologies are over blown. But when the tech is actually revolutionary they just add noise to everything - until most sane people believe completely insane things about the current state of play.
Anyone who thinks recent machine learning advances aren't A MASSIVELY BIG DEAL™ is either emotionally invested in their failure or listening to these doom and gloom puppets.
The author also failed to discuss the many actualv advances in AI this year. Neural nets on 3D point clouds has become much better, image based transformers are doing some fascinating work, massive improvements to training efficiency and scaling models.
2020 from the perspective of AI progress was pretty solid, lots of good wins.
Why would the burn rate make the results unimpressive? The author just made a comparison to MIT, which has a burn rate 3 times as large, and no one questions if their results are impressive or not each year.
It also seems like a cheap criticism of AlphaFold to say it hasn't solved protein folding by citing the NP-hardness of the problem. There's widespread expert consensus that DeepMind made a spectacular advance in the area, and the author characterizes it by saying, "well, they didn't solve an NP-hard problem..."?
The author seems pretty abrasive to be honest. This is a disingenuous "AI Update" for 2020.
Given that this is an initial, barely-functional release, I look forward to what 2021 brings in improvements to the algorithms.
I remember as a teenager in the 00s, I thought robotics works like in the sci-fi movies and roboticists are working on things like Asimov's laws etc. or that AI actually simulates real brains etc.
My point is, some people are naive fanboys, others are bitter know-it-all naysayers, some are hyping salesmen, some are "it's complicated" experts...
I've been paying attention to HN views on AI over many years an it's an inconsistent roller coaster ride. People are one upping each other, try being contrarian without actual expertise, Dunning Kruger etc.
To really appreciate if progress is more or less than expected, you need to have a good intuition for what is expected. Knowledge of the history of the field, knowing what and why is hard or easy, which may not be in line with laypeople's intuition.
Yes, some people oversell it. On the other hand there is real progress. Sure, this is a very uninformative summary.
But it's just ridiculous to see over and over this "where is my flying car, you damn lying researchers" attitude. And equally the Star Trek fan, Tesla fan, Bitcoin fan "how do I start hacking", ifl science cheering crowd who never get their hands dirty. Okay I guess I found a way to feel superior to both.
There's just so much noise and playing telephone through popular press and lay discussions, posturing etc.
The more I do and hear from others about real world AI (through university study and industry) the more I see the disconnect between online blogosphere + confident comments vs reality. I used to be so naive...
And yet : https://blog.waymo.com/2020/10/waymo-is-opening-its-fully-dr...
I guess "the really (sic) AI winter" may have to wait! I commend the power of reading (and grammar) to people who write blogs...
It’s not just about whether it’ll work on day one but about how sustainable the development model is. We rarely if ever manage to produce software without bugs and edge cases. Adding AI to the mix has a somewhat unknown outcome at the moment. I’ve seen some laughable outcomes on simpler models in the finance sector.