In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level.
The Deepmind team did this with ;
"We trained this system on publicly available data consisting of ~170,000 protein structures from the protein data bank together with large databases containing protein sequences of unknown structure. It uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks, which is a relatively modest amount of compute in the context of most large state-of-the-art models used in machine learning today."
So it wasn't out of reach for academia, pharmaceuticals, or others with a bit of resources.
These years of research involved trying many different architectures, many of which received as much or more compute time than the final system.
The price of training the final architecture is meaningless. Researching and training AlphaGo was expensive but it enabled the ideas and development of AlphaZero which is more computationally tractable.
To have any chance, an academic team would need the same compute resources as what the DeepMind protein folding team used during the whole development of the architecture during the last few years, not only the resources used to train the final system. And I bet this funding is not available to most if not all academic teams.
The research is the giant shoulders you stand on, the compute cost is the price of the tool you need to do the present-day work.
Both are relevant but the shoulder’s of giants are generally more accessible, particularly if we’re talking about published research and not proprietary tech.
A competing team is not starting from the same place the DeepMind team started at 5 or 10 years ago.
> I don’t think we would do ourselves a service by not recognizing that what just happened presents a serious indictment of academic science. There are dozens of academic groups, with researchers likely numbering in the (low) hundreds, working on protein structure prediction. We have been working on this problem for decades, with vast expertise built up on both sides of the Atlantic and Pacific, and not insignificant computational resources when measured collectively. For DeepMind’s group of ~10 researchers, with primarily (but certainly not exclusively) ML expertise, to so thoroughly route everyone surely demonstrates the structural inefficiency of academic science. This is not Go, which had a handful of researchers working on the problem, and which had no direct applications beyond the core problem itself. Protein folding is a central problem of biochemistry, with profound implications for the biological and chemical sciences. How can a problem of such vital importance be so badly neglected?
In short, academia got utterly schooled by a small group at Google spending a relatively small dollar amount on compute, using techniques that in hindsight are fairly described as "simplistic". There's no way around it.
If you were trying to get across the Atlantic, this would be like getting upset at a group of bridgebuilders for trying to solve the problem by building a bridge across instead of by inventing the airplane. The approaches are that different.
>The approaches are that different.
I'm not sure if that analogy applies here. DeepMind wasn't the first group tackling structure prediction with machine learning. Their success lies in the innovations that they implemented (predicting interresidue distances as opposed to contacts, for example).
Meaningless in historical terms, but meaningful in future terms. It's meaningless how long the training took because there were countless resources spent to get to that point. It's meaningful in the future, because we know that training times are fairly short, and iteration can be done fairly quickly.
That said I think the academic and pharma communities had engineered themselves into a corner and weren't going to see huge gains (even thogh they are exploring similar ideas) for a number of banal reasons.
What I meant to focus on was that I think DeepMind has less of a pure money/scale advantage in this area than in some others. In something like Go or Atari game-playing, there are many academic groups researching similar things, but their resources are laughably small compared to what DeepMind threw at it. So you might argue that they got good results there in part because they directed 1000x the personnel and compute at the problem compared to what any academic group could afford. In biomed though, their peers in academia and industry are also pretty well-funded.
Here's an exemplar of how I think it evolved well in a cloud world: https://gnomad.broadinstitute.org/
that project adopts many concepts from google and others and greatly improved our analytic capabilities for large-scale genomics.
How much does hiring a deepmind-like team cost though? (massively more than the TPU resources?)
Still within reach of pharmaceutical industry I guess, but maybe not so easy for academia.
And they had income around 100 million in 2019 but it's all against Google, so looks like a 2 billion +/- 0.5 operation so far, and who knows if they pay for compute.
Other articles place the runrate at 500 million per year in 2019.
Which means 500 million * 6 years = 3 bn + 0.5 purchase price. = 3.5 bn. So somewhere in the 2.5 - 3.5 billion range its seems likely as total cost so far.
Nevertheless doesn't seem out of reach for a multinational.
Remember, we are looking in hindsight that it seemingly paid off. A few years ago, this was just an educated bet; only the richest companies with money to burn (from selling ads) would be willing to take on that kind of a risk.
There is barely any multinational that has the freedom Google had of planning to spend 3.5B with no ROI. Their shareholders would sue and vote the managers out.
Excited to see how follow on use of these models, by many more teams, researchers, and companies plays out over the next two decades.
This is a foundational advance!
Much like political dictators, they can be exceedingly efficient and have resources (and authority) to do things in spite of opposing interests.
People who faced with the narrative that countries have a monopoly on a number of aspects of life find monopolies are not a BAD THING(tm), but that they are bad for a consumer market - as a monopoly eventually blockades aspects of the market.
I feel like that is not too far from saying it makes one reconsider communism because good things can happen with authoritarian control.
Big companies can suck up all the air in the room by monopolizing talent and making it harder for startups to pay the kinds of salaries needed for top tier AI research. Xerox PARC came up with all kinds of groundbreaking inventions that were never commercialized (by them). For every invention that comes out of a big company, it's worth thinking about whether it might have actually come out faster if it was borne of competition instead of a side project. Or in the grand scheme of things, if corporate taxes were higher and the money was given to a university research lab.
I think the best results may come from the middle ground. Smaller/medium companies are so worried about staying afloat or hitting their quarterly earnings that they have trouble making long term investments. Large companies are diverse and profitable enough that they can afford to blow money on things that might not pan out, but they don't have the same drive -- and in fact have some pressure to avoid being "too" innovative because it could cannibalize their existing products.
Most people just show up to vote once every 4 years (or less) and make their decision based on the party affiliation or the wedge issue du jour, and the rest of the time pretty much ignore what's going on or don't have the power to do anything about it, which gives a lot of leeway for special interests to slide things in under the radar.
See https://ilr.law.uiowa.edu/print/volume-100-issue-5/all-i-rea...
This was a/the "holy grail" problem of molecular biology, long thought to be an automatic Nobel. It's somewhat unfair to characterise developments prior to this as insignificant. In fact by the time I was working on it, that "automatic Nobel" was no longer assumed, because the field had made quite a bit of progress, in many tiny steps by many different groups, and the assumption was it would continue in this slog until reaching some state of sufficiency for practical applications without ever seeing the sort of singular achievement that would be worthy of praise and prize.
Far more went into this breakthrough, obviously, than those TPU-hours: the development of those TPUs, for example, and assembling a team that can make use of them. The protein folding problem requires very little knowledge of biology or physics to understand and was always pre-destined for some outsider to sweep. Indeed, there was game that allowed people to solve structures by intuition alone, and, IIRC, some 13-year old Mexican kid cleaned everyone's clock some years back.
Why didn't some research group do this first? Most of them just don't have the budget. We were five people, total, IIRC, and felt pretty rich because we were computer-people getting the same budget for materials as everyone at our institution, which was all wetlab, otherwise. So I was a student being paid $20/h but with a $50,000/p.a. hardware budget. How many false start does it take before you do that run with 128TPUs "for a few weeks" that works? If you blow your budget on one gigantic Google invoice, what's going to happen to you when it doesn't pan out, and the whole institute laughs at you? Etc...
There are quite a few rather good things this problem has inspired over the years, though. Among them is CASP itself: the idea of instituting a yearly competition that gives unequivocal feedback on the state of the field and every group working on it is rather rare, I believe, and it's been successful. Indeed, it would seem that CASP was necessary to attract outside groups like Deepmind, i. e. deep-pocketed industry groups striving to prove themselves on a clearly defined problem. Chess, Jeopardy, CASP: maybe it would be worthwhile to explore not <solving x>, but <stating X as a problem that attracts Google/IBM/etc.-scale money> as a superior strategy in some cases.
There was also folding@home, pioneering the distributed-donated-computing model, and the aforementioned gamification of the problem, and hundreds of the most intricate, custom-tailed, more-or-less insane ideas people devoted months and/or careers and/or careers of their most promising post-docs to that didn't pan out.
Like cellular automata. They don't work for this, trust me. (Great hit for interactive poster sessions, though)
This is a big issue that most people miss. Having easy access to vast computational power makes such a difference for experimentation.