Computer chemists win Nobel prize
bbc.co.uk
bbc.co.uk
Martin Karplus' group is behind CHARMM (http://www.charmm.org), which was the first(?) molecular dynamics package (you can think of it as a precursor to folding@home, though it's still under active development, so that isn't a totally fair statement).
Michael Levitt has done a bunch of things, with no one big software package, but he was one of the earliest people trying to do ab initio protein structure prediction. He was also one of the first people to really start categorizing protein structure in a way that allowed for computational modeling -- back in the 70s and early 80s.
Arieh Warshel is probably best known for bridging the gap between quantum mechanics and the (relatively quick-and-dirty) molecular mechanics work. He's done a whole bunch of work modeling enzymatic reactions, coming up with better electrostatic models, and other things where quantum mechanics does a good job, but is way too slow to be used on giant molecules like proteins.
These guys have collectively done so much stuff that it's impossible to give a fair accounting of it all in a brief post.
This Nobel is hits very close to home. Much of my PhD work was done in CHARMM and I come from the Karplus lineage, and at one point in time I had pretty much read every paper Levitt and Warshel had ever written.
It's kind of weird that forcefields haven't improved all that much. Implementation has improved drastically (LAMMPS, Gromacs, etc.) and QM level theory has too (FCIQMC, exploiting symmetries to eliminate the sign problem), but MD forcefields don't seem to have made a whole lot of progress from CHARMM, AMBER, and UFF. I suppose some of the reactive potentials (AIREBO and Tersoff) are a step up, but still...
Actually even for biological molecules there are still advances being made. Water is an excellent example seeing how there is a new potential released once every 6 months or so. Its difficult to say whether these actually show a quantifiable improvement over old potentials, but its still an area of constant effort and progress.
Yeah, that's what I was meaning with regard to force-fields in general. What are we, like TIP23P for water now? ;)
All that aside, Anton is a wicked cool piece of hardware able to accelerate their approximation of traditional molecular dynamics by a factor of 50 over GPUs and up to a factor of 500 over CPU clusters, hitting milliseconds of simulated protein time in a single trajectory.
And since it's the only hardware capable of simulating a single monotonic trajectory out to that timescale, it's challenging to compare it to other approaches.
Similarly, a lot of the folded proteins were used to develop the same force fields now used to simulate their folding. I'm not dismissing this data, but I am saying I think the jury is still out on the models until we have a true test set/training set dichotomy, a separation that was an absolute revelation for ab initio structure prediction algorithms in the 1990s.
That said, I think we both agree that undersampling is the biggest offender. And it only gets worse with results published for larger systems simulated at the same timescale as much smaller ones.
Folding kinetics are surprisingly insensitive to detail. For example, it's been understood for a while that even ridiculously simple models can predict transition state energies for simple proteins:
http://www.ncbi.nlm.nih.gov/pubmed/10322214?dopt=Abstract
http://www.ncbi.nlm.nih.gov/pubmed/10500172?dopt=Abstract
But to answer your question, I'd say that "predicting the correct structure of a protein" is the gold-standard benchmark of forcefield accuracy, and MD forcefields are really bad at it.
(You could reasonably add other great measures, like: "does the simulation tend to fly apart without hacked-up pseudo-physical constraints?", but that feels like piling on.)
What? It's not like the "decade-old" paper has become less correct over time. I mean, it's great that MD is finally close to predicting something that could be predicted a decade ago with much simpler methods, but that isn't saying a lot. Plus, as the other commenter pointed out, there's a complicated cross-validation problem that biases the results you're reporting.
In truly blind tests (i.e. CASP), MD force fields just aren't that good at predicting structure from sequence. The usual counter-argument is that they aren't meant to work with non-MD methods (fair enough, I guess), but you don't have to try very hard to find reasons not to trust them. MD simulations have always been very finicky things, requiring lots of manual intervention to get "right".
I'm glad to see computational research is becoming more popular and accepted. There's a (quickly diminishing) subset of scientists who think computational work is too theoretical, inaccurate, and inapplicable to real-life. This was the case when the field was developed, but it is no longer true.
The exponentially increasing computational power is allowing discoveries that simply can not be performed experimentally because laboratory technology just isn't advanced or capable enough. Who needs to actually study reality if you can just simulate what you need -- and the end result is the same?
I'm not sure many people know this, but our understanding of the laws of physics is advanced enough nowadays to describe almost perfectly everything we observe in everyday life. (Exceptions include things like quantum gravity which don't really matter [ducks to avoid physicists]).
The problem is that if you want to simulate these laws, it requires a lot of computational power. Brute-force approaches are simply ineffective and so simplifications and clever techniques must be developed to reduce the computational effort while giving increasingly accurate results. I think it's the combination of improving computational resources and improving simulation algorithms that are really driving this field.
We calculated heats of formation in our chemical thermodynamics course, but we were always given ample data, and we knew what we were trying to calculate.
If you think you can do the same thing for ethene -> butadiene, you'll find yourself to be horribly wrong, because of extended conjugation networks.
So while for some cases it works, it is not always simple to go from known empiricial results to more complex structures using tables and addition and subraction. And in the case of the molecules I care about, there is pretty good reason to believe that the simple linear methods will fail.
We then built an instrument which could measure the results experimentally. Since this system was so well described by simulation and theory, we used it as a calibration sample before doing real science. Except the results didn't match the computations at all.
I spent years tracking down problems in the instrument (some of which were real) before finally setting in on the fact that the theory and calculations were simply wrong. Just as you're always experimentally limited in what you can measure, you're resource constrained in what you can calculate. Sometimes, you need to drop down another physics abstraction level before you can get the right answer. However, experiments don't care about your abstraction level and get it right every time.
Performing physical experiments is a waste of time and money if you can just simulate the results. However, you don't know if you can simulate it until you've performed the experiment. Every scientist I've known has, at least once, run a simulation and later found that it wasn't even close to the experimental results. Sometimes it's a leaky abstraction in the simulation (e.g. Ignoring the anti-reflective coating on a mirror). Sometimes it's a leaky abstraction in the experiment (e.g. ignoring the pH differences between H2O and D2O). Sometimes it's new physics, but not very often. Sometimes it's exactly like you predicted. Until you leave the keyboard and enter the lab, you won't know.
(sorry for the sheepish comment but this thinking is dangerous; look into some of the current problems in http://en.wikipedia.org/wiki/Quantum_biology)
This is a gross understatement and fallacious. Many of the problems being studied today would take millions/billions/trillions of years of computation in order to model... with approximations. Using even the largest clusters on the planet we can't even find the ground state of even the smallest protein (with ab-initio). DFT scales cubically with system size, and DFT is an approximation to actual first principles calculations that are impossibly enormous and can never be fully calculated.
The number of systems we can accurately model is minuscule, the number of open problems in computational chemistry and materials science is enormous.
I wouldn't go so far as to say that.
A quantum computer can handle the fermion sign problem by scaling polynomially with the number of particles instead of exponentially. Estimates for if/when such a practical device will be created vary wildly but I would think the possibility of this, along with new techniques that take advantage of the redundancy inherent to certain categories of problems and efficiently diagonalize the Hamiltonian could accelerate the rate at which we can handle larger and larger systems. It's kind of hard to predict what breakthroughs will be made, but I'm staying optimistic.
EDIT: Then again, looking at your posts on here, I suspect you already know all that ;)
Bonus points if they observe real world reactions that agree with the model after the simulation has run.
1. You take a lot of empirical observations of known reactions, build a predictive model, then apply it to contexts you don't have empirical observations of.
2. You program in the low level rules of chemistry (e.g. quantum mechanics), and see how the scenario plays out at a higher level (molecular reactions).
I think in their work they are more number 1. A predictive model of proteins without having the simulate the (exceedingly expensive) low level details of QM
You know the main mechanisms at work, but that does not mean you can study all the processes. Sometimes it's just that the sheer size of the mathematical model itself is basically impossible to study without a computer. Oftentimes, though, it's the inadequate formulation of a model that doesn't allow you to simulate it.
A good simulation is an immensely useful tool. There are phenomenons which you know and can simulate, but you cannot easily measure their parameters while they are happening (e.g. you risk disrupting the phenomenon because of your measurement installation). I imagine this is even more so for chemists.
Real-life example from my own research: we worked on tools that helped engineers who designed really high-frequency ICs (think tens of GHz) study things like cross-talk through the substrate. The mechanism itself is basically well-understood, but save for really, really simple structures that are nothing like those in an IC, you can't solve that by hand. Of course, ours were modest achievements, but the point is that this kind of research can open new gates and shouldn't be considered "second-rate".
Also software in chemistry can be used to simulate reactions and processes, which can take years to analyze with conventional methods(if at all possible), like folding protein chains.
http://en.wikipedia.org/wiki/Markov_chain_Monte_Carlo
Similarly, David Baker's lab has focused on producing the folded conformations of proteins in hours to days using simplified models that are designed for that task:
Today is a great day for computational chemistry. It is good to see the detailed Nobel announcement acknowledge other stalwarts of the field like Peter Kollman (who created AMBER, a package similar to CHARMm by Karplus et al). I have no doubt that Kollman would have received this prize had he been alive today.
Chemistry is applied physics. Chemistry answers that H2 and O2 can be joined to form water under certain conditions, but doesn't answer the "Whys"
Physics answers that. And the answer goes around orbitals, energy states, electrons, etc.
Chemistry does answer the whys beyond product yields, why certain elements react is vastly covered throughout organic, inorganic & physical chemistry . It is the basis of these sub-fields, and how this is determined is looking at orbitals, energy states etc...
So, Pauli's exclusion principle is Physics or Chemistry? I'd say it's "both".
I never said reductionism is a one way street.
But in simulations you usually work bottom-up (in abstractions). Of course, sometimes you can't do that, because it's overly complex.
My field (electrical engineering) has a lot of physics as well and I don't mind if those who invented the techniques I use were physicists or engineers (or something else), or if such and such things are in the "engineering" domain or the "physics" domain, I just use them. The point is moot.
TLDR: Approximations are everywhere, make sure you work that into your conclusions.
But the previous generation of scientists came of age when Molecular Dynamics was a disappointing tool because of limitations in the Newtonian models and the lack of computational firepower to do sufficient sampling. The latter has been addressed by a combination of Moore's Law and the ongoing migration of molecular dynamics codebases to GPUs, but the former issue remains - a Newtonian approximation to quantum chemistry.
What's surprising is how much one can get out of these simple models despite these limitations, something that was echoed back in 1976 by one of Michael Levitt's papers that led to today's Nobel Prize:
http://csb.stanford.edu/levitt/Levitt_JMB76_Simplified_repre...
Kieron Burke has a nice take on the matter.
http://www.tddft.org/TDDFT2008/talks/KB.pdf
see also http://www-fourier.ujf-grenoble.fr/~joye/simon2006.pdf, esp. its ref. 93