In a First, an Entire Organism Is Simulated by Software
nytimes.com
nytimes.com
"For their computer simulation, the researchers had the advantage of extensive scientific literature on the bacterium. They were able to use data taken from more than 900 scientific papers to validate the accuracy of their software model."
This is not (necessarily) a model of M. genitalium - it's a model of our understanding of M. genitalium, and as such incorporates in everything that current technology in biological sciences allows us to look at. It takes a huge amount of data from many sources and tries to bring it together on a scale not previously done. However, that data may have significant flaws, biases and ultimately considers only what we're good at looking at (technologically/scientifically speaking). It's awesome, and 100% the right direction for the field, but equally it is not a, "synthetic life form being simulated" as much as, "a very very complicated model which uses huge amounts of multidimensional data to try and replicate the behavior seen in that data".
The exciting thing is that with this model, they can now rapidly iterate between the behavior of the model and the behavior of the real organism, see where the gaps in knowledge are, and work to fill them in.
(I'm happy to see M. Genitalium getting attention, since I made a Java genome display for it 14 years ago http://www.righto.com/java/genome/MG.html)
Models can be predictive, indeed that's one of the main reasons to create a model to observe behaviour that otherwise might not have been expected.
A theory in physics for example is a model with proven predictive powers, the model [theory] 'knows' some aspects of what can be observed before those things can be confirmed empirically.
But as you indicate new empirical data can shows flaws in a model just as it can in a scientific theory.
Also, and this is something I personally feel very strongly about, the paper is incredibly well written and accessible. There is no point making science difficult to understand. Science is (itself) not difficult - it's complicated, encompassing, and confusing, but nothing is actually difficult. There is a trend to use big words and complex language in scientific papers, which I worry must be incredibly off-putting for those less used to them, and are almost always totally unnecessary. This paper was just beautiful (I devoured it on my subway ride home) and certainly one I'll be recommending to friends interested in the systems biology field.
The best representation of a thing is that same thing in the state you're trying to represent. The cat is not a model, a dead cat is not a great model of a live one.
They develop a number of submodules that each use an appropriate algorithm or branch of mathematics to best represent the interaction. To integrate them, results from the previous time-step feed each model as appropriate. The derivation of the correct mathematics and algorithm to use and the best way to connect them is anything but measuring a bunch of sticks.
Plus they were able to do some predictive modelling and pointed a direction for a novel result. From the abstract: "experimental analysis directed by model predictions identified previously undetected kinetic parameters and biological functions." This is not about simulating organisms, it is for accelerating experimental discovery of interactions, dynamics, parameters etc. This is a proof of concept that a technique which may one day contribute to saving lives is viable.
Awesome, maybe. But "100% the right direction for the field" is a stretch. This is press-release bait. What hypothesis did it address? What theory is it advancing? (other than "we theorize that we can simulate a cell in a computer...ohlook, we did it!", of course.)
This is interesting from a technology perspective, perhaps, but it's hard to call it science.
As a physicist and scientist myself, I feel extremely inclined to say 'potato, potahto' at this one. This technology allows us to test many more hypotheses at a much quicker pace, which is good enough for me to call it science.
I think you're seriously overestimating the quality of the model. This isn't particle physics -- it's not as if there's an equation that precisely predicts the results of any particular molecular interaction. We don't even know if we know the full set of possible interactions in a system of this size. How can we possibly simulate it?
This is an example of "we ran the fancy machine for a while, it spit out some data, and we cherry-picked 'interesting' results from the output. Some of them even show up in experiments!" These sorts of breathless announcements are common in computational biology, but they usually don't amount to much. The best work is extremely reductionist. As others have already noted, it's easy to spend 10x the computational resources simulating a single protein on timescales far, far shorter than the ones simulated here.
(Edit: there was a whole bunch of stuff here, but it's pretty much irrelevant given that we seem to agree on the basic points.)
"This is an example of "we ran the fancy machine for a while, it spit out some data, and we cherry-picked 'interesting' results from the output. Some of them even show up in experiments!" These sorts of breathless announcements are common in computational biology, but they usually don't amount to much. The best work is extremely reductionist. As others have already noted, it's easy to spend 10x the computational resources simulating a single protein on timescales far, far shorter than the ones simulated here."
That's fair enough. I checked your profile and saw that you have a PhD in computational biology, which is not my field of expertise, so I'm inclined to take your word on that one. In that case, my comments do, of course, not apply.
Now there may be a question as to whether we can ever truly test a hypothesis against reality or if we're stuck with models ...
This is the main point of interest for me personally.
Currently it takes about 9 to 10 hours of computer time to simulate a single division of the smallest cell — about the same time the cell takes to divide in its natural environment.
So, it's a realtime simulation! Cool, although wholly incidental.It's hard to tell from the article exactly how fine-grained the simulation actually was. Was it actually at the level of tracking individual molecules, or did each "cell object" just keep track of, e.g., concentrations of different types of molecules?
Also, I wonder if this will have any long-term impact on things like the Open Worm project ( www.openworm.org ).
I agree modelling entire organisms will ultimately require large collaborations, potentially through projects like Open Worm.
Source code and training data available online; written, of course, in MATLAB. Very refreshing, and I'm looking forward to dissecting this first-hand.
Having the code in Matlab seems like a disaster as far as ever making this or similar approaches modular and so usable-by-others goes. And I do know by miserable experience that Matlab indeed what biologists generally use but if biology is ever going to interface with larger scale software construction, it seems like it is going to have to change it's standard operations a bit.
Edit: And this isn't saying Matlab is generically "horrible". It is great at what it does but horrible from my perspective, as a programmer whose task usually is putting pieces of software together.
I think MATLAB itself isn't so bad, so much as the way its typically used -- poorly commented, organized, and untested. We put a lot of effort into clearly commenting and organizing the code. We used matlab-xunit and hudson to manage testing, which worked quite well. http://www.mathworks.com/matlabcentral/fileexchange/22846-ma..., http://www.mathworks.com/matlabcentral/fileexchange/33971-xm....
> We're starting to move toward python for future work.
Thanks a lot for commenting. As somebody who does a lot of scientific coding in Python and part-time contributor to a couple of projects, I'm of course curious -- what's your reasoning behind the move and where do you anticipate to see the most friction?I don't see any friction, except if one needed to port old code. It seems that a lot of people are moving toward SciPy/NumPy these days.
> the code in Matlab seems like a disaster as far as ever
> making this or similar approaches modular and so usable-
> by-others goes.
As always, it depends. I've seen very well maintained MATLAB code bases, and I've seen the opposite (with the latter greatly outweighing the former). We should give these guys the benefit of the doubt. Somewhere Karr mentions Hudson CI, so they don't seem fully removed from good practices. Interfacing with MATLAB from C is reasonable.In an ideal world, this would be a NumPy/SciPy prestige project, but neither the community nor Python for that matter are quite there yet.
Consider that if biological simulations are going to go to a larger scale, you won't want to simply call a bunch of Matlab simulations from a single C program but rather have a bunch of distinct programs that would be modified to call each other (with each of these programs running as they do now in their Matlab instance).
I'd also like to know how you think numpy/scipy are short compared to Matlab. I've not run into their limitations yet.
*XXXXXists make huge breakthrough on YYYYY*
"I can't believe they used programming language Z!"
Domain specialists generally have more market leverage than the random programmer geek that scoffs at their "primitive" tools, and the code they write is there to do a job, not satisfy the peculiar preferences of a specialist in some other domain (like programming).Now, when programmers work to make domain specialists' job easier, that gets some attention; thus the tidal levels of interest in MDA, DSLs, and so on.
http://www.sciencedirect.com/science/article/pii/S0092867412...
Thanks, Elsevier.
The further question with any project for full organism simulation would be how many lines of code are going to be produced and what would the process of maintaining that code look like?
You're probably right about the hierarchical simulation though. I have trouble believing that it's actually useful to model an entire body at a macro-molecular level. It would be like modeling electrons in circuit design software, rather than using the abstractions of voltage and current.
Human-designed systems need to be human comprehensible, so they're laid out in neat hierarchical layers, each abstracting away the complexity of the underlying process. Evolved systems aren't restricted by comprehensibility.
http://www.damninteresting.com/on-the-origin-of-circuits/
>Finally, after just over 4,000 generations, the test system settled upon the best program. When Dr. Thompson played the 1kHz tone, the microchip unfailingly reacted by decreasing its power output to zero volts. When he played the 10kHz tone, the output jumped up to five volts. … And no one had the foggiest notion how it worked.
>Dr. Thompson peered inside his perfect offspring to gain insight into its methods, but what he found inside was baffling. The plucky chip was utilizing only thirty-seven of its one hundred logic gates, and most of them were arranged in a curious collection of feedback loops. Five individual logic cells were functionally disconnected from the rest– with no pathways that would allow them to influence the output– yet when the researcher disabled any one of them the chip lost its ability to discriminate the tones. Furthermore, the final program did not work reliably when it was loaded onto other FPGAs of the same type.
I think voltage etc is airtight abstraction mostly because each electron is guaranteed to be both simple and the same. Cells are both complex and distinct from each other (based on both genetics and internal physiology and so-forth). So macro-configuration of cells would seem to be a more leaky abstraction.
Roger J. William's classic text Biochemical Individuality describes how much the parameters of even very basic physiological functions varies from person to person. So unlike a chip which starts with simpler building blocks and is designed to depend on discreet inputs as much as possible, the simulation of an organism may not have a better solution than a bottom up design with perhaps a variety of clever shortcuts.
The point is that circuits can be simulated at many different levels of detail, with low-level simulations more likely to get things exactly right, but with high-level simulations much faster, easier to write, and easier to understand.
Likewise, it will be interesting to see with cell simulations how much complexity can be abstracted away and still have a useful simulation. To get all the protein interactions right for example, you'd need to simulate individual atoms, which is insanely slow. So to simulate a cell, you're probably running at the level of protein and chemical concentrations and known interactions, which is faster but introduces error. For instance, how much do local concentrations matter? And to simulate a multicellular organism, you're probably going to make the cells fairly abstract.
Personally, I think the key area for biology is going to be dealing with cell state. Cells hold state in a lot of different ways over many time frames (eg epigenetics), and I think computer scientists have a lot to offer biologists in understanding state. Someone else mentioned the few hundred cell types in the human body, but the internal state makes a huge difference. (Not to mention distributed state, such as how the brain stores information.)
Sort of like how image compression algorithms store an abstraction of the image for large, uniform areas and get more granular for complexity, but applied to simulation instead
Even when we look aside from the level of DNA/RNA there are huge differences in morphological organisation of eukaryotic cells when compared to most prokaryotes: dynamic compartmentalisation of cytoplasm, different types of cytoskeleton, vesicle trafficking, complex signal transduction networks instead of usually simple two-component regulatory systems... So the simulation of whichever human cell type could be much more complicated than one could initially thought.
I don't want to sound too much pessimistic, as someone with background in both CS and molecular biology I'm truly excited about this, but I still had to cool myself down a little bit after reading the article. I can't wait to read the original paper.
This is only one example of many cellular processes which are substantially harder to understand in eukayotes than bacteria. The lesson is that not all cells are created equal, different kinds of organisms have very different cells.
Also note that modeling interactions between cells makes things much more complicated in a very nonlinear way.
Sort of like a Hashlife[1] for real life, as it were.
And let me just say that I sure hope we're not still coding in 80 years. Fun as it may be, we'd probably be better off letting machines handle that sort of thing. Get started on that AI!
The computational complexity aspect could be side-stepped if quantum computers were invented before 80 years. One of the few things quantum computers can do exponentially better is simulate quantum mechanics. Biology is chemistry which is "just" a bunch of quantum many body problems. With a QC we can model Protein interactions, RNA, signaling, transport etc much more efficiently than a regular computer.
You could imagine a hierarchy of simulations which dynamically adjusted the level of detail. Do you simulate every cell division to see if cancer develops, or do you just do "perfect" division with a roll of the dice for a copy error? Depends on your needs.
You could imagine organs like your skin or liver have lots of redundant computations, where you can elide away details more easily. While say your brain probably depends on the idiosyncratic behavior of individual neurons bubbling up to a macroscopic level, requiring a more detailed simulation.
Having multiple sets of same cell would only speed up the simulation under some circumstances. In the most general case, it wouldn't. For example, if a cell itself is a Turing complete computer whose computations matter, then you definitely need to individually simulate each cell.
I can't even imagine how much parameter tuning and hacks went into a model of such staggering complexity. Paraphrasing one of my old academic advisors, the curse of models is that you can always make them look good.
http://wholecell.stanford.edu/ contains a link to the code (written in matlab)
Imagine if you could get the DNA from a cancer cell in a human patient, as well as the DNA of a normal cell, and then test the effect of a million different randomly-generated molecules until you find one that kills the cancer cell, but not the normal cell.
If you could scale the performance of this type of system and allow it to simulate Eukaryotic cells (much more difficult) it might let you cure most cancers!
Downloaded code zip file and unziped to to "WholeCell" directory
find WholeCell -type f | awk '{FS="."; print $NF}' | sort | uniq -c | sort -n
...
2 java
2 lib
2 log
2 mexa64
2 mexw32
2 pdf
2 swf
2 vbs
2 xlsx
3 exe
3 fsa
3 mexw64
4 jpg
4 svg
5 desktop
5 license
5 sql
5 TXT
6 col
6 sh
6 tmpl
6 tpl
7 ico
7 xml
8 bat
8 pl
8 z
9 json
15 dot
16 gif
18 map
19 css
23 dll
24 dat
25 p
52 mat
62 txt
114 jar
238 png
277 js
427 php
531 m
1096 htmlKeeping in mind that my knowledge is zilch, here are some exotic languages that appear to be designed for this sort of thing:
http://en.wikipedia.org/wiki/X10_%28programming_language%29
http://en.wikipedia.org/wiki/Chapel_%28programming_language%...
http://en.wikipedia.org/wiki/Fortress_%28programming_languag...
https://simtk.org/project/xml/downloads.xml?group_id=714#pac...
I went to a talk by Dr. Covert a few months ago and it was fascinating, all the moreso as a chalk-talk with an animation flip-book handout.
As an aside, since the topic of CS folks helping with biology research comes up often: apparently a Google engineer went on sabbatical to the Covert lab and helped re-architect and optimize the system to the point that it was feasible for this research. There are many other projects that would benefit from this type of expertise.
So in this case, it sounds like they're simulating the interaction of genes and molecules, since they think that's sufficient to model cell behavior (and/or it's the best we can do). But it doesn't really matter what technical level of detail they went to -- the only useful definition of a "proper" simulation is whether it behaves the same as the real thing in the context you care about. For example, this simulation would be totally insufficient if I wanted to model a hydrogen bomb -- but totally excessive if I wanted to model gravitational forces on independent objects in space. If it's good enough to tell us anything new about the actual cell, that'll be pretty cool.
In my experience, physics researchers are not familiar with GitHub or similar OSS hosting sites. The omission of a Github project is likely through ignorance or apathy, rather than disapproval. I'd be happy (delighted!) for you to take my own physics simulations and play with them, for example.
For one thing, if the protein interactions are abstracted and what's chewing things up is the modeling of them, then it would make more sense to use a lot of powerful GPUs; then you could spread the calculations out over a bunch of Xboxes.
Secondly, despite the huge amount of data generated, you should be able to monte-carlo outputs for inputs on a thing like this. In essence, use it to generate a rainbow table of states, inputs and outputs so you can then run a multi-cell organism much more quickly. Assuming every cell has the same DNA.
Similarly, the simulated bacterium is a series of modules that mimic the various functions of the cell.
Pretty cool that a relatively mass-market outlet like the Times thought it was worth mentioning OOP. Even as a software developer, it's a stretch for me to visualize what it means to simulate an organism; what a difficult job for the Times to distill it down into a couple of paragraphs for a lay audience.
Audacity, yes we do need more of that. pg says that tenacity if the single most important trait of an entrepreneur. Maybe audacity is a multiplier that brings us the big innovations.
Though it's probably getting close to the point where getting the right algorithms is the bottleneck; if you could somehow bring the techniques and optimizations we'll inevitably learn over the next 100 years back to today, a half decent simulation would probably already be possible on a cluster of commodity hardware.
Unless the techniques revolve around a new method of computation, of course - maybe memristor logic helps with that kind of thing.
[1] http://bionumbers.hms.harvard.edu/bionumber.aspx?s=y&id=...
What if we're just living in a giant simulation? What level of computing power would it take to run the planet Earth?
I gotta say, if I was going to be working on that code, that statement would make me rather uncomfortable.
Owned by U.S. Navy, issued 1987 so it's free now.
I am really just scratching the surface of what was covered, but it was a great book and I would recommend it for anyone interested in simulated life.