CS 522: Machine Learning Approaches to Decode the Human Genome
cs522.stanford.edu
cs522.stanford.edu
Any chance the course has room for the other side of the coin? Namely, how neuroevolution and genetic strategies inform deep reinforcement learning?
Am also interested in learning about the state-of-the-art in cloud based packages. I noticed recently Google released a tool called DeepVariant for use on their genomics platform.
https://github.com/google/deepvariant
Creating a universal SNP and small indel variant caller with deep neural networks
Genetic algorithms and the like are pretty much all terrible. They're ways of approximating your gradient, and fall to the curse of dimensionality. The only reason Uber and open.ai published their papers on evolutionary strategies (something pretty different from what people think of as genetic algorithms) is that current policy gradient methods are really bad as well, allowing what is effectively random search to do well.
It's kinda like how Bayesian hyperparameter optimization is pretty terrible and 2x random search almost always beats it easily.
This may sound like a naive approach, but is there anything special hindering us from building such a simulator or is it just that scientists will not find it useful as it will take too long before we know which nucleotides are important for our goals?
This is hard. To my knowledge we can't even write most basic components of a cell simulator yet. One of the obvious requirements would be a protein folding simulator. Nobody has been able to come up with a working one of those yet.
> when I have a simulator I can try changing them and see what happens?
If you write a simulator that does this you will be a billionaire and probably win a Nobel prize in medicine while you're at it.
If anyone is interested in this work, check their job site. Last I heard they were looking to hire a programmer to help write the sim. Cool team + science.
It was funny to notice that they seem to value a Bachelor's degree at a rate of higher than 1:1 with career experience:
> 4-5 years of experience in a software development team AND a bachelor’s degree in computer science or a related field | OR 15+ years of relevant working experience...
> 6-9 years of experience in a software development team AND a bachelor’s degree in computer science or a related field | OR 20+ years of relevant working experience
That is a pretty heavy premium to put on those years during undergrad. I certainly wasn't as good by undergrad + 5 years as at 15+ years in the field, but maybe they know better than me.
Is there anything better than DFT method for quantum chemistry? You can look it up and see how much effort it takes to simulate just a few molecules.
I'm pretty sure full cell level simulation would revolutionize biology and medicine.
This is so interesting. I cannot imagine what kind of language evolution chose to build on top of DNA. There must be some paradigm it maps to, and it's going to be incredibly interesting to see if a compiler/interpreter can be made for DNA, along with higher level languages that compile down to it.
DNA also contains regulatory segments, errors, “dead” code left over (and carried through generations) long after it stopped being transcribed, DNA once inserted by viruses (both active and inactive) etc etc...
If you want an analogy with computer code: it’s most like a spaghetti-code hair ball of assembly coding it’s own compiler, IDE, vim, a few games, and a neural network ten magnitudes the size of anything tensorflow can do. It does it all on hardware that works only probabilistically. And it is constantly starved for resources, leading to hacks such as DNA sequences that code for two completely different functioning proteins depending on reading it either forward or backward, or starting to read at an offset (what’s called a “reading frame”).
There’s a thick book on molecular biology by Alberts et al. It’s the most phantastic deep dive into this, any many other, insanities. I believe Larry Page used to recommend it to all new googlers.
https://www.amazon.com/Molecular-Biology-Cell-Bruce-Alberts/...
It's standard reading at the very least as intro grad/senior undergrad student in the biosciences.
Great book, well written, well curated.
Parent commenter is correct, DNA->function is massively complicated. The main wiki article to start with is:
https://en.wikipedia.org/wiki/Central_dogma_of_molecular_bio...
As the article notes, there are endless exceptions and edge cases. It links to various examples of those.
Not that I have a book to recommend instead...
The average MBOC provides is enough of a basic understanding of molecular biology to move on to more advanced work including papers that start to get into the nitty gritty details. Note that many people haven't been even exposed to the basics!
Not only that, but it's also (quite literally) radiation hardened: the genome has evolved to a state where most changes to the DNA do very little. Or, if the environment is changing frequently, it'll have evolved to a place in the genome space where it has a higher chance to obtain more beneficial mutations (for more info on this sort of stuff, look up articles by Hogeweg, Colizzi or Crombach).
As a computational biologist, I can tell you that evolution works in many ways, but most likely not in the way you'd expect it to.
Now, programmers, by our neurophysiology and training recognize things that look like abstractions in what evolution produces, and near-abstractions are useful to reduce our cognitive load when we're studying biology, but a biologist always keeps in mind a list of exceptions to those abstractions in mind. If you don't know of exceptions, it is almost always a fruitful research question to find them.
https://www.damninteresting.com/on-the-origin-of-circuits/
tl;dr - evolution takes advantage of the entire solution space without any respect the the abstraction layers we've created in our minds.
> after just over 4,000 generations, test system settled upon the best program. When Dr. Thompson played the 1kHz tone, the microchip unfailingly reacted by decreasing its power output to zero volts. When he played the 10kHz tone, the output jumped up to five volts.
> Dr. Thompson peered inside his perfect offspring to gain insight into its methods, but what he found inside was baffling. The plucky chip was utilizing only thirty-seven of its one hundred logic gates, and most of them were arranged in a curious collection of feedback loops. Five individual logic cells were functionally disconnected from the rest— with no pathways that would allow them to influence the output— yet when the researcher disabled any one of them the chip lost its ability to discriminate the tones. Furthermore, the final program did not work reliably when it was loaded onto other FPGAs of the same type.
> It seems that evolution had not merely selected the best code for the task, it had also advocated those programs which took advantage of the electromagnetic quirks of that specific microchip environment. The five separate logic cells were clearly crucial to the chip’s operation, but they were interacting with the main circuitry through some unorthodox method— most likely via the subtle magnetic fields that are created when electrons flow through circuitry, an effect known as magnetic flux.
This nicely illustrates a major advantage of evolutionary processes: they can use any resource in the environment, whether you know that resource exists or not.
The program's crippling overspecialization ("the final program did not work reliably when it was loaded onto other FPGAs of the same type") is also typical of evolutionary processes.
To answer your question, if you have your genome and a dataset of known genomes marked with functions according to regions, then you could probably perform an interesting analysis..