Introduction to Genomics for Engineers
learngenomics.dev
learngenomics.dev
https://www.biostarhandbook.com/
I have learned so much from it.
It is an introduction into what is like to do genomics in a scientific environment. The content at the link the OP posted appears to be an oversimplified, high level and naive overview
https://www.edx.org/course/introduction-to-biology-the-secre...
by Professor Eric Lander
> Introduction to Biology - The Secret of Life
> Explore the secret of life through the basics of biochemistry, genetics, molecular biology, recombinant DNA, genomics and rational medicine.
It's really well done and genomics is the focus. I took many dozens of edX and Coursra courses over the years, this is one of the top 5% of the courses there I would say.
I don't understand the phrase "from a programmer's perspective", or "for Engineers" in the title on top.
As a programmer whos studied CS but also took numerous life science courses throughout my life. You want to learn biology you study biology, what does a "programmer's view", or an engineer's, have to do with it? You use the correct tool for the job, and having a background in both, I don't see this working out well, more like the opposite actually.
The point of looking at biology for an engineer or programmer should be to broaden ones horizons, not to use ones internal models build for a completely different field in another one that really is not like that at all. IMO it's best to forget all computer metaphors here.
----
By the way, since there was something about this yesterday, there also is this course: https://www.edx.org/course/principles-of-biochemistry - it too is very good. A good knowledge of organic chemistry is a prerequisite, but there are plenty of equally interesting course resources for that available too, including even Khan Academy (https://www.khanacademy.org/science/organic-chemistry), or to give a(nother) random link, https://ocw.mit.edu/courses/5-12-organic-chemistry-i-spring-...
Biology becomes a lot more fun with this foundation already established in ones head.
1) The Moderna vaccine was made with the help of illumina genome sequencing. They were able to sequence the virus and send that sequence of nucleotides over to moderna for them to develop the vaccine - turning a classically biology problem, into a software problem, reducing the need for them to bring the virus in house.
2) Illumina has a cancer screening test called Galleri, that can identify a bunch of cancers from a blood test. It identifies mutated dna released by cancer cells. This is huge, if we can identify cancer before someone even starts to show symptoms, the chances of having a useful treatment dramatically go up.
Disclaimer: I work for illumina, views my own.
I wrote some more about why genomics is cool from a technical point of view here (truly big data, hardware accelerated bioinformatics) : https://dddiaz.com/post/genomics-is-cool/
"The Amazon CloudFront distribution is configured to block access from your country."
Having Turing complete programmatic control over biological systems has an absolutely endless list of transformative applications.
Imagine being able to program bacteria that can "infect" the patient and attack tumor cells, or act as fodder to keep autoimmune disease in check.
Or let's say we could program stem cells into "liver repair mode" to go and differentiate into new liver cells.
Then the implications for things like drug synthesis with the ability to programmatically control enzyme levels to compile more or less arbitrary biosynthetic pathways into fast growing photosynthetic algea, turning CO2, water and sunlight into medicine.
It's still a long way off being at that level of applicability, but man oh man it's gonna change everything.
Imagine a software heisenbug, but instead it's a life form that you can't kill -9.
The idea of tailor-made medicines in a vat is awesome, but as far as creating a bacteria to "specially target" certain cells seems like a disaster waiting to happen.
For instance, it might be possible to use ECC to get around transcription errors. It could also perhaps be ensured that any rogue "clinical biocomputer" could be easily treated with antibiotics or specifically engineered bacteriophage virus.
Like I said, the technology is very far off from having real world applications like this. At the moment it feels like we're in the analogue of the 40s and 50s for conventional computing. The field is still just inventing the very basic building blocks. It's going to be very limited in use, wildly dangerous(look up mercury delay lines) and unreliable for decades to come.
Chemo/radiation only kills the patients it was given to ( not entirely true - if the treatment caused mutations in the germline and the patient subsequently had children, the effects of the treatment might be passed on - but still very limited ).
Bacteria are the scariest when they've had the time to develop resistance to multiple different antibiotics.
Additionally, a bacterium that's engineered to be almost completely harmless evolving into a deadly strain in vivo is fairy unlikely in itself, especially if transcriptional errors can be reduced several orders of magnitude like GGP suggested.
Adding to that the option of hospitalisation or even home isolation to reduce risk of transmission, the risk of this resulting in some huge lethal epidemic must be pretty miniscule.
Software is essentially a cleanroom in the sense that the environment tends to be deterministic and man-made, and that is still riddled with unexpected accidents. Fortunately we can turn it off, fix the bug, and redeploy and the people involved in that tend to survive.
> Additionally, a bacterium that's engineered to be almost completely harmless evolving into a deadly strain in vivo is fairy unlikely in itself, especially if transcriptional errors can be reduced several orders of magnitude like GGP suggested.
The proposition was to engineer a bacteria that targets and infects a particular type of human cell to kill it. Creating medicines in a vat (like insulin) is different from releasing infectious agents in the wild. I was under the impression that this was obvious, but apparently not.
I never said we're at this stage now or even close to it, in fact I explicitly said the opposite:
>>>>Like I said, the technology is very far off from having real world applications like this. At the moment it feels like we're in the analogue of the 40s and 50s for conventional computing. The field is still just inventing the very basic building blocks. It's going to be very limited in use, wildly dangerous(look up mercury delay lines) and unreliable for decades to come
As for comparing creating medicines in a vat to using bacteria as an active treatment, you're the only one making that comparison. The paragraph you responded to wasn't about in vitro drug synthesis at all, so I'm not sure what your point is here. Yes, it's obviously different. I never said otherwise. It's perfectly possible to target bacteria to specific tissues; wild bacteria already do this.
My point was that a bacteria engineered to target a malignant tumour, to be very treatable with antibiotics or bacteriophage, and to have a strongly reduced rate of mutation, is extremely unlikely to evolve into a pandemic, and is likely to be much safer than chemotherapy and radiation.
Did you know bacteria have horizontal gene transfer? ie antibiotic resistance isn't just evolved and passed to children ( vertical ), it can be passed to peers horizontally.
It also happens outside bacteria - but bacteria have active mechanisms to enable this - that's how antibiotic resistance spreads - not just from parent to child, but peer to peer like a meme :-).
Safety is a complex topic, and you'd need to consider on a case by case basis - PhD students engineer bacteria every day ( something that had a self imposed ban in the 1970's I think ) - however that's within the context of standard platforms and each and every one should have a risk assessment.
Don't get me wrong, I think it could be done, but there is a Genie and a bottle here and it's best to think twice.
I'd like to see both a kill switch ( beyond antibiotics ) and some sort of growth external dependency - ie they need something you provide to survive as well.
^Most recent discussion I’ve seen.
I worked in genomics, left this year because you’re underpaid and often disregarded “IT-help” that assists wildly over-educated and underpaid people driving the actual research in 95% of cases.
Work is no more meaningful than anywhere else. It’s a big “selling point” for the industry, but it’s just a way to get people to get paid less (yay you’re making the world better than all those garbage people serving coffee or healing sick people or keeping your lights on or optimizing the routes of the goods you have delivered). If you want sustainable systems, trying to be a martyr and work for less only screws this up long term.
Code-base is not open-source. It’s biotech R&D, there is zero culture of sharing outside your organization within industry. You can present high-level things at conferences and such, but you’ll have to rip the raw data out of their dead hands…not happening.
I’ve been in too many conversations about building software to serve larger groups in this industry. It can happen, but it can’t currently and nobody wants it. Confident someone will find a solution, but everyone wants their own home-grown solutions in their own walled-gardens that no one has access to.
Data and the things it can/can’t tell you are held tightly in these companies. I was at a pharma a couple years ago where researchers were explicitly told they COULD NOT test certain compounds in a certain way because they did not want a trace of this data to exist while they were trying to push compounds through the FDA.
I got a real imposter syndrome from the whole project and my housemate went on to be a famous microscopist (but later rage-quit to leave for industry). I build microscopes using 3d printers and other easily sourced bits. But Home Depot is a terrible place to source materials. They are the lowest quality.
A quick trip to Lowe’s to get PVC fittings and pipes, and I suddenly made my own scientific equipment and saved me some time!
But it's a fun subject, and as the technology develops, middle layers will disappear and then the money from expertise will become better.
The number of people that are both capable software developers and has a good understanding of cellular biology are quite few and will probably remain so for the foreseeable future.
In biotech, the end goal is a physical product or a service performed by a doctor or another highly paid professional. Those don't scale as well as software. The ratio of users to developers is also low. You are likely developing software for many niche tasks, which does not scale either.
And if you are considering roles in the academia, your productivity is not going to be high enough to justify a competitive salary. Productivity, in monetary terms, is defined by the amount of money you can bring in. Either directly on indirectly. In the academia, that usually means grants. You may be able to argue successfully to a funding agency that one software engineer is worth two postdocs, but not four.
[1] – https://dantelabs.com/
at best you waste your time, at worst you will find all kinds of things that are not there
it is the Silicon Valley hacker mentality that thinks the life is some sort of computer where you can fiddle with parameters
learn some biology first, then you can marvel at it and realize just how absurdly simplistic is to think you can read anything out of some random letter
I wanted to play around with BAM files and it is much more engaging to play around with my own BAM file versus downloading one from a website.
It is also about data ownership. The value of a fully sequenced genome is limited, sure, but I still want that value without having to give my genomic data to 23andMe.
But Dante labs sure will!
Then there are also engineers from XKCD 1831 https://xkcd.com/1831
For me coming from a SWE background the computational skills are very easy to pick up especially if you work with bioinformaticians you can ask questions. It’s the genomics knowledge that is very difficult for an engineer to acquire.
And yeah it shows - contrived example after another, and honestly not a great description of anything.
If you want to truly understand genomics you have to understand how biology works. And honestly it’s great info for anyone even if you’re not getting into genomics or whatever.. why would you not want a working model of how life is put together? In that case I’d just recommend dusting off a biochem or cell bio text book and reading just the first 5-8 chapters. Typically they lay it out very simply from basic principles and the authors have far more experience and understanding and writing help than this weird tutorial course thing.
I once tried reading a few chapters of a bioinformatics book explaining DNA, RNA, protein creation, etc. The basic idea seems very simple but to my mind they explained it non-systematically with too many words. There seems to be an internal information structure in these RNA- and DNA- related processes that was not being concisely presented and it seemed that if the writers presented the material in terms of computer-science concepts, so much time could be saved.
There's nothing about sequencing by synthesis, how blocking nucleotides are added one after another, pictures of the fluorescent nucleotides on the flow cells are image analysed, etc.
This site looks like an ELI5 kind of treatment.
For example, the central dogma of DNA transcribed to RNA translated to protein seems simple, but it's not.
In almost every instance, there are vague 'rules' and many many exceptions to these rules. For example, often coding regions in genes start with an ATG, but sometimes they don't. Sometimes splice sites (where the non-coding parts of transcripts called introns are chopped out) can be predicted, but a portion of the transcripts are not spliced at predicted sites for no obvious reason. Sometimes the predictions are just wrong. Sometimes the generated proteins are modified at specific locations which impacts their function, but again, sometimes not. Even whether the gene itself is 'switched on' (i.e. able to be transcribed) is impacted by many many things, such as unidentified transcription factors, or whether the chromosomal location itself is accessible or not. There are many many other things that impact the process.
There is no simple underlying concept as the system is not designed, it evolved and is quite different among different organisms, and even in different tissues or timepoints in the same organism. As long as it works and provides enough benefit to avoid negative selection, that's enough.
It's a mess, which makes it interesting.
You are absolutely correct, there’s an information theoretical underpinning of genomics and systems biology that’s rarely if ever tackled in text books but (a) neither does this course tackle it, and (b) you can’t just skip on biochem basics and Jump to that. That’s like trying to become a physicist without learning math.
Amusingly that's literally like 80% true. Water is just a really big deal in biochem.
I saw a recent Lex Friedman podcast where the guest talks about "bioelectric patterns" and somehow getting a worm to grow a second head by messing with those patterns. I would absolutely start on this course now if it was a realistic pathway to doing something like that.
There is no REPL for the cell. No tinkering allowed.
When Marvin Minsky was growing up in New York, neighborhood pharmacists owned fluoroscopes. He said those fluoroscopes were like “great black boxes” to him and that “those kinds of black boxes don't exist for kids anymore.”
Unfortunately, each step ends up being extremely challenging, and there's tons of noise, and the cost of each Read, Eval and Print is far higher than in a programming language. Further, the "system" is running 38,000 other "threads" all of which have direct read and write access to your data, some of those threads consider your data to be the enemy and cut it up, while others are just randomly spamming your console with uneccessary debug log messages.
We have actually reached the point where some scientists have synthetically created a novel chromosome, and used a preexisting cell to bootstrap the new genome so that the cell eventually contains only protein from the new genome. To me, that represents a step beyond tinkering: it means we can create synthetic lifeforms with exactly and only the details we want, which makes studying them and engineering them far easier.
Interestingly, even though this tech exists, nobody has found any interesting use for it and it's not even really used to probe biology.
A better example would be gene therapy, which has been developing slowly over decades. A single person died in the a trial in the 90s and stopped development (that's the regulation part you're referring to) for decades. In other trials that don't include gene therapy, patients routinely die and they're just a statistic.
I should have clarified, by “REPL for the cell”, I meant one accessible to kids with an interest in science.
Nevertheless, you gave a great computational description of the cell! Wonderful analogies.
Genomics stretches vastly beyond this - assembly and annotation to start with.
I'd argue the most interesting problem space for software engineers is outside of what is covered in the document.
To get that basic biology foundation, another post mentioned an EdX Intro Biology course, that would be a terrific start, or just get a recent university-level intro biology textbook. It's not terribly difficult material and you'll be in far better shape than reading a biology-for-laypersons pamphlet.
The pay is exactly where market is. There’re ton of wet-lab people wanting to get into “data”. And the industry is less lucrative than showing ads like Google does.
> This Guide is written specifically by and for computer scientists and engineers.
How many engineers already have a bio phd. Who on earth is this actually written for.
A new generation of bioinformaticians and computational biologists are using rust, go, and the web to create, share and deliver.
Checkout nextclade.org
As soon as you start touching science, everything is important.
How do you handle one genomic variant affecting dozens of different rna transcripts and isoforms? How do you handle tissue-specific expression? LD haplotype blocks? Frequency across populations and reference choice? Sample handling affecting read depth? Mixed direction of effects in phenotype-genotype? The critical (and beauty IMO) feature of bioinfo is requiring an understanding of how your dataset can rarely be considered clean and as simple as _observation name_ and _observation value_. To succeed it is usually critical to know a lot about the observation meta data which is not collected in the dataset. Hopefully in the future it will be better curated and less esoteric.
- the dna that doesn’t code for proteins but makes up the vast majority of human dna
- the intron regions of genes that are translated into RNA but then sliced out of the RNA and not transcribed into protein and are 5x larger than the coding parts
Those two things alone are absolutely critical to understand to interpret a genome sequence. Of course there is much more.