Genomics – A programmer’s guide
gist.github.com
gist.github.com
CoFest is a collaborative two-day working session. The only requirement for attendance is that you have an interest in open source software and solving scientific problems. We will have contributors to open source bioinformatics tools present to collaborate with, and we welcome new attendees who want to learn and contribute to open source code, documentation, workflows, or training.
This year CoFest is July 26-27 in Basel, Switzerland. There is no registration fee to attend in person or virtually.
Disclosure: I'm on the organizing committee for BOSC and CollaborationFest
https://www.biostarhandbook.com/
The book is inspired by "Biostars" the StackOverflow like Q&A of genomics:
https://github.com/quinlan-lab/applied-computational-genomic...
There's a lot of subtlety in genomics analysis, and learning about all of it is a deep dive into chemistry, biology, and statistics, as well as a lot of literature search.
One fun fact about this is that most of the data in genomics is some form of a plain text table. As this makes interoperability between different programs easier. Many people have tried to make this more efficient, but their application is usually only limited to very specific use cases.
Both Gil and Peter have done some incredible things in this space, but the language of scientists can be difficult to understand by a lay audience.
I mean has anyone ever actually read the documentation of the GATK? It is famously dreadful. And that's professionally maintained.
Honestly a nice addition here would be a "so you want to" with snippets of raw FASTQ or VCF data and working code for various operations, maybe with an accompanying Docker container.
My experience has been translating domain data into spark has a 100X improvement in data analysis.
Reference for "famously"?
I was taught a decade ago that rolling your own in genomics isn't as bad of a decision as it seems.
Famous last words.
As for some resources, these books have been the most helpful for me.
[0] https://www.amazon.com/Molecular-Biology-Cell-Bruce-Alberts-... // I've linked and have seen mentioned here in the past, a great intro to cell biology.
[1] https://www.amazon.com/System-Modeling-Cellular-Biology-Conc... // Modeling with math. I would consider this high-yield as it gave me a great deal of insight into what different code bases are actually attempting to do.
[2] https://www.amazon.com/DNA-Nanoscience-Prebiotic-Emerging-Na... // Blew my mind when I first read it but the content is pretty standard as far as genetics course material goes.
If you're also interested in working on this stuff, shoot me an email ;) blopker@23andme.com
You may (at your own risk) take a look at
For your first question it depends on your definition of "companies like 23andMe". There are numerous companies that'll do a whole genome for you, but I don't know if any of them do the writeup about it that 23andMe provides. 23andMe did at one time offer an exome product, but stopped that a while back.
The largest hurdle is cost. Whole genomes, even exomes, are significantly more expensive than a SNP chip. As most would be users don't know enough to care it doesn't make much economic sense to offer those to the masses at the moment.
Also a guide to how to usr all that to interpret medical nonclemature of mutations, like c.345G>E would be handy
Those mutation descriptions are called HGVS (Human Genome Variation Society) nomenclature. In the example you give, "c." means that it's in a (protein) coding region, 345 is the position within the region, and G>E would be the change (although E isn't a valid "letter" in DNA sequence, even if you allow ambiguity codes -- you'd normally see something like G>T there instead).
Complications include:
1) You need to know which gene this is relative to.
2) The "coding sequence" for the gene isn't always perfectly defined, due to splice variation and different versions of the annotation. Ideally, you'd see this code relative to a specific splice variant (which might have an ENST identifier, from http://www.ensembl.org/). But it depends...
More at http://varnomen.hgvs.org/ if you're curious.
* RSIDs (from DBsnp https://www.ncbi.nlm.nih.gov/snp/) * HGVS as mentioned. * Ensembl chrom-pos-ref-alt (CPRA). * Variant key (Nicola Asuni)
As dasmoth says, there is no fixed coding sequence for a gene or location in the genome.
I'm a cs student, in my thesys I'll be working on a NGS C++ application. I need at least a brief introduction to "basic" sequencing but I'm struggling to find something accessible. Every book I find seems superspecialized. Now I'm reading "Insect Molecular Genetics : An introduction to principles and applications" but I'd like to read just a book chapter a little bit more advanced than the contents shown in this video https://youtu.be/ONGdehkB8jU
Any suggestions?
On the other hand, as a person who’s worked on sequencing software I’ve found the biochemistry knowledge to only be incidentally useful - though I may be underestimating some of the “basic” assumptions that were used day to day.
I have the same feeling but I'm uncomfortable working on something knowing so little about it. I'll check out the book, thanks!
The reason I'm recommending it is the quality of its interfaces. It can seamlessly handle (input or output) virtually any kind of file you throw at it (SAM, BAM, CRAM). I can't say the same for a lot of other software I have run into in this space.
[1]: https://github.com/samtools/htslib
[2]: https://gist.github.com/PoisonAlien/350677acc03b2fbf98aa
There are several next generation sequencing technologies:
1) short read - Illumina - dominates most next-generation sequencing 2) long read - nanopore or pacbio.
These have very different analysis methods, have measurement errors that are very different, and even have different file formats, etc.
Short read is far more common, so you're probably in the "Data Analysis" of this:
https://www.youtube.com/watch?v=fCd6B5HRaZ8
But you need to know about the adapters and indices (how multiple samples can be sequenced at the same time).
But as another commenter mentions, knowing some particulars about the project would really help know what sort of tutorial would be appropriate. You'll need to also know about the biology of the application, in addition to understanding the sequencing technology.
As another commenter said, I don't need superdeep sequencing knowledge because my work will mostly be on the programming side (enhance performance, not adding new functionalities) but anyways it could be useful to have a clear picture of the process.
Thanks for your help
https://www.youtube.com/watch?v=Wq35ZXyayuU
At about 1:30 there's a cartoon of the data signals that get processed into sequencing data.
It's been a while since I looked at long read data, but last time I did, the individual base calls in FASTQ files (A, C, G, T) have a fairly high error rate, and there are systematic biases in the errors, which makes it harder to correct them. Most of the processing of these data is trying to correct these errors, either by looking at a known reference sequence or by sequencing many times.
https://www.bioconductor.org/packages/release/BiocViews.html...
Here is a video by the author - https://www.youtube.com/watch?v=BfVo8EkeDVI
What is the relationship between chromosomes and the human genome?
I somewhat get that a chromosome is somewhat of a partition of the genome, but how does the „two copies“ phenomenon of the human genome and the „two copies“ thing of chromosomes fit together? Are those one and the same concept?
Are there two copies of the XY chromosome, too?
Animals tend to have the same number of chromosomes at all times, and they tend to come in pairs that are nearly identical. There are various ways of mapping chromosomes that yield unique fingerprints that are stable under the level of variation we typically see in a species (see restriction mapping for example), so we can take a particular fingerprint and call it chromosome 1 or 2 or whatever. Animals have two copies of chromosome 1 and two copies of chromosome 2, etc. There are individuals who don't, who have a single copy of one or extra copies, and this causes problems, such as Turner syndrome. Similarly, when animals reproduce, each parent produces a germ cell (sperm or egg) that has one of each of the chromosome pairs. One of the reasons that many hybrids like mules are sterile or nearly infertile is that their chromosomes, coming from different species, aren't in pairs, so when they pass on half of them, there may be necessary hunks of DNA that just aren't passed on.
Other species have different numbers of copies of chromosomes, and may vary. Depending on point in lifecycle and conditions, some plants range from two copies to hundreds. Dinoflagellates tend to have four for interesting reasons that I believe are related to the Byzantine generals problem.
There is no XY chromosome. Females in mammals follow the normal pattern with the X chromosome: they have two of them and they are passed on like any other chromosome. Males are weird. They have one copy of X and a copy of a shrunken chromosome called Y. Note that this is only mammals. Birds and reptiles have a totally different set of chromosomes for sex determination, and in some vertebrate orders the chromosomes don't fully determine sex. Incubation temperature often changes it.
The red chunks are its DNA. But they're 'chunky' because the 'whole genome' is partitioned into 23 chunk - each chunk is a 'chromosome'. And each chromosome comes as a pair. And in the movie you can see the cell split the pairs, where one set goes to one cell, and the other set goes to the other.
If you notice, in the cells just prior to 'condensing', the nuclei (red stuff) looks kind of brain-like in its topology. Those are the chromosomes as well, just relaxed and spread out.
Each chromosome is either an autosome or sex chromosome. The autosomes are chr1 to chr22. All people (excepting those with chromosomal disorders, like trisomy 21) have two copies of each autosome, with one copy coming from each parent. Then, for the sex chromosomes:
* If you're male, you get an X chromosome from your mom and a Y chromosome from your dad.
* If you're female, you get an X chromosome from your mom and another X chromosome from your dad.
So, yes, "two copies of each chromosome" and "two copies of the genome" are the same concept, since the genome consists of chromosomes.
Which phenomena are you referring to exactly? That we have two copies of each chromosome and if they mean we have two copies of the human genome?
> Are there two copies of the XY chromosome, too?
Each of our non-sex cells[1] contain two sex chromosomes: one from our father and one from our mother. Since your mother always inherits her X chromosome, your sex is determined by which sex chromosome you got from your father. If you are a female (XX), your father passed on his X chromosome. If you are a male (XY), your father passed on his Y chromosome.
This rule makes for some interesting inferences. For example, your father in turn got his X chromosome from your grandmother and his Y chromosome from your grandfather (both on your father side, of course). Your mother, on the other hand, got his X chromsome from both your grandparents on her side.
So if you're a male, your Y chromosome was passed on from your grandfather on your father's side. If you're a female, one of your X chromosome comes from your grandmother on father's side, but your other X chromosome may come from either of your grandparents on your mother's side.
You can trace this Y-chromosome lineage back to what's called the Y-Chromosomal Adam, which is the last universal common ancestor of all currently living human males[2]. You can make a similar inference using your mitochondrial genome[3] and arrive at what we call the Mitochondrial Eve[4].
Our sex cells[5] are different, since they only have one copy of our chromosome set. The number of the chromosome set we have is called ploidy[6] and so our sex cells are haploid cells, as opposed to our non-sex cells, which are called diploid.
If you're a male, a single mature sperm cell in your body contains either the X or Y chromosome. For females it's different, since they only have the X chromosome, all their mature cells contain only one copy of the X chromosome.
[1] https://en.wikipedia.org/wiki/Somatic_cell
[2] https://en.wikipedia.org/wiki/Y-chromosomal_Adam
[3] https://en.wikipedia.org/wiki/Mitochondrial_DNA
[4] https://en.wikipedia.org/wiki/Mitochondrial_Eve
I hope it was useful. It is just an introductory jargon-free guide. You can find more on Ensembl and Wikipedia.