The nightmarishly complex wheat genome finally yields to scientists
arstechnica.com
arstechnica.com
Having the source code... I'm not sure what that would be analogous to, but it would probably quite a long way off.
(To see the analogy, substitute file = chromosome, and routine = gene)
And that's not even taking into account non-DNA things that encode and propagate information, such as microbiota or anything linked to epigenetics.
Or is it opposite?
And, by the way, running programs are also the result of their source code and the environment in which they execute. It’s just that the environments which computer programs run in are very homogenous and predictable. But flip some bits in memory or have a user thrash the UI and you’ll see that a running program interacts with its environment in much the same way as a biological organism.
-I like to think that HN is filled with people who take interest in many things, especially science-related subjects, and would hope that many of them on this site would like to know more than an analogy-based understanding of genomics
-For better or for worse, the recent bloom and advances in genetics and biology in general have taken over the modern world and been hailed (probably with reason) by all kinds of circles, from the media to governments or the tech world, as a promising new era that's going to revolutionize our understanding of life, disease, what it means to be human, cognition, what have you. It does feel pretty cool as an insider to know that you're working in the 'hot new thing' but a side effect of that is, I have seen people develop a pretty bizarre fascination with DNA, whether it be their own or other people's. And it's not something that can be pinpointed to entirely rational reasons. I don't want to go off tangents any further, but let's say it can't hurt to occasionnally remind people what DNA is and what it isn't, what it's like and what it's definitely not like. I dislike pedants as much as the next person but seeing so many (presumably very educated) people take that analogy in the same thread made me think that the reminder wasn't that out of place.
Perhaps the mapping will yield improvements against blight and pests, improving crop yields and reducing human hunger?
I'm not sure genetic mapping is really the core of the issue anyway. What we need is the equivalent of a reverse vaccination, since celiac is essentially an immunity to wheat. Such a concept might do for chronic disease what vaccination did for acute.
As far as future treatments, this clinical review https://www.gponline.com/coeliac-disease-clinical-review/gi-... concludes that
“Future potential treatments for patients with RCD include the development of genetically detoxified grains, oral and intranasal 'coeliac vaccines' to induce tolerance, inhibitors of TTG, and detoxification of immunogenic gliadin peptides via oral peptidase supplement therapy.”
A single "reference" genome can serve as a high quality basis for interpreting new genomes. This is called resequencing. Small pieces of the new genomes are sequenced with very cheap techniques and mapped into the most-similar bits of the big genome, in effect allowing for a kind of guided reconstruction of the new genome. The short read lengths make it difficult to reconstruct a new genome de novo, but that's OK because if we have a reference we've already paid that price once for a related individual.
This process has bias but it's also expedient, say when you have thousands or millions of genomes in a study (as happens in agriculture).
One big problem in the field now is this obsession with reference genomes and resequencing against them. Researchers forget how much the reference genome can bias their resolution of a new genome. This has resulted in a lot of fussing and fixation on high-entropy regions of genomes and small variations (point mutations, or SNPs) within them. There is a growing body of evidence that suggests that large changes are probably more important for adaptation, but they remain somewhat underappreciated because of this short read resequencing modality.
The end result, though, is that now that we have access to a high quality picture of what sequences belong to what chromosomes, we can compare other varieties to that reference, see what structural differences (that really means 'sequences moving around') occurred over the course of evolution, when they occurred, why (in evolution terms) they may have occurred, and so on. Understanding that is key to do all kinds of cool experiments (or from an industrial perspective, optimizations) that weren't accessible to us previously, and the stakes are very high when dealing with such a staple crop that feeds billions of people.
Also, many varieties of wheat are hexaploid - their chromosomes come by six, instead of two for us humans. This isn't that exceptional for plants, but it means sorting out the chromosomes with the right alleles when almost each sequence is more or less replicated six times is, as the article title says, a nightmare of complexity. So even if you skip what others and I have said before, it is a technical prowess and it sets a very hopeful precedent for future projects involving other species of similar complexity, and many crops and plants happen to belong to that category.
You can view the human genome at UCSC genome browser [1] and zoom in to the "base" level to see the actual sequence. It has lots of "annotations" where projects have marked where genes are (most of the genetic code does not encode proteins, genes are the parts that do, the rest is less well understood). You can see data there that shows you which parts of the genome vary from one species to another (generally mammals will be very similar to us in genes and less so between genes where the DNA is "less important" and so changes there don't result in death, so looking at where things are "conserved" by evolution tells you where the most fragile and important parts of the genome are, to a first guess).
What "they don't tell you" is that it's still not really complete. For example, there are bits of genetic sequence that we know go somewhere in the genome, but we don't know where. The way we sequence DNA, you only get to read little pieces at once and then have to string them together to reconstruct the whole thing. But one problem is repeats: sometimes the genome just goes "ATATATATATATATAT...." and if there are more there than the length of each piece you can read, it becomes impossible to know how many there are in total. Ambiguities like this exist because our DNA isn't just random strings of data, it has lots of structure that repeats from one place to another. So it can be impossible to tell where exactly a piece of DNA belongs, given the limited information we have. Also the annotations of where genes are is very much not finalized.
There are other problems with the sequences. For example, they'll include "N" to mean "some base, we don't know which" - there are millions of N's at the start of the chromosomes I've looked at. In reality, we all have some actual code there, but it's never been sequenced for anyone. Other structures confuse us: there is DNA encoding for ribosomal RNA which is used in ribosomes which are used to go from genetic code to actual proteins. So your body needs lots of ribosomes to keep everything being made so it needs lots of copies of the ribosomal DNA to produce enough ribosomal RNA - like 500 copies in a row. And it has those not just in one chromosome but in multiple. But the human genome project doesn't show it like that as it really should be: it has about half of one copy in the right spot. And then the rest of it is just spread in short bits throughout the rest of the places the code occurs. This causes problems if you are doing research and don't know what it is you're looking at, since it looks out of place. It's a complete mess. But despite that, it's good enough to do so much with.
Here's one use for a fully sequenced genome: RNA sequencing (which, I'll reiterate, I just started working on this a month ago so I've probably said at least one false thing already). RNA is made from DNA and proteins are made from RNA. So if you want to know what proteins are in a cell - and they're the things that actually do the work - then you can look at what RNA is present. This is easier these days than checking the proteins themselves. But RNA on its own doesn't tell you a lot, so you "align" the RNA to a reference genome: given a little piece of RNA (like 100 nucleotide bases long), where does it fit in the genome? It often has a unique spot it could have come from, so now you know what DNA was turned into what RNA. If there's a gene there, then you know what gene is being "expressed," meaning turned into protein via RNA. Genes don't do much if they're not expressed, so this is very useful to know to understand what is happening in a cell. Looking at RNA sequencing versus DNA sequencing is like looking at what software is running on your computer versus looking at what code is stored on the hard disk. Both tell you a lot and understanding what code is on the hard disk will help you understand what you're looking at when you're looking at what code is running, like "Oh, these operations belongs to the bash executable - so bash must be running".
[1] https://genome.ucsc.edu/cgi-bin/hgGateway (hit GO if you just want a piece of human DNA, hit "base" button in zoom in if you want to see actual sequences but that's too close to see much of the structure that the annotations provide)
If it has a solid stem, wouldn’t it no longer be in the grass family by definition?
I mean, don't get me wrong, you CAN develop phylogenetic trees based on traits. That's how Linnaeus and all subsequent taxonomists did it before they could just look directly at DNA. But it takes a lot of traits, and has been preempted by DNA based methods anyway.
What about the rhyme “Sedges have edges, rushes are round, and grasses are hollow right down to the ground?” You’re telling me scientists are now saying that DNA testing is more accurate than just using the rhyme?
For example, if you looked at just the ability to fly, you would group birds and bats on a same group. But that would be wrong because bats are clearly more closely related to flightless mammals than they are to birds. This becomes obvious if you account for more characteristics of the animal (presence of fur, bone shape, pregnancy, etc) or if you examine the DNA.
NRGene has a proprietary algorithm that uses a large amount of Illumina sequences to build a pretty good genome assembly. This technology has been used in a few organisms with large genomes, as it has been used in v1.0 of this assembly.
However there have been previous, less high quality wheat genome assemblies based on Bacterial Artificial Chromosomes (BACs), and that information has made its way into this assembly too, along with long-distance information from HiC and genetic maps and Bionano technology.
https://en.wikipedia.org/wiki/Gene_bank
e.g.
Digital versions can be replicated and distributed at 0 cost.
> In October 2016, the seed vault experienced an unusually large degree of water intrusion due to higher than average temperatures and heavy rainfall .... The vault was designed for water intrusion and as such the seeds were not at risk.
The article's source: https://www.popsci.com/seed-vault-flooding
We should indeed continue with seed vaults. But I'd feel better with a backup for that in the form of digital data.
Frankly, I think storing data as ASCII in pits on a substrate is a great long term solution, as it can be read with a simple optical microscope, and ASCII is forever.
One of the cool things about digital data is it can be copied forward limitlessly. Make the seed genomes public domain and put them on a server - people will make copies, like they make copies of wikipedia to store in their prepper bunkers.
Human knowledge is more then English alphabet characters and not all concepts translate properly.
What we can do today is to manipulate DNA and re-inject into an egg. Although not nothing, this is not "just its genome."
In 100 years? Probably.
is it possible to recreate a running program from just it's source?
(lacking a compiler, development and production environments, and perhaps user inputs)
so my feeling is, not really, to an approximation
And, no, it is not possible (today) to rebuild a multi gigabase hexaploid genome. The best that's been done so far is mycoplasma (500kb genome, which is many orders of magnitude smaller).
Also, the "genome sequence" presented here is not really complete. Thus far only the simplest organisms have had their genomes synthesized from a digital template and resulted in a viable cell. We rarely can produce a "complete" genome for anything other than bacteria and viruses. And for those, the diversity in the population makes the idea of establishing a single genome sequence somewhat limiting.
Having a monoculture of crops is a terrible vulnerability. Look at what is happening with our current monoculture of bananas.
Personally, my sensitivity to gluten has just about ruined my life - right now, my life is on hold and I am staying in a hotel in Rochester while obtaining treatment from the Mayo Clinic’s celiac specialists. Not sure that anyone can say that about quinoa or millet.
sorry for leading with a negative, I know it can come across more strongly than I intend.
There was a scifi story about such a scenario, I think it was entitled "No Blade of Grass" by John Christopher.
Additionally, while many people are not familiar with cereal grains other than the close wheat family, there is really nothing that different about their flavor or how they taste in cooking. The only useful, unique thing about wheat is that the protein provides a texture in baked goods and sauces that people find pleasant. When forced to explore alternative grains for a couple of years, and having eaten my mothers alternative grain cooking before that, I learned that the gluten-free foods often actually taste better. Wheat is necessary and flavorful like Windows is user friendly and required for your computer to function.
If you put my comment in context, I am not intending to advocate any sort of eradication program for wheat. I’m responding with mild hyperbole to the notion that it is particularly critical to preserve wheat. The main idea that I am promoting here is that there’s nothing special about wheat, other than that it is particularly harmful to some people.
As a child I was allergic to all dairy products with serious eczema from head to toe, even a drop of dairy would initiate it, entire body covered in scabs, arms in wet bandage inner-wrapping, dry outer bangage should I get exposed to even a little milk product. Yet I do not want all dairy products to be banned, they bring valuable nutrition to diverse groups of people, I was just unlucky, but very lucky to be born in a country with a health service that cares.
Wheat is a wonderful food, extremely nutritious, that has been the cornerstone of the development of multiple civilisations.
Why eradicate it? Just avoid it.
If the global food industry was capable of keeping wheat out of their food, that would be fantastic. Unlike dairy, almost all dry foods are contaminated with gluten unless you go to lengths to obtain food that is not.
If we're going to optimize for anything, it's for feeding as many people as possible -- there are far fewer celiac sufferers than there are people who go without enough food. Optimizing the food chain to protect you would result in many more people dying from starvation, because you'd be eliminating an efficient, widely grown crop that feeds many. Wheat is the second most important crop in the world, behind only rice.
There is nothing special about wheat. All of the resources spent on growing wheat could be expended on many other crops which would provide equal or greater nutritional benefit. The reason wheat is a globally important crop is simply because people have made it that way for arbitrary reasons, like raising European cattle in North America.