In Newly Created Life-Form, a Major Mystery
quantamagazine.org
quantamagazine.org
As @sixQuarks has already written, finding the minimal amount of genes when there are 175 unknown ones and you don't know anything about their dependencies and relationships seems to be pretty much impossible.
> In fact, there’s no single set of genes that all living things need in order to exist. ... They found that not a single gene is shared across all of life.
That's the most interesting point to me, I deeply believed that organisms share the same basic set of genes.
The definition of life itself is not really obvious to me either, some sources count viruses as life other don't.
As to life without DNA. The methodology is somewhat in question for this but here is a more in depth link: http://science.sciencemag.org/content/332/6034/1163.full Now, this is really 1:1 with DNA just using slightly different chemistry, but it's hard to call something DNA when you swap out one of the building blocks.
Other examples are hard to locate in large part because you can't direct detect this stuff.
Anyway, my point was while a virus is 'stuck' using DNA as it needs to infect things that use DNA. However, they can use RNA internally.
PS: For a really out there example, prions seem really close to life.
Yes, it really wants phosphorus. But, a strong preference does not create phosphorus when there is none to be found.
Not that single cells really want or try etc, but you get the idea.
[1] https://en.wikipedia.org/wiki/RNA_world [2] https://en.wikipedia.org/wiki/Panspermia
It is definitely a function vs identity. Many mutations of a single gene can perform the same function -- there is a lot of redundancy in genetic code.
>If he had done the same set of experiments with a different microbe, he points out, he would have ended up with a different set of genes.
This is very interesting. I expect also, that if they proceeded with their knockout in a different order or different sets at one time they would end up with a different set of genes. Maybe they have some massively parallel way to knockout genes, but I doubt they have explored the entire set of knockouts from the original functioning natural organism. They very likely could be at a local minima of genes required (eg no single gene can be knocked out, but no where near the absolute minimum).
There is a lot of redundancy in biology, which allows for biological systems to be pretty robust, adaptable and fault tolerant. Of course this also means it's really hard to fix that biological process when things go horribly wrong, such as in many forms of cancer.
Paper: “Comparing genomes to computer operating systems in terms of the topology and evolution of their regulatory control networks.” By Koon-Kiu Yan, Gang Fang, Nitin Bhardwaj, Roger Alexander, Mark Gerstein. Proceedings of the National Academy of Sciences, Vol. 107 No. 18, May 4, 2010.
Perhaps a good moment to refer to this one:
Cancer tumors as Metazoa 1.0: tapping genes of ancient ancestors http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3148211/
Such environment should also be 'minimal' - i.e. having lowest number of resources required to construct it and lowest amounts of those resouces.
The point isn't the minimal genome, it's to identify the genes important for a minimal genome.
Genetics is all about filtering the signal from the noise. Now we know that these 175 genes are important, we can start focusing efforts on them.
The problem isn't that we don't know how to figure out what a gene does, it's that there are so many goddamn genes we don't know where to start.
Many groups of organisms do. All organisms is a tall order, if you think about it.
Maybe all is hidden behind getting blurry because of Heisebergs uncertainty principle.
Firstly what constitutes "minimal" depends on what you feed the organism. There's a set of bacteria called phytoplasma which are plant parasites that are missing genes to make nucleotides (even mycoplasmas have those) and they've adapted by sucking nucleotides from their hosts.
Some insider information: Most of the mystery essential genes are vague cell wall proteins. Probably what is going on is that if you knock out too many of these genes, you lose cell wall turgidity and the cell becomes nonviable. So what is important is not so much which of these genes you have, but how many of them you have.
Second insider information, for fun: The "hypothetical minimal genome" (syn2.0 is referred to as HMG in the paper) was actually a backronym because we called it the "hail mary genome" but decided that was inappropriate for publication.
That hints that you could just duplicate the same cell wall protein gene instead, right?
One in-house ORM I worked with was called TFP, because it was entirely written on a plane ride to Houston, and they named it TFP because the authors liked to say they wouldn't touch it again with a ten foot pole. Their company later thought about trying to sell it, with the name "Technique for Persistence."
So maybe the environment needs to be changed to support an organism with a smaller genome. The world today is certainly different than when life began. It's probably difficult to evolve the environment, so they'd need to understand what goes wrong with gene deletions and figure out how a different environment could help.
I'm not exactly sure what went wrong, but I think you're really misunderstanding something dnautics said.
M. mycoides certainly does not survive on that. Mycoplasmas need cholesterol and lipids because they lack the HMGCoA reductase pathway and FAS, so all the media you grow them on are enriched for that (started as FBS, but we found horse serum worked great and was cheaper). In general, I think you need a few more things than just that for most basic life (nitrogen, phosphorus, metals)... Is using glucose cheating? Because you could use light for energy and CO2 for carbon, but that adds a ton ton ton of genes...
So this becomes a nitpicky definitional conundrum.
There are no bacteria with smaller genomes today, that was the point of making this one. I was not talking about the global environment either, but the number of options for what substances and concentrations may need to be present in the environment around a cell is certainly very large - making trial and error difficult.
Sounds like dependency hell. Gene A only works in the presence of Gene B which only works when there's Gene C which needs Gene A. Knock any one of them out and everything breaks.
Random mutations are like that. It's almost like no-one was planning out how these things should work, and whatever did work stuck.
Weird, huh?
Isn't the current theory that all life on Earth has a common ancestor? Wouldn't this point put that into question and introduce the possibility that life has developed multiple times independently of each other or would this lack of shared genes eventually appear through divergent evolution? It seems like the former option would greatly increase the fl component in the Drake Equation meaning alien life is even more likely than previously thought.
If that's the case, it isn't that unlikely that every evolutionary branch managed to improve (or discard) every part of the original machinery in some way.
Not really. For one thing there's too much commonality across life in terms of things like the genetic code, i.e. what a sequence of genetic letters "means". The genetic code is arbitrary but essentially universal with minor variants.
For another thing, the organisms now present can be separated from each other by as much as 8 billion years of evolution (4 billion on each branch) meaning there's no particular reason to expect some shared gene to be present everywhere. New genes arise as mutations from old ones, old genes become obsolete and get removed, etc.
It's vaguely like how, if you fork a coding project, and the two projects are allowed to continue indefinitely, eventually there might be no code in common, even without a complete rewrite ever happening.
They do. What you're responding to says there's no one gene shared by all organisms. The genetic intersection of any two organisms is large; the genetic intersection over all organisms is empty.
this seems to support the idea that instead of evolving from a single ancestral life form, all beings were created with care
edit: or at least would seem to support the weaker assertion that singl life form ancestry evolutionary theory is inaccurate.
if this is logically incorrect please feel free to tell my why. downvotes are not for disagreement!
At face value the statement "if not all living things share a common ancestor, that makes them appear more like they were intentionally created" (although that line of reasoning, i.e. the watchmaker argument, has been debunked time and time again) seems reasonable.
However the statement seems far more grandiose than it actually is.
If two very "primitive" organisms (i.e. operating at a much lower complexity with much smaller and fewer moving parts than even, say, an earthworm) don't share their sets of DNA that may be unexpected (because whatever step there was from "not having DNA" to "having DNA" had to happen twice, separately, successfully).
If we found out that, say, all higher lifeforms shared no common DNA with other lifeforms outside their group and those groups closely aligned to the Christian "kinds" (which don't really map to biologically distinct groups in meaningful or consistent ways) THAT would be astounding.
But even that extreme case wouldn't make the "creation hypothesis" (if you even want to call it that) more plausible because it presumes the existence of an unexplained organism that is far more complex than anything it supposedly created. It doesn't explain anything -- it just shuts up questions about the "origin" by giving an answer that is impossible to analyse further.
There may have been a creator. But it's impossible to make a consistent logical argument for it. And it's entirely impossible to make an argument that that creator would in any way resemble something modern mainstream religions are worshipping.
I think this is a pop-science misunderstanding of what the authors were trying to say. They couldn't find a single unique "minimal set" of genes- but there are genes that we know all organisms share, like the rRNA genes and homologous core ribosomal proteins. All organisms will also require some kind tRNAs and their genes from some source, even if the tRNA genes themselves are not highly homologously conserved.
[1] https://www.amazon.com/Creation-Life-Make-Steve-Grand/dp/067...
Is there any way at all at simulating the tests? I'd like to apply metaheuristics to search for which combinatorics work
Yes, isn't there a universal gene that encodes the RNA polymerase enzyme? I.e., the machinery that transcribes DNA into RNA?
RNA Pol is very diverse across clades.
It should seem obvious in retrospect. Turing machines have many incarnations as different instruction sets in modern CPUs. I don't see why cells wouldn't have analogous multiple expressions, unless your deep belief was that there was only one way for protein chains to replicate.
[1] https://www.cmu.edu/biolphys/deserno/pdf/can_a_biologist_fix...
Got a better idea?
First, when you can't manipulate things directly, it becomes more difficult to tweak systems slightly to learn how they work. This is especially true if you start with something made from things you don't understand, and for which you may not even have a complete catalog. The radio paper raises some issues and acknowledges these challenges, but then sort of glosses over them, saying, "it's complicated, but that's just because we don't understand and don't use the right language."
Second, biologists don't just use subtractive methods, they try to use additive approaches to complement and test the hypotheses that come out of "it's broken when I take this bit out." Biochemists also try to reproduce complex behaviors by mixing only the purified components of the system in question. Biochemists and biophysicists also often measure very precisely physical and chemical properties to mathematically define the behavior of biological components.
The point is biology is hard, and we barely know anything about it. We have to start with simple experiments that point us to what's important and then progressively do better defined experiments to figure out how it works in detail. The next step here is: now that we know what proteins are important, let's figure out how they work exactly. We couldn't have gotten there before trying to knock out as many parts as possible.
TL;DR: this isn't the wrong approach, it's the right one, and given the limitations of the system, realistically the only one we have at this stage. It's also only the first step, now we get to try other approaches on what we've learned.
They use various neuroscience techniques to see how far they can get reverse engineering an Atari.
When you're talking about 175 unknown genes, the combination of all of these is a huge number. It's like finding a needle in a haystack the size of the solar system.
I don't think this brute force approach is going to work, we need a different way to figure this out, but I'm confident that once figured out, it will seem simple looking back on it.
Why not? Couldn't we use machines to automate this and analyze the results?
http://coolconversion.com/math/factorial/What-is-the-number-...
Perhaps a quantum computer can tackle this kind of problem, but nothing we have today can brute-force the solution.
That's Mendelian genetics. Geneticists focused on DNA as a physical molecule rather than a conceptual unit of inheritance use the same word to refer to a region of DNA that produces a protein.
They're both important concepts, but they're not the same concept.
In fact, I think that term is great for comparison: "phenomenon". Asking for the "minimum genes for viable life" seems like asking for the "minimum number of natural phenomena to put a man on the moon". Are those two phenomena really just different manifestations of "gravity" or are they two things in their own right? "It depends."
Genes are not independent functional modules. Their placement and arrangement on the genome matters. Did they only mess with coding features (genes)? Or did they also mess with other genomic features? Or do such things just not matter with bacteria?
But he reminds us that the subtractive process entirely depends on the starting cell.
I would like to hear more about the difference between additive and subtractive methods -- it wasn't entirely clear from the article.
As far as I can tell, they synthesized the molecule.
A line from the paper...:
We used whole-genome design and complete chemical synthesis to minimize
I did fret for a second over what meaning of "remove" daemonk was using (from the design? From the instantiation of the genome?) but decided I at least wasn't posting misinformation.
Subtraction method: After reading a bit of the paper, they utilized Tn5, which causes a DNA sequence to be arbitrarily inserted into the DNA (Random locations). This randomly disrupts genes, causing them to likely not function correctly.
Here's the logic: If all genes were necessary we would expect no cells to live that had mutations.
If no genes were necessary we would expect all cells to have mutations at about an equal rate relative to the space they occupy. (Assuming no bias by the Tn5, but that's a nuance)
What they instead found was that some cells grew, but there were certain genes that were not mutated, meaning that they are likely necessary.
They also classified a few things like studying the growth of the cell (Slow growing, but still viable was classified as "quasi-essential") and implanting their minimal genome in another species, which failed giving evidence that a subset of genes that are sufficient for survival are not necessarily sufficient in all cells.
Otherwise, it would be like cutting out instructions from an assembly language file, without knowing whether one of the instructions does something like adding a constant integer value to the program counter. You might think the code that was removed was critical to the operation of the program, but you could have instead replaced all the instructions with an equivalent number of no-ops, and it would still work as expected.
I believe they wanted to modularize the bacteria by placing cassettes of common functionality (e.g. all the tRNAs together) in parts of the genome. I don't know if they are still are working on that, or how it went.
In yeasts, the tRNAs have to be in the same orientation as the replication bubble, but there didn't seem to be that restriction in mycoplasmas.
If you take someone's 4000-5000 line program and whittle it down to 473 lines which are still somehow useful, "newly created" doesn't apply in full honesty, let alone if you don't know what a third of those lines do.
This is just excellent science. It seems like it should be very easy to get these unknown genes to reveal their function now. Very exciting times.
Can you expand on this comment? What would such a layer look like? Have we missed layers before and then found them?
edit: pedantics.
See http://www.sciencedirect.com/science/article/pii/S1534580708...
Love it when linear algebra pops up in unexpected places!
"Second layer of information in DNA confirmed"
A gene is a code for making a specific protein. Each gene is a sequence of three-letter "words" called "codons". There are special "start" and "stop" codons which mark the beginning and end of a gene.
[1] http://genetics.thetech.org/about-genetics/how-do-genes-work
Also what are the raw/compressed sizes for a human genome?
I have wondered this for a long time but never seem to find a concrete answer.
A nucleotide is one of [A,T,C,G]. So that means it encodes 2 bits of information.
473 * 1200 * 2 = 1 135 200 bits, or 141.9 kilobyte.
Of course, this is coming from a software developer using numbers from Google, so I might be wildly wrong.
[1] "Molecular biology of the cell" by Bruce Alberts
you can download the human genome from here http://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/hg38.... (it is recommended to use ftp, see: http://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/). it is an ~800mb download.
if it takes 800 mb to store 3.3 billion bases, i would assume it takes (531k / 3.3b) * 800 mb = 0.141 mb to store this genome.
genomic regions with higher variability between individuals are annotated separately and are not represented in the main chromosomal sequence, but can be aligned to it as such.
http://hgdownload.soe.ucsc.edu/goldenPath/venter1/bigZips/
However, the human "reference" genome (the updated main product of the Human Genome Project) is a patchwork of about 20 donors' genomes. Many(most?) of the donors are anonymous volunteers from Buffalo, NY, USA.
edit: appears to be the full article (scroll down past summary)
To answer a few question that have come up:
How did they synthesize the genome? By creating short sequences of DNA and joining them together step by step to make larger sequences. They chemically synthesized short DNA sequences and assembled them into 1.4 thousand base pair (kbp) fragments. Five fragments were assembled into 7 kbp cassettes, which were assembled in yeast to generate 1/8 chunks of the genome, and these chunks were assembled in yeast to create the full genome. (See figure 2.) To try out deletions, they could replace a 1/8 chunk, rather than synthesizing the whole genome.
Why can't they delete all the non-essential genes? Bacterial have a lot of redundant genes, where two different genes provide the same essential function. Just like redundant disks, you can remove one, but not all of them. And you'll get different of minimal genomes depending on which gene you keep.
One interesting thing is that because the growth medium provides almost all the necessary nutrients, they could remove a lot of the metabolic genes, but needed to keep a lot of genes to transport molecules across the cell membrane. You can imagine a minimal cell constructed the opposite way.
The article doesn't mention that in their first synthetic bacterium (2010), they encoded text (actually HTML!) in the DNA to provide a secret watermark. See http://www.righto.com/2010/06/using-arc-to-decode-venters-se...
I found about the book from HN, and have since bought every single other book by him, almost done with Life Ascending now, which is also amazing
Trying to produce genuinely novel and distinct biological lifeforms sounds cool enough on its own though.
Manufacturing a living cell from purely chemical ingredients is a concrete enough goal.
And almost certainly it lies far, far off in the future.
Thinking if it like: Just remove files from the OS until it wont start :)
It is the activation key God puts on each living being... Ain't gonna work without it. =P
After a while that got old.
http://science.sciencemag.org.sci-hub.cc/content/351/6280/aa...
About Sci-Hub: https://en.wikipedia.org/wiki/Sci-Hub