"Maximum mutational robustness in genotype–phenotype maps follows a self-similar blancmange-like curve"
https://royalsocietypublishing.org/doi/10.1098/rsif.2023.016...
"Maximum mutational robustness in genotype–phenotype maps follows a self-similar blancmange-like curve"
https://royalsocietypublishing.org/doi/10.1098/rsif.2023.016...
I submitted two papers about this, both were rejected. Not surprising I guess, my first papers ever!
It was also fascinating to me how complex (and how much larger) plant genomes are compared to animals. Plants have enormous genomes, almost making me wonder, what do they know that we don't?
I think the answer could be--a lot! :)
As a consequence, there are just a handful of papers that answer great questions.
This has probably held back science a bit.
Yes, the genome size of the common frog (Rana temporaria) genome is 4.1 Gb. https://www.ncbi.nlm.nih.gov/datasets/taxonomy/8407/ which is larger than that of mammal genomes.
On the other hand, the smallest frog genome is the Ornate burrowing frog at 1.06 GB, which is 1/3rd the size of the human genome. https://en.wikipedia.org/wiki/Genome#Genome_size . https://en.wikipedia.org/wiki/Ornate_burrowing_frog says:
"This is an adaptation to the desert environment where it lives. Because the ponds where they breed dries up fast in the desert, the tadpoles has to go through metamorphosis as fast as possible, which can occur just eleven days after the eggs were fertilized. A small genome gives small cells, and the smaller the cells are, the faster the tadpoles transform into small frogs and can escape the shrinking ponds"
Xenopus laevis (African clawed frog) is 2.7 Gb, also smaller than humans - https://www.ncbi.nlm.nih.gov/datasets/taxonomy/8355/
Thus, warm/cold-blooded cannot be the main reason for the difference.
From https://en.wikipedia.org/wiki/Genome_size :
"genome size is not proportional to the number of genes present in the genome ... In eukaryotes (but not prokaryotes), genome size is not proportional to the number of genes present in the genome, an observation that was deemed wholly counter-intuitive before the discovery of non-coding DNA and which became known as the "C-value paradox" as a result."
From https://en.wikipedia.org/wiki/C-value#Variation_among_specie... :
"Variation in C-values bears no relationship to the complexity of the organism or the number of genes contained in its genome; for example, some single-celled protists have genomes much larger than that of humans. ... C-values correlate with a range of features at the cell and organism levels, including cell size, cell division rate, and, depending on the taxon, body size, metabolic rate, developmental rate, organ complexity, geographical distribution, or extinction risk ...
Perhaps a naive view but, plants have just a limited nervous system[1] and no brains. Doesn't it make sense that they compensate by more preprogrammed behavior in the form of a more complex genome?
The range of genome sizes for plant genomes is very large. "In animals they range more than 3,300-fold, and in land plants they differ by a factor of about 1,000" and "genome size is not proportional to the number of genes present in the genome", quoting https://en.wikipedia.org/wiki/Genome_size .
See also the chart "Genome size ranges (in base pairs) of various life forms" from that page, and the table at https://en.wikipedia.org/wiki/Genome#Genome_size .
Arabidopsis thaliana has about 135 MB base pairs. Humans have about 3.1 GB base pairs. Paris japonica have about 150 GB base pairs.
I knew about non-coding or junk DNA but had forgotten how extreme it was.
Ended up learning about the onion test[1] and this interesting article[2] which suggests classifying non-coding DNA as either junk DNA, spam DNA or both, to better capture the differences that exist.
I also found this article[3] providing some interesting overview and discussion on the topic.
[1]: https://en.wikipedia.org/wiki/Onion_Test
Unfortunately (as I am discovering) the intersection between biology/genetics and some of these more computer-science-specific ideas gets far less interest than it should.
At the time, it hurt, but now I don't care. I don't even remember where I submitted it! I don't care.
If they'd accepted it, I probably wouldn't be working in software. I'd probably still be in some kind of research, I guess. I don't know!
It's so many years ago, I just don't care. Time changes everything I guess. I suppose they missed out! Ha ha! :)
For example, quoting "Sublinear growth of information in DNA sequences", Giulia Menconi, https://academic.oup.com/nar/article/32/suppl_2/W628/1040725... from 2005:
> The Lempel–Ziv complexity measure is based on text segmentation; we have termed it a ‘complexity decomposition’. It may be interpreted as the representation of a text in terms of repeats. Initially, this approach was implemented for analyzing DNA by Gusev and coauthors (13,14).
13 is Gusev,V.D., Kulichkov,V.A. and Chupakhina,O.M. (1991) Complexity analysis of genomes. I. Complexity and classification methods of detected structural regularities. Mol. Biol. (Mosk). 25 , 825–834.
14 is Gusev,V.D., Nemytikova,L.A. and Chuzhanova,N.A. (1999) On the complexity measures of genetic sequences. Bioinformatics, 15, 994–999.
Google Scholar has 326 matches for "Lempel-Ziv power law dna" at or before 2010, for example, "Entropy and predictability of information carriers" from 1995 with "The capability to describe the structure of information carriers as DNA, proteins, texts and musical strings is investigated.".
I was wondering in particular whether anyone looked at "junk" DNA with gzip, etc., as a test of how randomised it was.. Do you know if that ever happened?