The increase of genetic complexity follows Moore’s law (2013) [pdf]
arxiv.org
arxiv.org
First: Moore's law is just exponential growth (with a doubling time t). Drop the Moore's Law terminology from the paper.
Second: C-value paradox isn't. It's based on false assumptions. Given there are frogs and trees with genomes 100X the size of their nearest species neighbors (this is likely due to some completely boring mistakes), and that data is just copies and copies of the same DNA over (inlcuding duplications of genes, duplication of regulatory regions, and duplications of hanger-on DNA. Calling it an enigma (in the cited paper: http://www.ncbi.nlm.nih.gov/pubmed/19216716) is a bit better, because if we understood how genomes can tolerate this sort of expansion it owuld be great.
In short: the authors recognize some of the important problems, but go off in the wrong direction with it.
I'm not an expert in biology and have no idea how good this paper is. I was getting a PhD in math from West Virginia Uni in early 90s. Sharov was visiting at the time and had build a reputation of a pure genius.
"Don't worry about people stealing an idea. If it's original, you will have to ram it down their throats."
- Genetic complexity roughly increases 10X every billion year.
- This implies life started about 9 billion years ago.
- Adjusting for potential hyperexponential effects, origin of life could close to start of universe itself.
- It took 5 billion years to get to the complexity of bacteria.
- Life had started long before Earth came in existence but it was very primitive (aka panspermia).
- If you follow this theory it means, there were no intelligent species in universe that had existed before us. This means, Earth wasn't seeded by another intelligent aliens.
There are some sweeping propositions that Kurzweil is wrong on whole singularity concept. There is also a mention that Drake's equation is likely wrong to calculate number of civilizations in galaxy based on this theory.
In a way, I think this paper actually points to many more possible places with intelligent life because it proposes primitive life had been floating in universe long before and planets like Earth got seeded with them. Unless we assume Earth as having the best possible conditions in entire galaxy, this means intelligent life has to exist somewhere else in large enough sample of planets. The only problem is that this life isn't going to be far more super superior (or more accurately exceed us in genetic complexity).
There is no point there, I see no logic in that statement, as I said we know nothing about the subject, this is pure speculation.
Some interesting comments from last year (mostly critical): https://news.ycombinator.com/item?id=5580334, https://news.ycombinator.com/item?id=5552381
>Thus we stick to the suggestion to measure genetic complexity by the length of functional and non-redundant DNA sequence rather than by total DNA length (Adami, et al. , 2000; Sharov, 2006).
Please God tell me they've at least tried compressing the DNA to account for actual fundamental string complexity rather than just length.
>There is no consens us among biologists on the question how variable are the rates of evolution.
Well yes, I don't see that their model isn't blindly averaging over punctuated equilibria.
Overall, a nice line of research to pursue, though. Deducing how much time evolution actually takes to optimize from lifeless particulate soup to the first life-forms to animals to people would certainly be an achievement.
There is selective pressure to reduce the length of the dna within an organism, since maintaining dna is costly for the organism. So one might suppose that on the average, over time that an organism will only contain the length of dna which is necessary, give or take.
Of course it might be that the encoding of the information changes over time, so that the organism is able to store more information in a the same length. Or possibly less information in the same length, for reasons of benefit to the organism.
This needs careful reasoning, I am not convinced either my approach, or yours is enough.
I would still caution against using any old encoding technique on a string representation of the genome and using the compressed length as any sort of meaningful measure of the inherent information contained within it.
Therefore I am not sure that compressing the string is likely to give you a sense of the information contained within it, at least information in the sense which we are interested in.
See for example http://www.illc.uva.nl/Research/Publications/Dissertations/D... 'Statistical Inference Through Data Compression'.
However I am still not convinced that this upper bound is going to provide us with the information that we need in this case. What we are looking for is a relative measure of the complexity between each genome. The upper bound will not necessarily give us this relative measure because it may not be able to compress the genetic code of organisms by the same factor. The compressibility of a particular genome, by a specific algorithm will be dependent on the method of encoding of information used by the organism. For instance the organism may repeat codes for redundancy, but it may permute the letters of the copy in a predictable way, for it's own reasons. The compression algorithm used will not pick up on this.
It is useful to use compressiblity by a range of algorithms as the input to a machine learning algorithm, or as part of the model in AIXI, but it is not useful for estimating algorithmic complexity (as far as I am concerned, I am open minded, but not convinced yet).
It may or may not. But an upper bound is still an upper bound, and turns out to be usable for many purposes.
So you think DNA tends to be algorithmically random?
I was thinking about this a couple of months ago, and I've been browsing a load of verbose OxBridge articles looking for just one aspect: most uses of compression algorithms are for digitally stored data which can be classed under a homogeneous class of 'ASCII files', regardless of file type. But genes have a very small alphabet, four letters, so I can fit four bases in a byte, four times more space efficient than old ASCII. On the other hand certain compression algorithms could pick up on this (disclaimed, I'm not an expert on compression algorithms).