Theory Suggests That All Genes Affect Every Complex Trait
quantamagazine.org
quantamagazine.org
1. Genes are not loci. The "one gene = one protein" dogma of molecular biology (expanded: "one trait selectable in breeding experiments = one compact locus of the chromosome") was a priori wrong, but it's taken us decades to undo the damage. We were led astray because there are traits that do map to a single locus, or single mutations that were found to be able to control a trait, which is not the same thing as the trait.
2. Distributed representation. If I point to a sentence and ask "where is the sarcasm?" (assuming the sentence is sarcastic), there is no answer. It's certainly a trait of the sentence, just as, say, being red headed is a trait of the organism. But a linear model showing each word's contribution to sarcasm isn't helpful.
3. Perhaps a corollary of (2), humans have the concept of abstraction. There are many situations where you can get something like an abstraction from evolution. If a human engineered it, we'd call it a leaky abstraction, but please don't get caught up on that. Evolution doesn't abstract. But there are structures that emerge that can have similar properties, especially from repeated exaptation. Consider the MAP kinase pathways. Lots and lots of cell responses involve them, in all kinds of subtly different ways. We don't try to claim that a MAP kinase is the gene for anything in particular, any more than we would claim that the interrupt for triggering a system call on 32 bit Linux is the cause of certain behavior that a program is supposed to have in a particular domain.
Complex/quantitative traits on the other hand have been known for quite some time. I don’t think it’s anything new, but the landscape is becoming increasingly clear thanks to large scale GWAS and analyses from many groups including Pritchard’s.
That is a very profound observation. It suggests a way that one can distinguish between something that evolved vs something that was designed.
This notion should be supported by the fact that we don't mutate in completely random ways, and the fact that most mutations don't kill us.
Could you elaborate on what you mean by ‘evolution doesn’t abstract’? I did not totally understand what the example given had to do with abstraction
If the evolutionary pressures on E (or C individually) are greater than on A, C could evolve completely out from under A to the point where A is no longer functional or foldable. As such, evolution can't "rely" on abstractions like "protein C" because their "source code" can be changed nonlocally and their function is dependent on very local conditions like temperature, pH, and surrounding molecules. Abstracting away g(h(i(x))) as f(x) in Mathematica is pointless when your packet sniffer can arbitrarily change the meaning of "h(x)" to "nop".
I'm not a biologist or geneticists but from everything that I've read about cellular biology and genetics I get the strong feeling that there is a lot of encapsulation going on. Sure, it isn't 100% percent perfect but there is still encapsulation going on. Evolution would probably be way harder without it.
One hint is the human body, there are lots of encapsulations: eyes, heart, kidneys, digestive system, etc.
Evolution can repurpose an existing system for a completely different use with just some tweaking.
This is a bit off-topic but I think paleo-ontologists fall into a similar trap, at least based on popular writing about e.g. sharks. Claims that they haven't evolved in 100M years because they're the same shape as they used to be completely ignore all the other aspects of the shark that could change without affecting the overall shape.
Of course there's no little physical switch. This is a giant mass of molecules bouncing around. A fellow I know did some interesting work on replacing simple checkpoint models with entrained oscillators. But this makes the checkpoints look like an abstraction.
That's our view of it, though. It's not an abstraction. It's something that looks like one.
This is a physicalist line to take which implies that the universe is entirely material, and abstraction is just something human minds perceive.
This itself is a physical phenomenon :)
Multiple genetic materials operating on multiple levels...
Thanks for this. Came here to post a comment and saw this Very similar view, with a different term.
Another way to define this is:
1. The higher level object/thing has traits/properties that the parts do not have
2. Understanding each part does not necessarily = understanding these traits
3. The parts are interdependent
4. The traits/properties of the higher level is determined by the interaction of the parts
... which is the definition of a system.
Why not? Humans are products of evolution too, and they abstract.
I mean, what we consider conscious (e.g. human activity) and what we consider accidental (e.g. the result of evolutionary forces clashing) are not very well defined. Heck, we haven't even solved the determinism vs free will issue yet.
What does this mean? Are cells, for example, not abstractions?
In the same way the banana skin is a great packaging solution, but it wasn't created to serve as packaging for our consumption.
If this is true of a simple evolved circuit - i.e. the connection between the circuit specification and the behaviour being complex, with the circuit specification already encoding physics beyond the digital gates - then I'm not sure what hope we should actually place in statistical attempts to connect genes with organism behaviour.
Pushing the link a bit more, just as the circuit representation went beyond the target domain of digital circuits, could what is encoded in our genes explicitly encode exploitations at quantum level? With this, the possibility of brain structures acquiring quantum error correction schemes doesn't seem too far fetched or crazy thought.
[1] https://www.damninteresting.com/on-the-origin-of-circuits/
It reminds me of the halting problem and modern semantic analysis methods for programs. IE some things you can determine without running the code (static analysis of code or dna), and to determine other things you have to run the the code and see (determine if a program will ever end or see if a trait is expressed).
.. much the same as all of what the FPGA does is captured by electromagnetism + material physics. Nevertheless, the encoding of the FPGA circuit, while it describes a digital process, can hack into the features of a level below. Similar to that, while gene sequences may encode protein production, I'm wondering whether the resultant system could similarly make explicit use of below-the-protein-level features. Our inability to achieve that yet is not an argument, much the same as a digital circuit designer will be at a loss to make a "stop" versus "go" detector and that too one without a clock.
1. The theory of "junk DNA" of which the authors received Nobel Price is increasingly looking like...well...junk.
2. The recent discovery of Epigenetics adds a whole other dimension for understanding traits and genetics. Lots of biological science needs to be rewritten/updated in the years to come.
As new studies emerge on identical twins/triplets/etc., it has become clear that even identical DNA does not at all have 100% clone of the person, even if they do in fact share the same DNA. Ranging from things like height differences, leg lengths, intelligence levels, and a host of other characteristics, "identical" DNA does not mean physically identical creatures that look and behave exactly the same, even if they are far more similar than different.
New studies weren't needed; this has been known forever. From The Blank Slate:
> Even genetically homogeneous strains of flies, mice, and worms, raised in monotonously controlled laboratories, can differ from one another. A fruit fly may have more or fewer bristles under one wing than its bottlemates. One mouse may have three times as many oocytes (cells destined to become eggs) as her genetically identical sister reared in the same lab. One roundworm may live three times as long as its virtual clone in the next dish. The biologist Steven Austad commented on the roundworms' lifespans: "Astonishingly, the degree of variability they exhibit in longevity is not much less than that of a genetically mixed population of humans, who eat a variety of diets, attend to or abuse their health, and are subject to all the vagaries of circumstance -- car crashes, tainted beef, enraged postal workers -- of modern industrialized life." And a roundworm is composed of only 959 cells!
The genome is more like source code that evolved over millions of commits, and has code from way back when, for different use cases. It also has "injected code" ie ancient viral DNA, and other weird artifacts that's not relevant to the functioning code base of today. Sure, it is useful from a meta-evolutionary perspective, those parts can be refashioned for another use... but so can real-world "junk" :)
Source: https://www.crick.ac.uk/news/science-news/2018/06/14/non-cod...
EDIT: I think the finding is significant enough to warrant a separate discussion. Submitted a link: https://news.ycombinator.com/item?id=17361545
The quote is all over the internet. E.g., see http://egnorance.blogspot.com/2013/02/richard-dawkins-on-jun...
Honestly, I have a PhD, doing human genomics, and to me Dawkins is a popular science writer. I don't know anyone in the field of genomics who is an actual working scientist who believes in the 95% number.
Note that the Encode Project defines functional as being transcribed, which is not a widely accepted definition of functional. The whole idea of "function" itself brings a lot of subjective notions into play.
As I said in the parent, I also think "junk dna" is a terrible term. What I prefer stating is, in agreement with Ohno/Kimura's theory of neutral selection, the vast majority of our genome is not under positive or negative selection - i.e. mutations occurring there are neither detrimental, nor helpful, for fitness.
If you look at allele ratios of variants across the genome, the vast majority show evidence of neutral drift.
"So you can think of the protein-coding genes as being sort of the toolbox of subroutines which is pretty much common to all mammals -- mice and men have the same number, roughly speaking, of protein-coding genes and that's always been a bit of a blow to self-esteem of humanity. But the point is that that was just the subroutines that are called into being; the program that's calling them into action is the rest [of the genome] which had previously been written off as junk."
This makes sense to me (as a programmer :). But I wonder if someone managed to figure out the language in which the "caller" code is written. How does it say: call subroutine at position X, then if (condition), call subroutine at position Y, etc... I was not able to find any info on this.
Genes are transcribed by a polymerase, and a couple of factors determine if the polymerase can be "recruited" there. Here are a couple of those factors:
1) Is the site accessible? That part of the genome needs to be "open chromatin" for it to be transcribed. Next question of course is what makes some chromatin open some closed..
2) Histone signals: DNA spools around wheel-like things called histones. Histones get marked in a variety of ways to indicate "status" of that part of the genome. Some of them are pro-transcription.
3) Polymerase is usually recruited there by DNA-binding proteins called "transcription factors". These proteins bind to the regulatory / "promoter" region upstream of a protein-coding gene. There are over a thousand different transcription factors, each with a specific "motif" they recognize.
The global state is detected in various ways (receptors on the cellular mebrane etc.), which triggers these transcription factors, which then go bind DNA near the genes, to activate a new subroutine.
Here's a paper: http://www.pnas.org/content/107/20/9186 Title: "Comparing genomes to computer operating systems in terms of the topology and evolution of their regulatory control networks"
To explain it to a programmer, it's better to avoid the terminology and specifics, but focus only on functional relations - e.g. type of cell at point (x,y) is a FUNCTION of types (or something) of cells (x-1,y), (x+1, y) and maybe something else. I can't find it anywhere.
BTW, I am not a believer in cellular automation stuff a la Wolfram, just trying to frame the question in a way I could understand.
(just in case you don't want to reply here, my fake email address is in my profile)
It is a lot more like a cellular automaton than anything else. From minute gradient differences inside the fertilized egg, recursively structure spreads out in very much a cellular automata like fashion. Cells respond to the bio-mechanical stress around them and to the signals they bathe in, to decide cellular fate, and in turn release new signals..
It's a very different paradigm of computation than the one we think of. Each cell is like a docker instance, containing the full source code. Perhaps these links explain a bit:
https://www.youtube.com/watch?v=RQ6vkDr_Dec https://www.youtube.com/watch?v=3mCgHK-X6lE https://www.youtube.com/watch?v=EaZOdrU1Du0
The term junk was intended to convey something apparently worth keeping (perhaps in the attic), as opposed to garbage, which is thrown away.
At the same time, it seems like it's fairly clearly falsifiable almost on its face, because there are many genetic disorders involving clear, substantial mutations (in a chromosomal topography sense) that are circumscribed in their effects, at least to one extent or another.
I think some variant of this might be true, where traits are influenced by a very very large number of genes (like thousands or hundreds of thousands or more), possibly interacting or in more complex chaotic effects. But this is basically a variant of the polygenic hypothesis, which has been a major paradigm for years. This model seems like the dominant paradigm pushed to its limit, rather than something fundamentally new.
Maybe I'm misrepresenting my sense of the field, but I doubt that many people doing GWASes believe in a relatively limited number of genes affecting traits. I think this has been true for several years now at least. I've been wrong in my assumptions about the field before though.
"The "one gene = one protein" dogma of molecular biology (expanded: "one trait selectable in breeding experiments = one compact locus of the chromosome") was a priori wrong, but it's taken us decades to undo the damage. We were led astray because there are traits that do map to a single locus, or single mutations that were found to be able to control a trait, which is not the same thing as the trait."
It seems to me that this is an example of a common pattern in biological research. A simple explanation is developed for a phenomena, but then as time goes on, it turns out that in at least some important ways things are much more complicated than originally thought.
Also, the article talks about the controversy between the view that all genes make tiny causal contributions versus some are core and other peripheral.
But perhaps it is, at least in some cases, a matter of what I would call complex conditionality. So for instance maybe disease X happens if there are mutations in genes A, B and C, or C, D and E, or G, H and I, but not any other combination. Certainly we find things like that in other areas of reality.
There was no science in it, so I took a master’s, worked in industry, and am now doing real science in a CS PhD program.
I'm really glad that this mental model I've built up seems to reflect reality. The idea of neatly editing specific genes to produce a desired result always sat with me as the pinnacle of human hubris. The more I learn about Biology the more I just throw up my hands and say "magic".
If you took a bug report and did a statistical analysis of what transistors contributed to the bug, in most cases it would tell you "pretty much all of them". In very rare cases, it would tell you there were specific broken transistors. Similar for OS instructions. Because it's just the wrong level to be analyzing problems with the ultimate behavior of a system that's mostly functional. 99.9999% of the time, you don't have a compiler bug or an OS bug or a hardware bug, because if you did, everything would be dead.
My kneejerk reaction* is that genes are like transistors or OS instructions and not like application program instructions, and it really shouldn't be surprising or mysterious if they don't correspond to specific behaviors. It's just a different level of abstraction.
*having not studied biology and genetics, and just assuming living organisms have to be like computers.
A CPU or a single computer is still a fairly homogeneous, consistent creation. Sure, it has a lot of transistors, but it's just the same pattern over and over.
Think about the Internet: many of the billions of things on the internet are similar enough to be called the same, but many of them are ... a little not. This computer is Intel, this is AMD. This one runs Microsoft, this one Ubuntu. This one Ubuntu 15.07 with a 5.3.2.2 Linux kernel, that one with a 5.3.7.1 kernel. This part uses IPv4, that IPv6. This HTML on this site is a little nonstandard, that one relies on an expired SSL certificate.
The body, and modifying it, is a lot like the Internet. You set up 1.1.1.1 as a DNS server and accidentally partition the Internet.
To me, the pinnacle of human hubris is thinking you can correct the more challenging ones (like the ubergenetic phenotypes described here) without having some sort of powerful technology that helps work with the irreducible complexity of probabilistic chemical systems.
See, for example: Gene Therapy in a Patient with Sickle Cell Disease https://www.nejm.org/doi/full/10.1056/NEJMoa1609677
"Imagine a programming language where every single instruction simultaneously executes and simultaneously alters the inputs of all other instructions to varying degrees."
Complex artificial neural networks and data flow models are closer to gene regulatory networks and the genotype->ribotype->phenotype causal flow but are still oversimplifications.
It is sometimes possible to hack individual genes to fix really glaring problems reducible to one or a few key mutations like sickle cell anemia. Those things are the low hanging fruit of genetic engineering. Beyond that it gets harder. Trying to do things like boost life span or median IQ will be exponentially harder and will require new cognitive paradigms IMHO.
Its too bad nobody studies this stuff anymore. There was a ton of really great research into complex systems and emergent behaviors in the 90s and 2000s and then it went out of vogue for some reason.
If true, this means that evolution could look more like stochastic gradient descent, and less like brute force luck plus survival of the fittest.
Some say that the view it is pushing (core vs peripheral genes) is simply wrong as evidence of it cannot be found in the data (and there is a lot of data), see the end of the article about Naomi Wray's article. Also see the many mentions of the paper's critics! I like the phrase "provocatively phrased extension of earlier ideas".
There are bigger issues that need addressing but all of them are much much harder to get right and they would probably not make it to Quanta magazine ;) For example, phenotype heterogeneity and patient stratification are two big thorns in precision medicine's side.
Put another way, if you were building something and you wanted it to survive, would you want a flexible multifaceted system, and not a dogma-esque checklist of X, Y and Z.
Also, a trait, e.g. a psychological trait, or a group of them ("values-centered, empathetic, philosophical" vs. "productivity-centered, logical, process-oriented") could be seen as a technology orientation that helps the system build technology of a specific kind.
And who knows what we have yet to find. But it's making more and more sense in a why didn't we see this sooner sorta way.
This is like showing someone the hex dump of an mpeg file and expecting them to know it makes a picture of a kitten.
Wait until we start seeing the results of the hyped up "gene editing" studies where they follow up for a few years and look at more than one/few outcomes at a time...
Whatever we have known so far, early phases of the input system (sensory) and later phases of the output system (motor) can be localized. Even there is an overlap between the input and the output system.
fMRI is not useful in settling the disputes about how brain makes the mind--esp higher cognitive processes.
CRISPR and similar systems could someday be helpful for dealing with complex traits, but we first need to understand exactly what genes contribute to them and how, and for many traits, that will take a very long time. In some sense, CRISPR is just a dumb cutting tool. Knowing where to cut was always the hard part, and it always will be.
Are there any communities like HN, but focused on evolution, population genetics, etc. that you would recommend?
What is a bigger problem: Even an average idiot must now realize, that GM food may be riskier than you think.
Yes and no. It may lead to dangerous expressions, but plain chemical testing and analysis prior to commercial release should ensure nothing toxic makes it into people's food.
If your apples suddenly start producing magnitudes more cyanide after modifying something, just don't sell them.
And you know exactly under what conditions the toxins are expressed. Not.