https://www.nytimes.com/2002/04/27/us/scientist-reveals-secr...
I believe more recent sequencing projects have used a wider pool of individuals. I think some projects pool all the individuals and sequence them together, while others sequence each individual separately. This isn't really so much of a problem since the large-scale structure is highly similar across all humans and we have developed sophisticated approaches to model the variations in individuals, see https://www.biomedcentral.com/collections/graphgenomes for an explanation of the "graph structure" used to reprsent alternatives in the reference, which can include individual single nucleobase differences, as well as more complex ones such as large deletions in one individual, to rearrangements and even inversions.
It is not perfect, as a references can be missing or have large variability in DNA regions. The goal of the Human Pangenome Reference Consortium (HPRC) https://humanpangenome.org/ is to sequence individuals from different populations to address this issue. We are also working to develop new computation models to support analysis of data across populations.
A gene sequence allows researchers to determine the amino acids that are coded for, and from those, which proteins match which genes.
This can be matched up with genetic diseases. If you know that damage to a certain location in a chromosome causes a problem with a certain biological process, then ergo, the associated protein is needed for that process!
So: genetic illness -> gene sequence -> protein -> role in the body
Without sequencing, that chain can't be built.
Is this why it’s so hard? This feels more like a healthcare records keeping people and less like an “actually reading the data problem”.
I can’t help but feel like some form of single payer healthcare is truly the way out of this problem. One where all disease record keeping is uniform and complete.
(Also our health system's IT is a hellscape, but one reason for that is that people would literally rather not have a working system at all, than one with less than impeccable privacy controls.
Personally I'd gladly sacrifice a fair bit of medical privacy in return for giving scientists greater insight into disease processes, but the average citizen here wants advanced healthcare without giving their data to research scientists. /facepalm )
I'm not saying a health IT system should have no privacy controls either. But the requirements for such controls need to be balanced against having a system that actually works, and that means having some people who actually understand the tech, and the workings of hospitals, having a role in requirements conversations. Instead it was dominated by MPs, "patient advocacy" groups and privacy campaigners, none of whom know or care anything about how to build a workable system.
No- A "gene" isn't an A/G/C/T- it's a sequence of 1000-1000000 base pairs. Each gene has a well-defined start/stop sequence called a start/stop codon. When people have genetic differences, one (an SNP- single nucleotide polymorph) of the tens of thousands of base pairs in that gene is different. Even for genes that are entirely "missing" in some people, they're really just different in a way that makes them nonfunctional.
Does that make it obvious how sequencing all those genes is useful, even if everyone has different genes? It tells us 99.999% of how proteins are coded, even if individual variation is the other .001%.
When we have baselines, we can compare different individuals, and eventually make predictions of how they are going to be based solely on the genetic code. If I know that a certain polymorphism is tied to some trait I want, I might not have to even bother spending the time growing a plant: I know that it's not what I want, and discard it as a seed.
With humans we are probably not going to see much modification soon, but just being able to detect genetic diseases, risk factors for other diseases that have genetic omponents, or allow for selection of embryos in cases of artificial insemination is already quite valuable.
It's not source code that we are all that good at understanding just yet, but there's already some applications, and we have good reason to think there's a lot more to come
That is if we transfer the DNA to an advanced alien civilization - would they be able to make a human.
You'd need a cell to start the process, with the various nucleic acids distributed correctly and proteins/energy with which to create further proteins using the information encoded by the DNA. Thus the civilization would need information about cells and a set of building blocks before being able to use the DNA.
Epigenetics is missing in this discussion about reproducing a human from just the DNA. These are superficial modifications (e.g. methylation, histone modification, repressor factors) to a strand of DNA that can drastically alter how specific regions get expressed into proteins. These mechanisms essentially work by either hiding or unhiding DNA from RNA polymerases and other parts of the transcription complex. These mechanisms can change throughout your lifetime because of environmental factors and can be inherited.
So it's like reading C source code, except there so many of these inscrutable C preprocessor directives strewn all throughout. You won't get a successful compilation by turning on or off all the directives. Instead, you need to get this similarly inscrutable configuration blob that tells you how to set each directive.
I guess in a way, it's like the weights for an ML model. It just works, you can't explain why it works, and changing this weight here produces a program that crashes prematurely, and changing a weight there produces a program with allergic reactions to everything.
If they just received the DNA without some information about the zygote, I don't think it would be practical for even advanced alien civilization (LR5 or LR6) but probably an LR7 and definitely an LR8 could.
Based on the way the person is using it, it does not seem to equate to the Kardashev scale, as my peer stated
Humans are currently LR2 (food security) and approaching LR3 (artificial general intelligence, self genetic modification). LR4 is generally associated with multiplanetary homing (IE, could survive a fatal meteor strike on the home planet) and LR5 with multisolar homing (IE, could survive a fatal solar incident). LR6 usually has total mastery of physical matter, LR7 can read remote multiverses, and LR8.2 can write remote multiverses. To the best of LR8's knowledge, there is no LR9, so far as their detectors can tell, but it would be hard to say, as LR9 implies existence in multiple multiverses simultaneously. Further, faster than light travel and time travel both remain impossible, so far as LR8 can tell.
“An Outside Context Problem was the sort of thing most civilisations encountered just once, and which they tended to encounter rather in the same way a sentence encountered a full stop.” ― Iain M. Banks, Excession
“Unbelievable. I’m in a fucking Outside Context situation, the ship thought, and suddenly felt as stupid and dumb-struck as any muddy savage confronted with explosives or electricity.” ― Iain M. Banks, Excession
“It was like living half your life in a tiny, scruffy, warm grey box, and being moderately happy in there because you knew no better...and then discovering a little hole in the corner of the box, a tiny opening which you could get your finger into, and tease and pull apart at, so that eventually you created a tear, which led to a greater tear, which led to the box falling apart around you... so that you stepped out of the tiny box's confines into startlingly cool, clear fresh air and found yourself on top of a mountain, surrounded by deep valleys, sighing forests, soaring peaks, glittering lakes, sparkling snow fields and a stunning, breathtakingly blue sky. And that, of course, wasn't even the start of the real story, that was more like the breath that is drawn in before the first syllable of the first word of the first paragraph of the first chapter of the first book of the first volume of the story.” ― Iain M. Banks, Excession
Or do you mean that a civilization at a particular level will always be unaware of civilizations above? That doesn't seem to make sense either; I see no reason why a LR4 civ couldn't have knowledge of a LR5 civ, for example.
I wonder if that information could ever really be untangled by a civilisation starting entirely from scratch without access to a cell
But… doing so from first principles without a mental model of how all (human) CPUs work? I guess it comes down to whether the recipients had enough context to know what they’re looking at.
In genomic science we nearly always use more cheaply available information rather than attempt to solve the hard problem directly. For example, for decades, a lot of sequencing only focused on the transcribed parts of the genome (which typically encode for protein), letting biology do the work for determining which parts are protein.
If you look at the process biophysically, you will see there are actual proteins that bind to the regions just before a protein, because the DNA sequences there match some pattern the protein recognizes. If you move that signal in front of a non-coding region, the apparatus will happily transcribe and even attempt to translate the non-coding region, making a garbage protein.
Coding DNA and non-coding DNA looks very different. Proteins are full of short repetitive sequences that form structural elements like alpha helixes: https://en.wikipedia.org/wiki/Alpha_helix
Once you've identified roughly where the protein-coding genes are it would be trivial to identify 3'/5' as being common to all those regions. You could pretty easily imagine a much more complicated system with different transcription mechanisms and codon categories, but earth genomes are super simple in that respect. Once you have those you just have the (incredibly complex) problem of creating a polymerase and bam, you'll be able to print every single gene in the body.
Without the right balance of promoters/factors/polymerase you probably won't get anything close to a human cell, but you'd be able to at least work closer to what the natural balance should be, and once you get closer to building a correct ribosome etc the cell would start to self-correct.
I'm really surprised that in all these responses to your question no one's mentioned the womb or the mother, who (at least with current technology) is still necessary for making a human.
That's not to mention the necessity of the egg.
We're not just DNA.
You could say “well that's the last 10% of the details, maybe 90% is in the DNA,” but I think I would be suspicious that it's that high, because one of the things we know about humans is that we are born with all of the ova that we will ever have, rather than deferring the process until puberty. I should think that if it could be deferred it would have been, “you will spend the energy to make these 15 years before you need to for no real reason” seems very unlike evolution whereas “my body is going to teach you how to make these eggs, just the same as my mother's body taught me,” sounds quite evolutionarily reasonable.
DNA is the blueprints. There are infinite possibilities what to do with them. The advanced civilization would need additional information, like that they are supposed create a cell from the components to begin with, and a lot of detailed information how exactly to do that.
Edit: improved clarity
ie. "software" without "machine" to run it on, is kind of a useless.
Or, you could build an embedding with far fewer parameters that could explain the vast majority of phenotypic differences. the genome is a hierarchical palimpsest of low entropy.
My standard interview question- because I hate leetcode- walks the interviewee through compressing DNA using bit encoding, then using that to implement a rolling hash to do fast frequency counting. Some folks get stuck at "how many bits in a byte", others at "if you have 4 symbols, how many bits are required to encode a symbol?", and other candidates jump straight to bloom filters and other probabilistic approaches (https://github.com/bcgsc/ntHash and https://github.com/dib-lab/khmer are good places to start if you are interested).
Except of course "genetics" and "environment" aren't actually separate things; sure, people's skin color isn't usually affected by their food, but only because most people don't eat colloidal silver.
Since genome files contain more data than just ATGC (typically a comment line, then a DNA line, then a quality score line), and each of those draws from a different distribution, DEFLATE on a FASTA file doesn't reach the full potential of the compressor because the huffman table ends up having to hold all three distributions, and the dictionary backlookups aren't as efficient either. It turns out you can split the file into multiple streams, one per line type, and then compress those independently, with slightly better compression ratios, but it's still not great.
In short, a lot of very painstaking genetics and molecular biology work has gone into characterizing the function of certain sequences.
In humans even though hervs don’t reactivate into infectious viruses they have been implicated in both harmful (senescence during aging[0]) and beneficial (protection from modern retroviruses)[1] activities in the body.
They might be up to 8% of the human genome.
If you had a list of all the genomes of all the people in the world, and all their phenotypes (height, eye color, hair type, etc), you could take all their genomes as input variables and treat all their phenotypes as output variables, and make embeddings or other models that mapped from genomes to phenotypes. The result would be a predictive model that could take a human genome, and spit out a prediction of what that person looks like and other details around them (up to the limits of heritability).
A good example is height. If you take a very large diverse sample of people, and sequence them, you will find that about 50% of the variance in height can be traced to the genomic sequence of that individual (other things, such as socioeconomic status, access to health care, pollution, etc, which are non-genomic, contribute as well). originally many geneticists believed that a small number of genes- tiny parts of the feature vector- would be the important features in the genome that explained height.
But it didn't turn out that way. Instead, height is a nonlinear function of thousands of different locations (either individual bases, entire genes, or other structures that vary between individuals) in the genome. This was less surprising to folks who are molecular biologists (mainly based on the mental models geneticists and MBers use to think about the mapping of genotype to phenotype), and we still don't have great mechanistic explanations of how each individual difference works in concert with all the others to lead to specific heights.
When I started out studying this some 35 years ago the problem sounded fairly simple, I assumed it would be easy to find the place in my genome that led to my funny shaped (inherited) nose, but the more I learn about genomics and phenotypes, the more I appreciate that the problem is unbelievably complex, and really well suited to large datasets and machine learning. All the pharma have petabytes of genome sequences in the cloud that they try hard to analyze but the results are mixed.
I spent my entire thesis working on ATGCAAAT, by the way. https://en.wikipedia.org/wiki/Octamer_transcription_factor is a family of proteins that are incredibly important during growth and development. Your genome is sprinkled with locations that contain that sequence- or ones like it- that are used to regulate the expression of proteins to carry out the development plan.
The really interesting approaches these days, IMHO, combine genomics and microscopic imaging of organoids, and many folks are trying to set up a "lab in the loop", in which large-scale experiments run autonomously by sophisticated ML systems could accelerate discovery. It's a fractally complex and challenging problem.
Statistics has been key to understanding genetics from the beginning (see Mendel, Fisher) and so at a big pharma you will see everything from Bayesian bootstrappers using R to deep learners using pytorch.
Are there any positions at Google/ companies you wold suggest me to look into? I'm coming from algortrading/ ML research with ML MSc.
[1] https://www.genomicsengland.co.uk/blog/data-representations-...
Would such a predictive model really be possible? As far as I'm aware there is contradicting research whether a specific phenotype distinctly originates from a SNP/genotype.
The model would be highly nonlinear and nonlocal, at the very least.
Almost all variations that humans have in their genomes (compared to each other or a reference genome) are tiny, mostly one base differences called single nucleotide polymorphisms (SNPs). These tiny changes encode who you are. The rest of it just makes you carbon-based organism, a eukaryote, an animal, a mammal etc, just like a whole load of other organisms.
A pangenome is a complex graph model that weaves together hundreds or more genomes/haplotypes—usually of one species, but the idea can extend across species too, or even cells within one individual (think cancer pangenomes).
On the idealized human pangenome graph each human is represented by two threads along each autosome, plus threads through Chr X, Y, and the mitochondrial genome.
How do we know this, if we have only sequenced the chromosome of one individual?
Ideally we'd have hundreds if not thousands of complete genomes which in total would reveal the population diversity of the human species as it currently exists, but this is a big ask. "Clincal variants" are of particular interest as those are regions of the genome associated with certain inherited diseases, although the promises of individual genomic knowledge leading to a medical revolution have turned out to be wildly overblown.
Since the paper is paywalled, there's not much else to say than that they have a (fairly arbitrary in origin, i.e. it could have been from any one individual or possibly even a chimera of several individuals) reference sequence to which other specific human Y chromosomes can be compared, eventually leading to a larger dataset from many individuals which will reveal the highly conserved and highly variable regions of the chromosome, population-wise.
Let's not pretend that this is not an end goal. It always was.
another "completion" happened 3 years ago, before this announcement. but this is the last one. I promise.