The complete sequence of a human Y chromosome
nature.com
nature.com
Ref: https://www.technologyreview.com/2007/09/04/223919/craig-ven...
35 years ago: https://www.nytimes.com/1987/12/13/magazine/the-genome-proje...
33 years ago: https://www.nytimes.com/1990/06/05/science/great-15-year-pro...
29 years ago: the competition gets fierce https://www.nytimes.com/1994/02/22/science/scientist-at-work...
24 years ago: the race is alive! https://www.nytimes.com/1999/03/23/science/who-ll-sequence-h...
23 years ago: draft is complete https://archive.nytimes.com/www.nytimes.com/library/national...
20 years ago: https://www.nytimes.com/2003/04/15/science/once-again-scient...
2 years ago: https://archive.nytimes.com/www.nytimes.com/library/national...
The celera approach using shotgun was tricky because assembling the bits required a great deal of computational finesse and horsepower; "The final assembly computations were run on Compaq’s new AlphaServer GS160 because the algorithms and data required 64 gigabytes of shared memory to run successfully."
At the time the GS160 was the monster machine, a classic big iron UNIX, which could be tightly clustered- unified IO/filesystem between all the machines. The primary author of the shotgun assembly was Gene Myers, who previously had written BLAST, and invented the Suffix Array with Udi Manber.
The public project had its own issues, as many of the teams assigned to work on it still had the "cottage industry/artisinal/academic" approach, then Eric Lander came along and turned it into an industrial process, and parlayed that into running the Broad Institute, a privately funded MIT/Harvard research institute in Boston.
These days petabytes of sequence data are generated every day and stored in clouds. The genome has been a fundamental tool for shaping our studies of humans, although its true potential for understanding complex phenotypes remains elusive.
So... um... BLAST processing?
As someone who is in the genomics world now as a software person, I find it amusing that I've gotten more into perl over the last year or so. It has its place, in the way that grep/sed/awk/etc does.
IMHO the person who "saved the genome project" was WJ Kent, who developed the assembler, BLAT, that the public project needed. I strive to point out that he wasn't a sole hero, nor was Lincoln. What I really like about BLAT is that while Celera was using Big Iron UNIX (massive 64-bit 64GB machines with 10s of terabytes of central high performance storage), BLAT ran on a cluster of linux machines, right around the time that people were waking up to the fact that linux was becoming a useful tool for scientific data processing. BLAT's design allowed it to work on a cluster of cheaper/smaller machines, while celera's algorithm really needed a massive shared memory single machine. it was sort of a microcosm of the larger battle being fought between Big iron UNIX and little intel linux at the time.
smaller organizations were realizing they could get slower computers power pretty cheap, and they could get a whole lot of them. There was a little flurry of activity with flocking algorithms, objective-c had a little renaissance due to swarm computing. probably the most famous and lasting was map-reduce, from alphabet, but they were called google back then.
there were a bunch of clever little tricks, like channel bonded network cards to make multiple cards look like one fast card, so you could double or triple bandwidth.
The beowulf cluster joke was kind of the spirit of the scrappy, make something cool out of junk approach while calling companies like sun, hp, compaq dinosaurs. and it continued for a while with people building stuff like this - https://ncsa30.ncsa.illinois.edu/2003/05/ncsa-creates-sony-p... I think there was some weirdness with export controls of ps2's for this kind of stuff.
I don't remember what pc chip was the new hotness back then, it was a good 20 years ago and my memory is dim.
but that, I think, captures the gist of the meme.
However, I got into the genomics world in the early aughts, and the perl hung on and on and on. I remember starting a new job in the mid-teens, and one of the first tasks I had was to port over a legacy perl script. Such is the world of scientific software. 99% of the people I know who have touched perl since ~2003ish are in the bioinformatics space.
And if I remember correctly BLAT was so useful because it could be run on machines with less cpu power by loading more data into memory…or was it the other way around?
> The primary author of the shotgun assembly was Gene Myers
No conflict of interest there!And it is worth noting that despite all the critiques of shotgun sequencing at the time as "cheating", chromosome walking sequencing is dead, and shotgun is what we use today. If anything replaces that it will be long read technology like that done by Oxford Nanopore sequencing, although we aren't really there yet (you still need to assemble that data and don't really get end-to-end sequences yet).
Mother: "Amy, dear, you hardly touched any of your vegetables. Most of them are still on your plate."
Amy: "That's not true, Mom; I completely sequenced them into my tummy using the shotgun method."
And the cost of a T2T using a combination of these technologies is already well under $10,000 in reagent costs. The assembly is still a complex art especially for a messy chromosome like Y.
Does that mean that every possible polymorphism is accounted for?
Does that mean that 100% of possible polymorphic sites are known?
What about weird edge cases. You have a tandem repeat that is known, but in some small population of humans, there's two tandem repeats with an island of something else in there.
Imagine it like putting together a jigsaw puzzle, except that the picture on the front of the puzzle has lots of repeating motifs, so you end up with multiple pieces that look identical. You won't be able to tell how the whole thing fits together, but you can assemble the bits that are well-behaved and unique.
Modern technology gives us larger jigsaw pieces, which allows us to distinguish between almost-identical parts of the puzzle better. But I would note that the project linked did a huge amount of sequencing using very expensive methods to be able to resolve the whole thing.
Modern sequencing techniques allow us o read 200 to 500 bases at a time. So after hat we need do find a way o arrange these short sequences into a single sequence. And this can be prety hard, especially when you are doing 'de novo' assembly[0].
Besides that, there is the fact that some regions of DNA are repeated[1].
[0] - https://en.wikipedia.org/wiki/Third-generation_sequencing#De...
My knowledge cut-off period in this domain is around 2018. So it's not suprising that things moved on.
Because the sequencing uses short reads, it will not be able to resolve parts of the genome that are repetitive with a repeat unit longer than ~150bp. You'll have all the jigsaw pieces from the puzzle, but you won't be able to reconstruct some parts of the picture. Long read sequencing can help with those, but that's more expensive.
> The decoding is now close to 100% complete. The remaining tiny gaps are considered too costly to fill and those in charge of turning genomic data into medical and scientific progress have plenty to be getting on with.
OP is being silly. Nobody was fooling anyone.
http://news.bbc.co.uk/2/hi/science/nature/2940601.stm
> The remaining tiny gaps are considered too costly to fill and those in charge of turning genomic data into medical and scientific progress have plenty to be getting on with.
another "completion" happened 3 years ago, before this announcement. but this is the last one. I promise.
A gene sequence allows researchers to determine the amino acids that are coded for, and from those, which proteins match which genes.
This can be matched up with genetic diseases. If you know that damage to a certain location in a chromosome causes a problem with a certain biological process, then ergo, the associated protein is needed for that process!
So: genetic illness -> gene sequence -> protein -> role in the body
Without sequencing, that chain can't be built.
Is this why it’s so hard? This feels more like a healthcare records keeping people and less like an “actually reading the data problem”.
I can’t help but feel like some form of single payer healthcare is truly the way out of this problem. One where all disease record keeping is uniform and complete.
(Also our health system's IT is a hellscape, but one reason for that is that people would literally rather not have a working system at all, than one with less than impeccable privacy controls.
Personally I'd gladly sacrifice a fair bit of medical privacy in return for giving scientists greater insight into disease processes, but the average citizen here wants advanced healthcare without giving their data to research scientists. /facepalm )
I'm not saying a health IT system should have no privacy controls either. But the requirements for such controls need to be balanced against having a system that actually works, and that means having some people who actually understand the tech, and the workings of hospitals, having a role in requirements conversations. Instead it was dominated by MPs, "patient advocacy" groups and privacy campaigners, none of whom know or care anything about how to build a workable system.
No- A "gene" isn't an A/G/C/T- it's a sequence of 1000-1000000 base pairs. Each gene has a well-defined start/stop sequence called a start/stop codon. When people have genetic differences, one (an SNP- single nucleotide polymorph) of the tens of thousands of base pairs in that gene is different. Even for genes that are entirely "missing" in some people, they're really just different in a way that makes them nonfunctional.
Does that make it obvious how sequencing all those genes is useful, even if everyone has different genes? It tells us 99.999% of how proteins are coded, even if individual variation is the other .001%.
When we have baselines, we can compare different individuals, and eventually make predictions of how they are going to be based solely on the genetic code. If I know that a certain polymorphism is tied to some trait I want, I might not have to even bother spending the time growing a plant: I know that it's not what I want, and discard it as a seed.
With humans we are probably not going to see much modification soon, but just being able to detect genetic diseases, risk factors for other diseases that have genetic omponents, or allow for selection of embryos in cases of artificial insemination is already quite valuable.
It's not source code that we are all that good at understanding just yet, but there's already some applications, and we have good reason to think there's a lot more to come
That is if we transfer the DNA to an advanced alien civilization - would they be able to make a human.
You'd need a cell to start the process, with the various nucleic acids distributed correctly and proteins/energy with which to create further proteins using the information encoded by the DNA. Thus the civilization would need information about cells and a set of building blocks before being able to use the DNA.
Epigenetics is missing in this discussion about reproducing a human from just the DNA. These are superficial modifications (e.g. methylation, histone modification, repressor factors) to a strand of DNA that can drastically alter how specific regions get expressed into proteins. These mechanisms essentially work by either hiding or unhiding DNA from RNA polymerases and other parts of the transcription complex. These mechanisms can change throughout your lifetime because of environmental factors and can be inherited.
So it's like reading C source code, except there so many of these inscrutable C preprocessor directives strewn all throughout. You won't get a successful compilation by turning on or off all the directives. Instead, you need to get this similarly inscrutable configuration blob that tells you how to set each directive.
I guess in a way, it's like the weights for an ML model. It just works, you can't explain why it works, and changing this weight here produces a program that crashes prematurely, and changing a weight there produces a program with allergic reactions to everything.
If they just received the DNA without some information about the zygote, I don't think it would be practical for even advanced alien civilization (LR5 or LR6) but probably an LR7 and definitely an LR8 could.
Based on the way the person is using it, it does not seem to equate to the Kardashev scale, as my peer stated
Humans are currently LR2 (food security) and approaching LR3 (artificial general intelligence, self genetic modification). LR4 is generally associated with multiplanetary homing (IE, could survive a fatal meteor strike on the home planet) and LR5 with multisolar homing (IE, could survive a fatal solar incident). LR6 usually has total mastery of physical matter, LR7 can read remote multiverses, and LR8.2 can write remote multiverses. To the best of LR8's knowledge, there is no LR9, so far as their detectors can tell, but it would be hard to say, as LR9 implies existence in multiple multiverses simultaneously. Further, faster than light travel and time travel both remain impossible, so far as LR8 can tell.
“An Outside Context Problem was the sort of thing most civilisations encountered just once, and which they tended to encounter rather in the same way a sentence encountered a full stop.” ― Iain M. Banks, Excession
“Unbelievable. I’m in a fucking Outside Context situation, the ship thought, and suddenly felt as stupid and dumb-struck as any muddy savage confronted with explosives or electricity.” ― Iain M. Banks, Excession
“It was like living half your life in a tiny, scruffy, warm grey box, and being moderately happy in there because you knew no better...and then discovering a little hole in the corner of the box, a tiny opening which you could get your finger into, and tease and pull apart at, so that eventually you created a tear, which led to a greater tear, which led to the box falling apart around you... so that you stepped out of the tiny box's confines into startlingly cool, clear fresh air and found yourself on top of a mountain, surrounded by deep valleys, sighing forests, soaring peaks, glittering lakes, sparkling snow fields and a stunning, breathtakingly blue sky. And that, of course, wasn't even the start of the real story, that was more like the breath that is drawn in before the first syllable of the first word of the first paragraph of the first chapter of the first book of the first volume of the story.” ― Iain M. Banks, Excession
Or do you mean that a civilization at a particular level will always be unaware of civilizations above? That doesn't seem to make sense either; I see no reason why a LR4 civ couldn't have knowledge of a LR5 civ, for example.
I wonder if that information could ever really be untangled by a civilisation starting entirely from scratch without access to a cell
But… doing so from first principles without a mental model of how all (human) CPUs work? I guess it comes down to whether the recipients had enough context to know what they’re looking at.
In genomic science we nearly always use more cheaply available information rather than attempt to solve the hard problem directly. For example, for decades, a lot of sequencing only focused on the transcribed parts of the genome (which typically encode for protein), letting biology do the work for determining which parts are protein.
If you look at the process biophysically, you will see there are actual proteins that bind to the regions just before a protein, because the DNA sequences there match some pattern the protein recognizes. If you move that signal in front of a non-coding region, the apparatus will happily transcribe and even attempt to translate the non-coding region, making a garbage protein.
Coding DNA and non-coding DNA looks very different. Proteins are full of short repetitive sequences that form structural elements like alpha helixes: https://en.wikipedia.org/wiki/Alpha_helix
Once you've identified roughly where the protein-coding genes are it would be trivial to identify 3'/5' as being common to all those regions. You could pretty easily imagine a much more complicated system with different transcription mechanisms and codon categories, but earth genomes are super simple in that respect. Once you have those you just have the (incredibly complex) problem of creating a polymerase and bam, you'll be able to print every single gene in the body.
Without the right balance of promoters/factors/polymerase you probably won't get anything close to a human cell, but you'd be able to at least work closer to what the natural balance should be, and once you get closer to building a correct ribosome etc the cell would start to self-correct.
I'm really surprised that in all these responses to your question no one's mentioned the womb or the mother, who (at least with current technology) is still necessary for making a human.
That's not to mention the necessity of the egg.
We're not just DNA.
You could say “well that's the last 10% of the details, maybe 90% is in the DNA,” but I think I would be suspicious that it's that high, because one of the things we know about humans is that we are born with all of the ova that we will ever have, rather than deferring the process until puberty. I should think that if it could be deferred it would have been, “you will spend the energy to make these 15 years before you need to for no real reason” seems very unlike evolution whereas “my body is going to teach you how to make these eggs, just the same as my mother's body taught me,” sounds quite evolutionarily reasonable.
DNA is the blueprints. There are infinite possibilities what to do with them. The advanced civilization would need additional information, like that they are supposed create a cell from the components to begin with, and a lot of detailed information how exactly to do that.
Edit: improved clarity
ie. "software" without "machine" to run it on, is kind of a useless.
Or, you could build an embedding with far fewer parameters that could explain the vast majority of phenotypic differences. the genome is a hierarchical palimpsest of low entropy.
My standard interview question- because I hate leetcode- walks the interviewee through compressing DNA using bit encoding, then using that to implement a rolling hash to do fast frequency counting. Some folks get stuck at "how many bits in a byte", others at "if you have 4 symbols, how many bits are required to encode a symbol?", and other candidates jump straight to bloom filters and other probabilistic approaches (https://github.com/bcgsc/ntHash and https://github.com/dib-lab/khmer are good places to start if you are interested).
Except of course "genetics" and "environment" aren't actually separate things; sure, people's skin color isn't usually affected by their food, but only because most people don't eat colloidal silver.
Since genome files contain more data than just ATGC (typically a comment line, then a DNA line, then a quality score line), and each of those draws from a different distribution, DEFLATE on a FASTA file doesn't reach the full potential of the compressor because the huffman table ends up having to hold all three distributions, and the dictionary backlookups aren't as efficient either. It turns out you can split the file into multiple streams, one per line type, and then compress those independently, with slightly better compression ratios, but it's still not great.
In short, a lot of very painstaking genetics and molecular biology work has gone into characterizing the function of certain sequences.
In humans even though hervs don’t reactivate into infectious viruses they have been implicated in both harmful (senescence during aging[0]) and beneficial (protection from modern retroviruses)[1] activities in the body.
They might be up to 8% of the human genome.
If you had a list of all the genomes of all the people in the world, and all their phenotypes (height, eye color, hair type, etc), you could take all their genomes as input variables and treat all their phenotypes as output variables, and make embeddings or other models that mapped from genomes to phenotypes. The result would be a predictive model that could take a human genome, and spit out a prediction of what that person looks like and other details around them (up to the limits of heritability).
A good example is height. If you take a very large diverse sample of people, and sequence them, you will find that about 50% of the variance in height can be traced to the genomic sequence of that individual (other things, such as socioeconomic status, access to health care, pollution, etc, which are non-genomic, contribute as well). originally many geneticists believed that a small number of genes- tiny parts of the feature vector- would be the important features in the genome that explained height.
But it didn't turn out that way. Instead, height is a nonlinear function of thousands of different locations (either individual bases, entire genes, or other structures that vary between individuals) in the genome. This was less surprising to folks who are molecular biologists (mainly based on the mental models geneticists and MBers use to think about the mapping of genotype to phenotype), and we still don't have great mechanistic explanations of how each individual difference works in concert with all the others to lead to specific heights.
When I started out studying this some 35 years ago the problem sounded fairly simple, I assumed it would be easy to find the place in my genome that led to my funny shaped (inherited) nose, but the more I learn about genomics and phenotypes, the more I appreciate that the problem is unbelievably complex, and really well suited to large datasets and machine learning. All the pharma have petabytes of genome sequences in the cloud that they try hard to analyze but the results are mixed.
I spent my entire thesis working on ATGCAAAT, by the way. https://en.wikipedia.org/wiki/Octamer_transcription_factor is a family of proteins that are incredibly important during growth and development. Your genome is sprinkled with locations that contain that sequence- or ones like it- that are used to regulate the expression of proteins to carry out the development plan.
The really interesting approaches these days, IMHO, combine genomics and microscopic imaging of organoids, and many folks are trying to set up a "lab in the loop", in which large-scale experiments run autonomously by sophisticated ML systems could accelerate discovery. It's a fractally complex and challenging problem.
Statistics has been key to understanding genetics from the beginning (see Mendel, Fisher) and so at a big pharma you will see everything from Bayesian bootstrappers using R to deep learners using pytorch.
Are there any positions at Google/ companies you wold suggest me to look into? I'm coming from algortrading/ ML research with ML MSc.
[1] https://www.genomicsengland.co.uk/blog/data-representations-...
Would such a predictive model really be possible? As far as I'm aware there is contradicting research whether a specific phenotype distinctly originates from a SNP/genotype.
The model would be highly nonlinear and nonlocal, at the very least.
Almost all variations that humans have in their genomes (compared to each other or a reference genome) are tiny, mostly one base differences called single nucleotide polymorphisms (SNPs). These tiny changes encode who you are. The rest of it just makes you carbon-based organism, a eukaryote, an animal, a mammal etc, just like a whole load of other organisms.
https://www.nytimes.com/2002/04/27/us/scientist-reveals-secr...
I believe more recent sequencing projects have used a wider pool of individuals. I think some projects pool all the individuals and sequence them together, while others sequence each individual separately. This isn't really so much of a problem since the large-scale structure is highly similar across all humans and we have developed sophisticated approaches to model the variations in individuals, see https://www.biomedcentral.com/collections/graphgenomes for an explanation of the "graph structure" used to reprsent alternatives in the reference, which can include individual single nucleobase differences, as well as more complex ones such as large deletions in one individual, to rearrangements and even inversions.
Ideally we'd have hundreds if not thousands of complete genomes which in total would reveal the population diversity of the human species as it currently exists, but this is a big ask. "Clincal variants" are of particular interest as those are regions of the genome associated with certain inherited diseases, although the promises of individual genomic knowledge leading to a medical revolution have turned out to be wildly overblown.
Since the paper is paywalled, there's not much else to say than that they have a (fairly arbitrary in origin, i.e. it could have been from any one individual or possibly even a chimera of several individuals) reference sequence to which other specific human Y chromosomes can be compared, eventually leading to a larger dataset from many individuals which will reveal the highly conserved and highly variable regions of the chromosome, population-wise.
It is not perfect, as a references can be missing or have large variability in DNA regions. The goal of the Human Pangenome Reference Consortium (HPRC) https://humanpangenome.org/ is to sequence individuals from different populations to address this issue. We are also working to develop new computation models to support analysis of data across populations.
Let's not pretend that this is not an end goal. It always was.
How do we know this, if we have only sequenced the chromosome of one individual?
A pangenome is a complex graph model that weaves together hundreds or more genomes/haplotypes—usually of one species, but the idea can extend across species too, or even cells within one individual (think cancer pangenomes).
On the idealized human pangenome graph each human is represented by two threads along each autosome, plus threads through Chr X, Y, and the mitochondrial genome.
And yeah, you're right: All modern large-scale sequencing is shotgun sequencing where the DNA is randomly broken up, each fragment is sequenced, and then the individual segments (reads) are assembled using a genome assembler.
We all have different DNA. So is "the human genome" some kind of "average" DNA, or is it the DNA of whoever they sampled, or is it maybe an overview of what is common for all of us?
If you want a population model of genomes you need a pangenome.
See "pangenome graph", "variation graph", and the human pangenome project.
The latest/greatest end-to-end T2T reference is entirely based on 'HG002' an individual from Utah, due partly to new information derived from long read technologies.
The T2T team published a preprint [1] last December and released the data [2] in March. However, due to the peer review process, the findings have only just been formally published in Nature. The publication timeline can indeed be slow, and in cases like this one, the question is: what's the point when all scientists interested in the topic already know about it and working with this assembly?
[1] https://www.biorxiv.org/content/10.1101/2022.12.01.518724v1.... [2] https://github.com/marbl/CHM13
In this context, I would say the point of a press release is getting the news out generally to non-scientists.
We understand some CNC instructions, but there's no good understanding of the control mechanisms or sequencing.
One day we'll identify the keys, but we are very far from the driver, and more importantly, from identifying what each key does.
Oh, we did discover that changing specific bits in the raw data causes the keyboard to fail in interesting ways.
Short-read sequencing data is a notoriously bad datatype for reconstructing the low-complexity / repetitive regions of genomes, so up until recently the most commonly used reference genomes have left many of these regions "dark". According to the preprint, the Y chromosome has the highest density of these low-complexity regions. It's also something of a bioinformatic nuisance when constructing a generic human reference genome, as it's only present in 50% of the population.
I wouldn't call random data 'complex', but it is easy to sequence when assembling short reads.
I did it with COI gene, which is just a short (1000ish base pairs with our snails IIRC) sequence of purely random ATGC base pairs. Lots of unique sequences make the short strands easy to match, just get a bunch of 10-15 BP bits and you can match the whole thing.
Now if your gene is 62M BP of repeating palindrome sequences, you can imagine how hard it would be to align random pieces sequenced as it will be very hard to find unique sequences to match.
By difficult, imagine a jigsaw puzzle. Difficult bits in the jigsaw puzzle are where you have the same sub-image repeated over and over again, or where the same image section is repeatedly scattered over the wider image. Puzzle pieces from these bits are difficult because you can't tell which part of the image such a small puzzle piece comes from, because it matches multiple places. It's technically impossible to resolve repetitive features where the repeating unit is larger than the pieces you are trying to assemble together. Modern technology gives us long read sequencing, where sequenced sections of DNA may be up to 100kbp or larger (maybe up to a 1Mbp), where older sequencing methods (that are still heavily used because they are cheaper) give us sequenced sections of DNA between 100-300bp long. (bp stands for base-pairs - one of [ACGT].) These larger puzzle pieces allowed the whole picture to be assembled without ambiguity.
However, this isn't why the Y chromosome was solved last. The reason for this is that the other chromosomes were analysed using a completely homozygous hydatidiform mole, which is where cells generate two copies of their entire genome from just one copy during conception, and therefore the two copies are identical. It makes the sequencing a lot easier if you don't have to deal with having two copies of the DNA that are slightly different. The side-effect of that is the hydatidiform mole doesn't have a Y chromosome, so they had to analyse a different sample later on to get a Y chromosome.
I am thrilled to see more chromosomes being mapped/sequenced. Please excuse my high-school level of biology knowledge here, but have we definitively progressed beyond correlation when it comes to genes, gene expression, and how they all interact?
Take "Blue eyes" as an example, we know that [1] OCA2 is responsible for brown/blue eye colour. BUT are we sure that none of the 20,000 others are involved/needed?
In laymans terms I would guess that to get 'blue eyes' would require several genes to ALL be on/active/present as well. More like a recipe, then a simple on/off it's located in 1 position.
[0] https://www.goodreads.com/work/editions/436071-y-the-descent... [1] https://www.nature.com/articles/s41433-021-01749-x
No, we are not completely sure how many chromosomes play a role in determining eye color. However, we do have a pretty good guess. Most recent estimates I found out the number at 16. OCA2 and HERC2 have the largest impact on eye color, but, there are many OTHER genes that also have smaller impacts on eye color. [2] That article I cited is actually amazing. But, it does touch on some more advanced subjects that are more introductory college-level biology or AP Biology than standard high school bio e.g. Gene Regulation, introns and extrons, etc.
To answer your more general question, in my (admittedly only mildly less basic, introductory college-level biology) opinion, it is unlikely that we will, anytime soon, reach a point where anything in genetics can be completly, 100% definitive. This is not to say that we haven't made amazing advancements in the field of genetics and biology more broadly. BUT, it would be a mistake to take for granted the complexity of the human genome. There is most definitely things we still do not understand about the genome, and will not for some time.
But, practically speaking, while other genes may have impact on some simple phenotypic traits, such as eye color, we can generally make an accurate guess based on only a few genes. In eye color, for example, one study was able to predict eye color with only 6 genes with about ~75% accuracy. [3]
The crazy part is, sometimes the genes that effect the phenotypic trait, don't actually store genetic material that determines the trait. In other words, they don't DIRECTLY determine the trait at all. Rather, the only effect ORHER genes which then effect the expression of the trait more directly. You can imagine that this could become very complex very very quickly when you have multiple genes effecting other genes which effect other genes, not to mention accoubting for environmental and demographic baises when doing these studies and you begin to see why this genetics is such a difficult field of study.
Apologies for the long post and rambling. Hopefully I was still able to provide you with some mediocre introductory-collage-level biology
[1] https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2791696/ [2] https://www.nature.com/articles/jhg2010126 [3] https://doi.org/10.1016/j.cub.2009.01.027
BTW. if you want to know the applications of this work, have a look at this ACM SIGPLAN Keynote: https://youtu.be/JTU3JYp3JYc?si=jOZz611ATQar3Gec (helped me understand DNA more than all biology classes at my high school)
2) This incomplete sequence has nevertheless been VERY useful, because these repeats mostly are outside of protein encoding regions ('genes').
3) 'The' human genome is of course non existent because you and I have a different one. What is meant is that this has been done for a 'bunch' of individual humans.
4) Because we're much more alike than we are different in our genome, this is still very informative.
That's insanely cool, but very far from what you are asking. An analogy: We have re-installed the operating system of an existing computer, when you are asking whether we can manufacture a computer.
It's like if you had a program to 3D print a 3D printer but you don't have a 3D printer that can read that program.
It's all ridiculously complicated. To get an idea about it you could learn about how simple RNA viruses like influenza or HIV replicate. It's easy enough for anyone to understand. But no single person can fully understand humans (or other mammals, plants etc, we're not special in that respect).
I presume our DNA are not identical. So do they sequence a particular person's Y chromosome? Who's?
How similar are two Y chromosomes from different persons?
The person they picked is pseudonymously HG002, who is an Ashkenazi man who took part in the project and consented to commercial distribution of his genome.
https://gnomad.broadinstitute.org/news/images/2018/10/gnomad...
Scientists: here you go
How can someone have their whole (!) genome sequenced already when so far we weren't able to fully sequence the Y chromosome. And this person seems to have a Y chromosome.
This is shaping up to be a very interesting decade to be in tech.
We might actually be able to generate new sequences and test them against all the diseases to remove them.
It's the only way that 2023 Gender Science can be true, and all good community members know that it is true.
All Glory to the Commune Leader!
/s