Ctcctcggcgggcacgtag is unique to both SARS-CoV-2 and Moderna's 2015/2016 patents
arkmedic.substack.com
arkmedic.substack.com
Edit: to elaborate, certain amino acid sequences are seen again and again, and others are practically impossible to see in nature because amino-acids fold into 3D proteins...and each amino acid is like a Lego piece. There is a lot more detail needed to understand fully how tertiary and 4try protein structures form...but it is also possible to understand it decently well after an undergraduate degree in biology, chemistry, biochemistry, etc. A PhD in the field knows this inside and out. If I had to guess, the author of this post has a PhD in something like statistics or computer science, and thinks that they can apply high school level math to two fields (molecular biology and biochemistry) that they do not even have a high school level of understanding of.
To make a very stretched analogy (I am a doctoral student in the life sciences and a hobby programmer), this blog post is like saying that some Java source code is stolen because four of the class names are the same between two projects, and then using the number of possible characters in the class names and the length of the name to run some statistics. The problem being that no one has class names like "AeNOQ92bA"...in fact the vast majority of 8 character sequences will never be class names. Just like the vast majority of amino acid sequences (likely) do not exist in nature. And then you dig deeper and find out the class name is something like "MainDashboardSupervisorTree" (I do not know Java so forgive me). Then you also find out that the author of both classes...is the same guy who moved companies and likes particular naming convection, but never meant to copy stuff word for word. Similar to how SARS-CoV-2 could have naturally incorporated HIV-1 RNA into its genome when co-infecting a host.
Hell, the entire genetic code of every variant of COVID-19 exists somewhere in the digits of Pi, but that doesn't mean that mathematicians created COVID. This reeks of a slightly more advanced form of numerology to me.
The reality is that average people simply aren't qualified to deal with the subtlety and complexity of ambiguous data, and the adversaries of public health can just make a video that lies and convince millions of people with seductive, but wrong, ideas. It's extremely rare to have a society which is scientifically savvy enough to understand these sorts of things.
(Preemptive) Synonymous mutations are able to be selected, as they affect speed of translation and protein folding due to stalled ribosomes.
Given the number of bats in the world, the number of human hosts, and probably most importantly, the number of individual virions within each infected host (and therefore the number of replication cycles) is...astronomically high. There are somewhere between 1 and 100 BILLION virions of COVID in each infected person. Now imagine a few thousand bats infected...we are already talking about maybe 1 QUADRILLION different virions (1,000 trillion). It only takes one virion to incorporate some very handy and fitness-increasing HIV-1 RNA into its genome and it is off to the races.
I truly have no idea what the answer is, but as someone well versed in economics, econometrics,and statistics, to me the relevant odds are not about independent coin tosses etc. I would like to know P(A|B) where:
A: sequence of 30 nucleotides appears in a virus identified in nature
B: sequence appears in a patent application that predates discovery in nature
Seems to me that would involve a whole bunch of arithmetic, but that ought to be calculable using this database.
So yeah. Credentials do matter, and it matters that this person is not willing to say this out loud, in public, without using a pseudonym.
The problem is that we've come so far that him not saying this out loud is not necessarily hinting to him being a charlatan. Plausibly, the cost/benefit analysis by that individum might not make it worth speaking the truth given how heated up the matter is rn. If this sentiment propagates to other topics I fear that larger (and important) parts of the world might drift towards soviet era incompetence
Reputable journals and other publication outlets use something called "Double Blind peer review" precisely to prevent that the reputation of a researcher could skew the peer review process.
If you want to review and cast a judgement for the points presented in the article, you should do it only by refuting or confirming the content of the text itself. Not because it was written by Einstein or Donald Trump.
Double blind peer review is INCREDIBLY UNCOMMON among high impact biomedical journals
Second, you're missing the key difference, here: peer review is for review and evaluation within the community of qualified scientists.
What I'm talking about is something else: the ability of the reasonably well-informed public to evaluate claims, even though they lack domain expertise.
Anyway the blinding, when it exists, is as much to protect the reviewers from retribution by asshole authors, as it is to avoid biasing the reviewers by the reputation of the authors.
(Musing: What is the probability that a genetics expert's opinion on the probabilities at play would make the discussion less well-informed rather than more well-informed?)
Why, according to you are they worth and not worth examining?
> It’s easy enough to change a single nucleotide (a single point mutation or SNP) or even insert or delete nucleotides (less common) but to insert 20 or 30 nucleotides with a code that works? Nope, that has to come from another virus or else it’s been done in a lab.
The original preprint which noted similarity of the HIV-1 sequences and those in the COVID spike protein, and on which this blog author's claims rest, was voluntarily withdrawn by its authors. [1]
Moreover, the assertions of uniqueness have been thoroughly addressed in the published article, "HIV-1 did not contribute to the 2019-nCoV genome" [2] :
> For any virus to obtain additional insert sequences from other organisms, it requires that it has direct interactions with other organisms, most likely through homologous or non-homologous recombination. [...] On the contrary, these motifs are widely present in various mammalian cells and so it will be more likely for bat CoV viruses to gain those motifs from the genomes of their infected cells if recombination indeed occurs.
That rebuttal article clearly suggests several feasible sources of the gene in Nature. That surely is more plausible than suggesting that Moderna or some other group nefariously created the COVID genome _in vitro_ before unleashing it upon humanity.
I deeply enjoy Hacker News for its ability to surface different ideas and sources of information that I would otherwise never find, but this article is simply ridiculous.
[0] https://en.wikipedia.org/wiki/Horizontal_gene_transfer
[1] https://www.biorxiv.org/content/10.1101/2020.01.30.927871v2
[2] https://www.ncbi.nlm.nih.gov/labs/pmc/articles/PMC7033698/
This is pretty common. The reason I take it as obligation is that if there aren't people fighting disinformationn, the disinformation will win, because it tends to be full of simple, seductive, but. wrong ideas that naive people fall for.
I'm not sure which particular thing public health people said that you're complaining about in the last part of the section- but I don't really think that public health leaders who communicate things which aren't 100% correct are dissemination disinformation. That's sort of a pedantic and uninteresting argument to have about the philosophical nature of information.
Even without spending time my heuristics is that I find biases in the claim and find out how referenced it is. This one had plenty red flags, including dismissing the RaTG13 results claiming that they were just put into the database after people started questioning, the adversarial tone, presumptions and so on. Conspiracy theorists use recognizable heuristics, but their motivation is to seek attention, for which there is large competition. They want to be relevant, but who doesn't. Conspiracy theorists are not really into publishing a paper either, a blog post, video or tweet is all you get, it's not about the truth and it is easier to avoid being found out this way, though social media these days have fact checkers, so they have to fly under the radar or keep to the social circles that are not bothered by fact checkers.
Knowing about biases and how to recognize them is a vital skill, should be taught in schools. Everyone needs a working bullshit meter.
To be fair, that sequence came from the Wuhan Institute of Virology
It is in the Moderna 2015 patent. 19 nucleotide match.
This would be a particularly useful discovery, because it would dramatically limit the candidates for who did the gain of function research to organizations licensed this Moderna lineage.
There also seems to be some circular reasoning the argument. Apparently we can ignore RaTG13 because it’s obviously synthetic, which makes SARS-CoV-2 look even more synthetic. It would be interesting to compare to the BANAL family of SARS-CoV-2-related viruses that are even more closely related to SARS-CoV-2 than RaTG13 [1].
I’m not sure why only viral genomes were searched for the furin cleavage site sequence. Viruses famously exchange genetic material with their host organisms. The “smoking gun” sequence also appears Mycobacterium smegmatis, for example [2].
[1]: https://www.nature.com/articles/d41586-021-02596-2 [2]: https://twitter.com/soychicka/status/1243547603746410500
Random things occur at random.
Let's try a more sensible approach. Let's see how often the sequence occurs in other databases. The author's 1/20^6 (1/64,000,000) probability calculation omits the fact that every database search considers many millions of possible alignments (a crude calculation would be 6 * the database length, or 6 * 1.6E7/20^6 ~ 10000E6/64E6 = 156 times by chance), and we're shown that in a 5,000,000 protein/1.6E7 residue database, 100% matches are seen several dozens of times, as expected.
How about other protein sets? Search the human protein set (19,000 proteins), there are no 6 amino acid 100% identical matches. The best matches are 100% of 5 residues, or 83% identity over 6 (with an E()-value of 22).
How many proteins do we need to search to find an exact full-length match? (we have dozens of matches in a 5,000,000 viral protein database. and none in 19,000). If you search all the NCBI landmark sequences (~500,000, non-redundant), there are 3 identical matches (E()-value: ~44) in Drosophila, Corn, and Strep. pneumoniae. So this sequence is not very special -- as the E()-value statistics indicate, it happens all the time by chance. (If it happens 3 times in a 500,000 sequence non-redundant database by chance, we would expect it to happen 30-100's of times in a 5,000,000 redundant viral database by chance, which is what we see.)
Six amino acid non-significant matches are evidence that properly done statistics are accurate, and pretty much nothing else.
- Humans and their pathogens
- Economically important organisms (e.g. crops and farm animals) and their pathogens
- Model organisms
For example, compare the number of Genebank sequences for HIV-1 [1] versus Feline immunodeficiency virus [2]. HIV-1 is additionally problematic because it has an extremely high mutation rate [3], making it more likely for a random sequence to match some sequenced HIV-1 genome.
[1]: https://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi?mod...
[2]: https://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi?mod...
So unless you go to a lot of trouble (e.g. the NCBI landmark database), most sequence databases are a mixture of 10's to 10,000's of very closely related sequences (not a random sample), from a non-random sample of organisms (we tend to sequence pathogens).
Protein and DNA sequence databases are very large, but they are not random.
BLAST does not contain all genetic material from all viruses or organisms in all of the natural world. There are countless (really, impossible to count) viruses of all kinds out there dancing around in the bodies of all kinds of creatures. There are literally hundreds of trillions of viruses in your body right now (not different kinds but individual viruses). Each of those hundreds of trillions of replications was an opportunity to mutate. Now multiply that by all of the individual animals human or otherwise who could harbor coronaviruses and you have many hundred billions of trillions of chances to generate the gene sequences in question. If I could get billions of trillions of lottery tickets, I'd win the lottery, and here it looks like SARS-CoV-2 won the "does a gene sequence in the virus exist in a patent by moderna" lottery.
And the point isn't that a virus mutates on every replication, but that every replication is a chance for a mutation.
The point is everything here hinges on an argument from probability.
There are more viruses on earth than there are stars in the universe, the astronomical figures are meet with each other. We're not talking different orders of magnitude between the opportunities for a certain mutation and the probability of a certain mutation occuring..
edit: Further, it looks like the sequence exists in some bacteria, meaning it isn't even a question of a random mutation but possibily the incorporation of the sequence from an existing source (which the author claims doesn't exist, but it does, in mycobacterium smegmatis...). this destroys the entire premise of the claim.
And even if they survive the selection pressure is immense.
That kind of worried me. Because even if you think sars-cov-2 is completely natural (and I have no problem accepting that position), have you thought about how many groups around would still have the capacity to make such a virus, if they wanted to?
Even just the fact that "some people believe the virus was made in a lab" can be dangerous. Because what if, say, Kazakhstan's dictator-emeritus Nursultan Nazarbayev believed it? How far-fetched is it that he, or someone like him even in the more nominally democratic parts of the world, would decide, "they fired the first shot, we can't allow there to be a bioweapon gap!" and commandeer the institute to do "dual use" research?
It would be easy to excuse too. They could say, "we're just designing these powerful virus variants to have a vaccine ready in case someone else comes up with the same powerful virus variant", and it would be a reasonable argument, at least as arms race escalation arguments go.
Dismissing everything as far I'd problematic. If you can't prove it and several scientists were baffled by the number of changes and we still donate an answer, you can't criticize someone from believing A while you believe B.
This is *published research*. Specifically observations 3 and 4 https://jvi.asm.org/content/jvi/82/4/1899.full.pdf
And so, fast vaccine production must be a priority. The development of drugs to treat viral infections and their manufacture must also be a priority. The ability to detect carriers and quickly shut down travel must also be a priority. And how do we enforce quarantine in populations that are resistant to it? All of these concerns are, I am certain, at the forefront in the mind of any Western government to be certain, but likely any government anywhere.
Because regardless of IF Covid-19 was created in a lab, it COULD HAVE BEEN created in a lab. And we all got caught with our pants down.
I propose that what we really see is that humans have a very hard time without someone to blame (hence victim blaming, hence religion and a whole bunch of other coping mechanisms).
I ain't a geneticist, so it goes without saying that I don't really know what I'm doing and I'm probably using BLAST wrong, but it seems to me like (EDIT: something very close to) CTCCTCGGCGGGCACGTAG is pretty dang common even within the human genome (let alone other species), and it doesn't seem far-fetched to me that SARS-CoV-2 might've yanked (EDIT: something very close to) that sequence from one of its hosts at some point.
EDIT: I forgot to check the completeness of the matches, and it does look like the hits tend to have one or two nucleotides that don't match. Still, it's close enough that the article doesn't seem like it presents much of a smoking gun.
¹: https://blast.ncbi.nlm.nih.gov/Blast.cgi?CMD=Get&RID=Y3TZW9W...
²: https://blast.ncbi.nlm.nih.gov/Blast.cgi?CMD=Get&RID=Y3UU86P...
After accounting for that, and after realizing that BLAST only returns the first 100 results, I searched on the "nt" database (where the article found "only" SARS-CoV-2 results) while excluding SARS-CoV-2 (taxid:2697049), "synthetic construct" (taxid:32630), and "other sequences" (taxid:28384). This¹ produced not just a lot of hits, but a lot of exact (i.e. 19/19) hits (mostly various microbes) - further suggesting that even the exact sequence in question ain't unique.
I also ran a search on "nt" strictly on SARS-CoV (i.e. the original SARS virus)², and while no single result completely matches the query, there are a lot of results that could be concatenated/combined to produce the query - and while I'm (again) no geneticist, it doesn't seem to me like some sort of impossibility for that concatenation to happen through mutations in the wild.
¹: https://blast.ncbi.nlm.nih.gov/Blast.cgi?CMD=Get&RID=Y3WR4BP...
²: https://blast.ncbi.nlm.nih.gov/Blast.cgi?CMD=Get&RID=Y3YCNK2...
Well this is a rough start. “This information is being suppressed, which you can tell by the fact that these linked articles are still available.” The utter lack of critical thinking doesn’t make me expect much from the rest of this.
I read this before all the lab leak theories came out, and I'm glad I did because it makes them all seem ridiculous.
https://www.scientificamerican.com/article/how-chinas-bat-wo...
"Also, Dr. Shi was trained by Ralph S. Baric of the University of North Carolina in building “chimera” viruses — taking, for example, the spike protein from a new virus and splicing it to the backbone of a known one like SARS. He invented “no-see-um” techniques that left no trace of the splice."
source https://donaldgmcneiljr1954.medium.com/how-i-learned-to-stop...
https://www.researchgate.net/publication/8119695_Development...
The spike protein was already a defining feature of SARS-CoV when it emerged and became "known" [1]. That's why it was classified as a coronavirus! [2]
Edit: Appears I have misread, see reply
The statement was about grafting a new spike in place of an old one. Where is it implied that SARS doesnt have a spike?
I dont even think that is the important takeaway. The no-see-um editing technique can apply to more than just spike proteins. I dont think its being implied that thats what happened.
The rest of the interview is how to move form using a hash table for counting k-mer frequencies to a vector using a minimal perfect hash, and compute k-mer frequencies quickly using a rolling hash. It's a great question because it always throws the CS people off to hear their questions asked with biological structures involved.
if somebody was going to go straight to proposing a probabilistic data structure like a bloom count and could code up a simple example in 20-30 minutes, that is also OK (and would lead almost certainly to an immediate hire).
The field of string algorithms, or "stringology" as they whimsically call it, is also heavily focused on DNA.
I was planning to study bioinformatics in my youth. Life events intervened and I had to do a regular software engineering degree instead, but I still have an unreasonable love for the Burrows-Wheeler transform.
I think the way I'd put it is, DNA sequence analysis benefitted greatly from the wide range of algorithms creatd for string processing- everything from fast string search (boyer moore) to dynamic programming to forward-backward alogirhtm for HMMs. But there's very little, if at all, stringology done that's purely DNA specific.
But it's a good example of the thing I say: Even for pure math papers related to BWT, odds are good they use DNA as an example. I wouldn't be "thrown for a loop" at all from your clever bio-themed interview questions, and there are a lot of people like me. Maybe not who love BWT as much, but who in exchange have actually studied bioinformatics.
I'm not sure we're better candidates for it, so you're bringing in some questionable bias.
But since you're a BWT nerd, you'll enjoy my story. I used to work at Google and was well-placed at a key time, sitting near Jeff Dean and Sanjay Ghemawat while launching Google Cloud Genomics. One day, I invited Richard Durbin (one of the innovators in using BW for genomics) to stop by and give Jeff and Sanjay a quick update on his work in sequence compression and indexing. it was great to be a fly on that wall.
At the end of the talk I mentioned that Burrows of Burrows-Wheeler was just down the hall if Richard wanted to chat with him as well. Unfortunately Mike wasn't in the office that day, otherwise I would have been able to cross "Be present when the inventor of an important and subtle algorithm meets a user who has applied it for state of the art research in an unexpectedly productive way" off my bucket list.
As for the bias, sure, but it's fairly small, and I can adjust on the fly to give every candidate a reasonable attempt to solve the problem. I can also give the problem without any reference to biology
Well that's just entirely false.
https://www.sciencedirect.com/science/article/pii/S187350612...
I'd draw your attention to section 2.3:
> 2.3. Furin cleavage sites are common in Betacoronavirus
Oh, also, if you click on Prashant Pradhan's paper it has been withdrawn:
https://www.biorxiv.org/content/10.1101/2020.01.30.927871v2
> Abstract
> This paper has been withdrawn by its authors. They intend to revise it in response to comments received from the research community on their technical approach and their interpretation of the results. If you have any questions, please contact the corresponding author.
1. Why do we care so much about a furin cleavage site? The author makes a big deal out of this and I'm not sure why. It seems like he specifically looked at the gene sequence at a furin cleavage site and then said: hey look, this occurs at a furin cleavage site! It must mean something! (Thanks to etaoins for finding at least some evidence that furin cleavage sites are not, in fact, unusual in Coronaviruses, although I have not read the paper: https://www.sciencedirect.com/science/article/pii/S187350612...)
2. What does HIV-1 have to do with anything? That is, if these sequences all matched virus XYZ instead, would that also be some sort of smoking gun? The top 3 sequences all seem to be very common among lots of viruses (rotavirus, measles, coronaviruses, and HIV all show up in the top 3 lists). So he seems to have found a single sequence that lots and lots of viruses have (at least variants of), and that one sequence is the first 3 rows of his table. Am I missing something?
3. If we ignore every result entered after Feb 2020, aren't we throwing out a lot of sequences that were entered specifically because people were suddenly all very interested in Coronaviruses? If he's making claims like "these don't appear in any other Coronavirus", the "except those we found after everyone started looking at Coronaviruses much more" somewhat detracts from that, doesn't it?
It feels a little bit like finding a passage "He froze, trembling with fear" and saying: the English language has 1 million words, so the probability of finding these 5 words in order is (1 million)^5. Therefore if we find it in two novels, one must have plagiarized the other. But not really though, because "he" and "with" are extremely common words, and "froze", "trembling", and "fear" are very highly correlated to each other. I think it's pretty clear his first 3 examples are all highly correlated.
His last example (and the title of this post) seem a little harder to explain to me as a layperson, but I'd note that for a 30-nucleotide sequence, there are 3^30 other sequences that are only a single mutation away. We again see Rotavirus on his list. How far away is that Rotavirus sequence, and do we think it prohibitively unlikely this sequence came from Rotavirus?
let's just say some number of scientists love jargon and they love to bandy about "furin cleavage site" and "smoking gun" because it sounds cool.
For your second question, HIV-1 doesn't have anything to do with this.
But then why is 40% of his article and 3/4 of the rows in the table devoted to the similarity between SARS-CoV-2 and HIV-1?
1. We care about a furin cleavage site, because it sticks out like a sore thumb. It doesn't exist in the closest family of coronaviruses to SARS-COV-19, yet exists in distant relatives. This is the exact sort of thing you fuck around with if you are doing GOF (pathogenic or otherwise) research -- you know it's likely to be tolerated by the parent scaffold because relatives feature it and aren't completely unstable -- and you see the vague outlines of a correlation with function.
Furin cleavage sites are also something a bioengineers like to engineer because we are familiar with them (we often create furin cleavage sites for protein production -- to liberate the purified protein from whatever scaffolding we use to build the protein -- or for analytical purposes, because you can get good evidence that you built what you built if it gets chopped in two).
Finally, we know that a nonprofit associated with the Wuhan Institute of Virology proposed specifically to engineer furin cleavage sites into coronaviruses (in association with WIV, but with the research done in the US), which proposal was rejected by DARPA.
2. HIV doesn't have anything to do with anything except that this is what the protein sequence of the spike matches up with. This is ~6 amino acids, which is admittedly not a whole lot, but not all furin cleavage sites look alike. If I am reading this blogpost correctly, the HIV-1 DNA sequence (which can represent the same protein sequence with worst case ~60% divergence) doesn't match up with the coronavirus, but the moderna patent sequence does (and it's exact). The moderna patent furin cleavage sequence known to be based on the HIV-1 sequence.
3. You ignore those results because NCBI will send you a hell of a lot of COV-19 sequences, because guess what, we've been sequencing the shit out of COV-19. But your criticism of the "statistics" that come out is 100% correct. It's possible that in the last two years have found a "missing link" that will be filtered out by those criteria, though one suspects that if such a thing were found, it would have made some pretty big headlines.
So, if we do a dispassionate analysis without considering faulty application statistics to the pre-2020-only sequence counts, and simply sticking to the existence or inexistence of sequences as per the database we are left with a few possibilities:
1.
COV-19 precursor protein intermixed with HIV-1 and did an uptake of the DNA (this is totally possible but rather rare, even given that HIV-1 is rampant in the wuhan-ish part of the world), then somehow converged on the moderna sequence (let's say 1 in 250 odds).
2.
Some (prior-to-2020 unknown or unregistered) DNA sequence that happens to exactly match up the moderna sequence, by chance, got merged with some wild Coronavirus sequence to yield covid-19.
3.
- someone came up with the idea of fusing furin into coronaviruses (true if you believe the leaked DARPA grant application, btw I have read the whole thing and the level of detail and its similitude to grants I have read and written would have been really hard, but not impossible, to fake, especially by a state actor.)
- someone went ahead and did the research on a slightly different coronavirus than proposed in the DARPA leak even without grant money (this does happen in biology, I'm probably on the tail end of the distribution but out of 10 years in science only 1 was working on previously proposed research, the other 9 were on research I was gathering preliminaries for to write grants)
- the postdoc or grad student assigned to this went and collected a bunch of furin cleavage sites, probably off of NCBI blast, and stack-overflow-copy-pasted them into the coronavirus, of which one was the site from the moderna patent.
- the moderna sequence randomly happened to be the "winner" and spread around.
4.
Moderna engineered covid-19. (I don't believe this, but I included it for completeness, because I don't have a good intuition of why I don't believe this)
I don't know what the statistics on 1 and 2 are, but my gut feeling says they are rare. But rare things happen, and we cannot discount that, and there is a bias to see things that look good. We can't say "it's too good of a coincidence that the furin cleavage site happened to be in the right spot" -- because those events are selected for by the results. For 3, it's bullets 3 and 4 we don't have particularly good priors on, but I think it is extremely irresponsible to claim "it could never have happened", because they are things that I would have done if I were given the directive to insert furin sites into coronavirus. My gut feeling, thus, is that 3 is the most likely, due to projection: I assume WIV researchers are as smart (and as dumb) as I was.
I've noticed that you are responding to many people that blatantly ridicule the original article and the author for writing behind a pseudonym. They are just trying to manipulate uninformed people towards ignoring this information.
The key to all this is that Coronaviruses have been studied for decades and there is no single coronavirus with that furin cleavage and HIV matches on the bindings.
After testing thousands of bats, no one has found an animal with COV-19 and, even worse, it is not easy to make bats to get infected by COV-19.
The fact that this short strings are in patents by one of the main manufacturers of the vaccines and the other ones are part of HIV which are directly connected to the Early Life of Fauci might be coincidental, but there is nothing that indicates that this is a natural virus.
Hold on. Let's be precise. Other coronaviruses do have furin cleavage sites. No other coronaviruses have this particular furin cleavage site.
> but there is nothing that indicates that this is a natural virus.
Let's also be clear here. You don't have to have anything that indicates that this is a natural virus; there is no reason why (as some of the other rabid commenters on the "other" side keep saying) this couldn't be natural. It would take some (very possible) stretches to get there. Biology does stretch sometime, there is no occam's razor, consider the platypus.
In the end I think one has to look at the proposed chain of events. My belief and let me be clear that this is a belief is that it's a lab leak. From when this started, I have half-jokingly said that this is how I would have mishandled the virus (I'm pretty clutzy in the lab - imploded no less than two dewars), and over time evidence has trickled in making the lab leak hypothesis stronger [0], while the observations you would expect have come out that makes the 'natural' hypothesis stronger have not surfaced, so in my mind, the trendline has been towards lab leak. Over time, the 'missing chain of events' that would have to have happened for it to be a lab leak are closing in. But sure, we could find a credible missing link for the natural hypothesis tomorrow and I would have to revise my beliefs. Honestly, the way the CPC has locked down information about what happened, we will almost certainly never know for sure and either way it will likely have to remain a belief.
[0] e.g. the leaked DARPA grant, information that the French refused to certify WIV lab, genetics on lao and yunnan viruses, similar lab leak happening in Taiwan, etc.
Nature favors sequences that work. Try a billion variants, keep the one that actually helps. We don't see the 999,999,999 failures.
That covid was made in a lab? Or did covid and moderna independently reach the same "conclusion" through parallel evolutionary pressures (in nature and in research)?
>this sequence
>this
>sequence
Furin clevage sites haven’t been found elsewhere in SARS-Cov-2’s Sarbecovirus subgenus.
You have two strings, A and B, and a list of strings L. Find the longest substring that appears in both A and B but not in any element of L.
Mutation? The same way we get every new pathogen?
(eg https://i1.wp.com/www.ghgossip.com/wp-content/uploads/2020/0...)
Then you'd have Substack's talking about how the virus could have been naturally occurring, that there's no reason it had to be made in a lab, that it's just the government blaming China, etc.
I wish there was a word for beliefs that supposedly make use of evidence, but can easily use the exact same evidence to argue for the polar opposite belief.
>Posted by 'version_five'
"This sounds like one of those "bible code" things. As someone who knows nothing about DNA sequences, what is the probability that two codes could be matched? (In a birthday paradox sense where you can go looking anywhere)"
>Posted by 'whoomp12342'
"isn't it well known that one of the main reasons why we got a vaccine so quickly was because we had done a ton of research when SARS originally broke out?"
An unrelated threatening comment would have gotten the unrelated comment killed, but not the main posting.
If the main posting was flagged, that is because some number of users with enough karma to flag clicked the "flag" link on the post.
The genome in question is about 30000 bases long, which isn't nearly long enough for the outcome space size issue to come into play. Assuming independence, the probability is still very remote, effectively zero.
However, there is in no way any expectation of independence. SARS-CoV-2 is obviously closely related to SARS-CoV-1, which is obviously closely related to 2015/2016 ongoing vaccine research. (And perhaps gain of function research.)