Reverse Engineering Source Code of the Biontech Pfizer Vaccine: Part 2
berthub.eu
berthub.eu
I wonder if the authors have also steganographically encoded something in there...
ViennaRNA is also available as a standalone package[1] if you prefer to run the secondary structure predictions locally.
You still need living things for the runtime, but hasn't AlphaFold basically solved that problem too?
Programming in DNA is like programming in assembly language, but a 7.5 KB assembly language program is well within the reach of a lot of people. Has anyone tried to write a 7.5 KB living thing or DNA-based tool from scratch? It doesn't necessarily even need to reproduce to be revolutionary, tiny genetically engineered transistors or structural fibers or chemical reactors might all be super useful and have super simple DNA programs that only have a few dozen lines of code. What makes this so hard?
Systems like AlphaFold2 only solve a tiny tiny part of the problem (i.e. the question "given this amino Acid sequence: how does this Protein look like in 3d"), but it doesn't adress any of the other problems (like: "how does X structure behave in Y environment", "how are the interactions between these N Protein structures", "How do Proteins form complexes?", "how do proteins interact with RNA?", "How do you determine the location of Amino Acid side-chains?" and, most importantly, "what does X do?" )
I liken this to the problems we have in automatic theorem proving: We have formal logics/Type theories that allow for automatic theorem proving, but we still need computer scientists or mathmaticians even for sometimes very trivially complex tasks, because the scale of the problems is way to big to handle automatically. My AI-professor framed it like this "If you have a problem that would take you 10 minutes to solve, a theorem prover will solve it in 50ms. If you have a problem that takes you 20 minutes, the theorem prover won't solve it". We are at a similar, if not worse position, in Biology: We don't even know the system we're working in, yet.
This paper gives background on designing vaccines to maximize rapid antibody responses post-infection, ideally before the virus enters cells:
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6936610/
-
This paper talks about how this is achieved specifically for SARS-CoV-2:
https://www.biorxiv.org/content/10.1101/2020.06.15.152835v1....
To maximize the effectiveness of the vaccine, its spike protein should stay in the prefusion state no matter what. The two mutations in the engineered spike proteins allow exactly that.
The most important mutation in the new VOC 202012/01 variant (N501T) translates to tighter binding to the ACE2 receptor, which should have no effect on its structure. However, that variant has also a double deletion (residues 69 and 70 are removed) that might induce slight conformational changes.
Another thing to keep in mind is that vaccines are polyclonal: the antibodies produced by them target several different regions of the spike protein. This implies that the virus would need to mutate several times in separate parts to be able to evade immunity effectively.
https://www.addgene.org/145032/
(seriously though, don't try this at home)
Words can't express how much more difficult it is to work with mRNA than DNA.
Pigs are hardy, but humans are too.
Besides that, the problem is overstating how much of an issue this actually is. For a medical product you need to cater to the long tail of humanity that might have a very adverse reaction to endotoxins. But most people, very nearly all of them, do not have such a reaction. People prepare their own IV drugs all the time. What kills them isn't a reaction to endotoxins (if there is any to react to in the first place, as it is quite easy to make chemically pure drugs) but side effects of the drug or infections from reusing injection equipment.
Mammals have been getting poked with sticks for millions of years.
Basically, Pfizer is responsible for logistics/distribution and Biontech does the specialized research and development. It’s a good partnership.
Also why you’re seeing manufacturing issues from the Moderna side by the way, as they tried to do everything by themselves, with no experience in logistics/distribution.
This is a naive probabilistic method that calculates probabilities of base change according to base position in codon, base, and amino acid of codon. Then it applies the probability to the viral sequence to generate a vaccine sequence.
The sequence generation is not deterministic. It appears to generate ~87% matching vaccine sequence most of the time.
*edit
I just realized the blog post was calculating % matching of codons, not the base-pair sequences. In which case, my method is ~66% matching.