Biochemical Pathway Maps
biochemical-pathways.com
biochemical-pathways.com
One thing I've observed is there seems to be no universal well-annotated structured database for biosynthetic pathways. Updates and additions to known pathways are published unstructured in papers, often in graphical form, presenting even more challenging for structuring the data.
Were this to exist, it would be possible to build an app that let you easily design DNA plasmids for testing all kinds of incredible biosynthesis experiments. For instance, it would be able to select an organism [E. coli], enter your starting substrate [glucose], select a desired metabolite [psilocybin] and out would come a DNA sequence for a plasmid that contains are the necessary parts, promoters, genes and terminators for transforming E. coli to produce psilocybin.
In my opinion, such a tool could profoundly impact science, industry and human well-being.
Have you tried applying your hobby towards reducing the cost of healthcare in America?
Could you somehow make monoclonal antibodies in your petri dishes?
Could you diagnose what kind of flu virus you have?
Because the phrase "self-learning synthetic biology" immediately rang a bell inside my mind that screamed "small business dedicated to cheaper healthcare", and is something that I have considered dedicating my free time towards.
That isn't my stated goal, but I would imagine that sharing how to genetically engineer an incredibly powerful organism to express heterologous genes for less than $3000 in a kitchen lab could open up many possibilities!
> Could you diagnose what kind of flu virus you have?
Is it possible to extract DNA from a bodily fluid, perform PCR and read the sequence results to determine flu variant? I'd have to check on if there is a known DNA sequence that barcodes flu viral variants, but presuming there is, then yes - all of this is possible from a home lab currently.
This isn't so much a technological problem, but more a politics/incentives issue, as evidenced by the comparably lower cost of healthcare in Europe. Startups might be able to make small reductions in the cost, but truly solving the issue requires re-organising the entire healthcare system.
> Could you somehow make monoclonal antibodies in your petri dishes?
This wouldn't be particularly useful. In order to use it in humans, you need a ton of quality control and regulatory approvals, and this is where a lot of the manufacturing cost crops up.
- the academic publishing model incentivizes groups to build their own tools
- the small-ish market for something like this has kept commercial software from taking off (Genomatica started by building a tool like this in the early 2000s, before pivoting to bioprocess development)
- it's really hard to specify pathways in concrete physical terms. Even a chemical like glucose is actually a collection of pseudo-isomers (alpha & beta D-glucose). And try firmly defining a "gene" in your database!
that being said, there's a ton of work going on in this field and many cool projects to follow: https://dd-decaf.eu, https://biocyc.org, http://bigg.ucsd.edu, http://metanetx.org
many papers mention KEGG and use its references and strucfture as resources. https://molecular-cancer.biomedcentral.com/articles/10.1186/... (I'm not saying there's anything special about this paper in particular)
Reactome: https://reactome.org/
KEGG: https://www.kegg.jp/kegg/pathway.html
Wikipathways: https://www.wikipathways.org/index.php/WikiPathways
BioCarta: https://www.hsls.pitt.edu/obrc/index.php?page=URL1151008585
SMPDB: https://smpdb.ca/
ConsensusPathDB: http://cpdb.molgen.mpg.de/
Regarding the issue of of parsing pathway figure data from published papers: pathway-figure-ocr (https://github.com/hiplot/pathway-figure-ocr/) is a project by the developers of Wikipathways which is trying to solve that issue.
WikiPathways supports advanced queries via their SPARQL API and UI. See [1] and [2]. I find WikiPathways nice because it lets logged-in users create and edit pathways, with a low barrier to entry.
I've been building a way to find related genes using biochemical pathways [3]. The source code linked there includes practical examples for fetching information on genes in those pathways, which you rightly note is needed for something compelling. That and other code there might help spark ideas for you on how to glue together various biochemistry and molecular biology APIs to achieve your vision.
I'm currently working on a way to drastically expand the set of organisms and pathways covered by WikiPathways. Yeast has 66 pathways there, compared to 1319 for human. By doing fast ortholog detection at runtime (using another SPARQL API, provided by OrthoDB [4]) I'm hoping to be able to convert relevant annotated pathways across organisms, e.g. human to yeast, mouse to rat, Arabidopsis to rice -- and vice versa.
[1] http://sparql.wikipathways.org
[2] https://www.wikipathways.org/index.php/Help:WikiPathways_Spa...
[3] https://eweitz.github.io/ideogram/related-genes?q=RAD51&org=...
Automating the process is a valuable effort nevertheless. I think because of it's complexity, it would require a combination ofknowledge-graph based reasoing, with AI.
There are databases such as Addgene: https://www.addgene.org/search/catalog/plasmids/?q=Cas9, in which you can search existing plasmids created by the scientific community. I think this could be a valuable resource for a machine learning approach.
(1) https://news.ycombinator.com/user?id=ssprang
(3) https://news.ycombinator.com/item?id=6665261
Your idea of a program where you enter starting point and desired product is great in theory, but it’s similar to saying “oh, creating DropBox is easy, it’s just an interface linked to Amazon Cloud”. It’s missing about 99% of the stuff in the background nobody sees.
- there is no documentation to start with
- you can only learn the language through trial and error and development of crude frameworks that are accurate 75% of the time
- you have no ability to look at the environment or operating system beyond “trying things” and making assumptions based on results
- the environment itself resists most attempt at modifying it through feedback mechanisms (but you can’t see these either, just the effects)
- no two computers + environments are identical. Code you write will mostly work in different computers, but occasionally won’t and sometimes will just brick (kill) the computer
- when you learn a rule about changing a variable to produce Y affect, that rule doesn’t apply in 9 out of 10 scenarios
* What your ultimate goals are (proof of concept versus lab/industrial-scale synthesis)
* Your synthetic vehicle, i.e. in which organism you will be recreating this pathway
* Whether the relevant pathway/genes already exist in some form in your vehicle, and may be affected by an introduction of foreign genes
* How you will detect expression of your foreign proteins (the fact that your genes are expressed doesn't mean the corresponding proteins are produced, and if you can't confirm protein expression, the inevitable troubleshooting will be five kinds of hell)
* What kinds of promoters you'll need to drive gene expression
* How stable the proteins in this pathway are, and if you need to modify them to increase stability
* Whether your proteins, after being produced, need extra tweaks (post-translational modifications, in the parlance) in order to work
* How much DNA you'll need
* How you'll deliver the DNA, i.e. what size/kind of delivery vector
* How you'll make the delivery vector
* Whether your proteins will fold properly once produced in the vehicle, and whether they will localize to the right parts of the cell
* Whether your proteins are toxic to your vehicle cell
*etc.
I'm looking for something more practical and hands on at this point. I was planning to start my home late in the next month or so, and start with some basics bacteria projects. So far I found the following to be useful:
1. EDX principals of synthetic biology: https://www.edx.org/course/principles-of-synthetic-biology
2. Coursera systems biology: https://www.coursera.org/specializations/systems-biology
3. Coursera Industrial biotechnology https://www.coursera.org/learn/industrial-biotech
4. SBOL https://sbolstandard.org
5. This GitHub: https://github.com/websemantics/awesome-synthetic-biology
6. https://barricklab.org/twiki/bin/view/Lab
7. Synthetic biology primer (book)
With regards to your tool idea, it is still very expensive to generate DNA from scratch. Designing a theoretical plasmid won't get you all the way there if you can't source the parent DNA sequence. We usually pay ~$0.30/base for DNA synthesis. Plus, I'm not sure what the metabolic pathway of psilocybin is but you may need several plasmids to recreate the whole thing in E. coli (generally limited to ~10-15 KB). An interesting idea but definitely no small feat of engineering.
We recently had a collaboration with AWS Open Data+Neptune to show how easy it is to combine Rhea,ChEBI and UniProt in your own database.
[3] https://academic.oup.com/bioinformatics/article/36/6/1896/56...
[4] https://aws.amazon.com/blogs/industries/exploring-the-unipro...
Showing them a map like this tends to clear things up pretty quickly!
the biologist, who knows better, will regard the chart as “our current best attempt at documentation, subject to change at any time at all”
My background is biomedical. This or a very similar chart was used in the last week of lectures on control systems as a lesson in humility.
It was meant to demonstrate both how complex things can be — and then, when the professor recolored the lines on a section according to the confidence in their correctness as a function of recent research, how uncertain our knowledge of that complexity really is.
He had a backup slide where an entirely new graph, with almost no topological similarity to the metabolic one, was created for a few of the nodes that participate in other nonmetabolic pathways.
I see you've not done much legacy software maintenance.. ;)
I'm a programmer and I do understand that these maps are "as much as we currently know" and the "we currently know" is expanding rapidly. I like to see biology as reverse engineering alien technology made by a much more advanced civilization. Of course that's not true — it all evolved over billions of years — but the resemblance is totally there.
The one thing that really strikes me about biology is how young of a science it actually is and how much progress we've made in so little time. It was only 1953 when the structure and function of DNA was discovered. In terms of history, it's not even yesterday. Yet now, 68 years later, we're literally injecting people with correctly encoded instructions to make a harmless part of a deadly virus to save their lives. This kinda blows my mind.
Be really careful here, though, otherwise you'll hoodwink yourself.
The two famous things that come to mind for me are these, both of which remind me to tread carefully:
https://blogs.sciencemag.org/pipeline/archives/2007/11/06/an...
https://www.cs.utexas.edu/users/EWD/transcriptions/EWD09xx/E...
These metabolic charts, like the results of spectroscopy, actually represent some of the best, highest quality scientific reference data that exists.
But yeah, subsystems don't need to be kept to a size that will fit inside a regular programmer's working memory.
In many cases you can trace pathways upstream until arriving at a supplement or nutrient that can be bought online. People erroneously assume that consuming that nutrient will achieve some downstream effects in certain pathways because all of the arrows connect. It's not that simple, though. For the most part our bodies are excellent at absorbing the nutrients we need and regulating everything as appropriate.
This one can be ordered as well.
0. https://www.roche.com/sustainability/philanthropy/science_ed...
> We believe all people should have access to important scientific information to help their research and education. [...] Roche provides copies as a free service to the public.
they are pretty cool to have, with very high quality two colour printing, and you also get a wee booklet with an index to the two posters to help you find things. there's also a book [0] by the same author with much more detail, and obviously much more portable, but sadly much more expensive ;(
so, i'd try asking for them again, if i was you.
I wonder what formalisms/representations are there to manage the complexity of metabolic pathways. For example, say that that figure is 100% accurate, and furthermore, "all metadata needed" in terms of reaction rates, etc. for each individual reaction is available. If one wants to stage an intervention (say, suppress the production of compound X without altering anything else significantly), what kind of program would be able to find a solution?
I have been trying to find an image of this poster without success. The detail level was like on these Roche posters, but the style was more like in these:
MIT 5.111 Principles of Chemical Science, Fall 2014 https://www.youtube.com/playlist?list=PLUl4u3cNGP63LOmB3_O0x...
Textbooks (free PDFs available on gen.lib.rus.ec):
Lehninger, Principles of Biochemistry https://www.macmillanlearning.com/college/ca/product/Lehning...
Molecular Biology of the Cell https://www.ncbi.nlm.nih.gov/books/NBK21054/
Glucose, Acetyl-CoA, Pyruvate, ADP, ATP, NAD+, NADH, NADP, NADPH
Protein Carbohydrates Fat
Your goal is to take reduced carbon (carbon bound to hydrogen) and metabolize it to oxidized carbon (carbon bound to oxygen). That process releases energy your cells can use.
The problem with oxidizing carbon is that it creates toxic products ("free" electrons) and involves splitting atmospheric oxygen (O2) into O-, which is very reactive. So then you need a whole host of other pathways to contain that reactive compound.
If you really want to understand it, follow the carbon as it comes in, and trace it through to CO2. That's the pathway, and everything else is to support that.
[1] https://www.elsevier.com/solutions/pathway-studio-biological...
EDIT: and see hint below for pdf: https://news.ycombinator.com/item?id=27545148
Just an important detail. It's not just randomness. It's randomness + natural selection.
Keep in mind that as soon as an element is added to the system - and it works - that element becomes sticky. Rinse and repeat, and you get more and more complex systems over time.