Call for a Public Open Database of All Chemical Reactions
pubs.acs.org
pubs.acs.org
> This conservative state of affairs in the chemistry community is unlikely to result from some intrinsic properties of chemistry as a science. Rather, it is likely to be the product of complex historical and sociological factors, which can be traced back at least to the Middle Ages and the secretive research projects of the alchemists in their quest for a recipe for converting vulgar metals to gold.
Yes, I blame Hermes Trismegistus, Abū Mūsā Jābir ibn Ḥayyān, and Isaac Newton :)
Unfortunately published data so far is highly questionable. I can totally confirm this quote from wikipedia:
"A 2016 survey by Nature on 1,576 researchers [..] on reproducibility found that more than 70% of researchers have tried and failed to reproduce another scientist's experiment results (including 87% of chemists, 77% of biologists, 69% of physicists and engineers [...]" https://en.wikipedia.org/wiki/Replication_crisis
You’d have one postdoc do the pipetting and the experiment worked. Another person tried, following the same procedures with the same equipment in the same room at the same time of day in the same environment … experiment failed.
My takeaway was that a lot of this bio stuff is inherently flakey and/or poorly understood. We kinda mostly know what’s going but it’s really lots of trial and error and chasing statistical variations/effects. Best you can do is “it works N% of the time”
You can have a system where 100% of researchers have failed to replicate another experimental result while 99.999% of results are replicable.
That's not to suggest that's what's actually happening here, because the replication crisis is well discussed, but it's not a good statistic to measure things with.
I absolutely agree that enumerating reactions would suffer from combinatorial explosion (as opposed to chemical explosions, hoho). However there have been some efforts in what I know find is called 'computational retrosynthesis' (I think):
https://www.chemistryworld.com/features/computer-guided-retr...
Of course, any 'prediction' of a reaction relies on some model to give you an idea of how feasible it is to do in the lab, so I'm less worried about . There's the classic Derek Lowe post (that I've seen on here before, but still) about 'FOOF':
https://www.science.org/content/blog-post/things-i-won-t-wor...
[1] - https://reactionmechanismgenerator.github.io/RMG-Py/index.ht...
RMG predicts likely reactions based on chemical kinetics, which might also make a useful database itself - or link to the experimental one, of course.
Of course, our entire global economic system is heavily weighted against doing something so useful in public, for the public good...
Please email me at surprisetalk@gmail.com if you (or anybody you know) would be interested in collaborating.
"Finally, one must address the concerns of various commercial stakeholders."
"A few chemical databases of reactions do exist, but these are commercial (e.g., Beilstein/Reaxys [Elsevier], SPRESI [InfoChem], and CAS [ACS]), and even when a license is purchased, the underlying data are accessible only through a narrow, one-query-at-a-time interface, completely stifling the application of powerful artificial intelligence and machine [...]"
This seems like the first point of approach, not the last. Can anyone comment why these data gatekeepers have not made a business model to give programmatic access to nerds?
And if they have financial reasons for not doing so, how will governments sweet talk shareholders into supposedly losing money?