This Chemical Does Not Exist
thischemicaldoesnotexist.com
thischemicaldoesnotexist.com
The model seems to have learnt chemistry pretty well because, as many have already pointed out, most of the molecules generated do actually exist (or are extremely likely accessible if they haven't already been documented). Even the ones with strange bond angles have otherwise perfectly normal number of bonds. The only time where I get molecules that cannot possibly exist are those with overlapping atoms that just defy known physics.
Addendum: it is worth noting that the model might actually have been trained with data that contain bond lengths, or even spatial information if no post-generation geometry optimisation is performed before a molecule is rendered.
I'm not a chemist, I've an undergrad degree in physics/mathematics. Intuitively, your answer sounds right, but I'm not in a position to judge for sure.
https://www.chemistryworld.com/opinion/chemical-space-is-big...
They could have limited SMILES (popular choice) to have a more stable generative space or they might have introduced validity into the loss. I think the coolest part is how the builder got all the rendering to work well!
Or it could also be a rule-based fragment model guaranteed to hit a valid structure, that works too.
After validating the output, it was easy to plot a (2D) skeletal formula. I never got around 3D renders, but I guess a SMILES -> 2D -> 3D pipeline with some molecular mechanics structure energy minimization for the 3D part is cheap to do.
I found the output surprisingly diverse. Model was very good in adding branched lipid tails that kept rambling on forever, though...
There was recently a link, I think on the front page here, to an article about how many chemical compounds there are [1]. Based on that link we're looking at probably trillions to quadrillions of potential structures with atomic weight under 300, which would cover the structures I saw in my few reloads of the page.
Chemistry is wild. For an example close to home, taking table sugar (sucrose, a single type of molecule) and applying heat to caramelize it results in hundreds to thousands of different end products from at least half a dozen qualitatively different classes of chemical reactions.
[1] https://www.chemistryworld.com/opinion/chemical-space-is-big...
If you ask a CS grad "What will happen if you run this program?" they should be able to predict it. If they've gone through nand2tetris they can explain it all the way -- compiler, OS, machine language, ALU / registers / bus, logic gates.
If you ask a chemistry grad "What happens when you apply heat to this molecule?" can they predict it? Can you explain it all the way -- from molecules to atoms to electrons to quantum fields?
If we can't predict "Okay this is what will happen if I mix these two substances together," how do we have a good scientific theory? I guess chemistry says we always end up with the same atoms we started with (unless you start to go nuclear by using energetic particles to modify the nucleus), but can we predict which of the zillions of possible rearrangements will actually happen? We know by experiment that H2SO4 is an acid, and that H2SO4 is a "legal" molecule in a way that HSO3 or H5S7O9 are not. Is there a way to figure this out from first principles? Can you figure out by inspecting the chemical formula that H2SO4 will be an acid if you didn't already know that ahead of time? Can you figure out that H2SO4 will be a "legal" molecule but H5S7O9 will not? Can you look at a reaction and tell whether it will "compile" and what it does, the same way you can look at a program and figure out if it will compile and what it does? If you can't, why not?
And what use is a theory of chemistry that can't make concrete predictions? If you just have a list of known substances and reactions, is that even a theory, or is it just experimental data?
When you look at just a molecule by itself “what happens if you apply heat” is somewhat simple. Covalent bonds just break because the molecule is vibrating too much - think of a covalent bond as a flexible strut, if you put too much pressure on it, it snaps. This can result in the temporary formation of unstable molecules that then recombine. You could predict which particular bonds in a molecule are unstable based on the total structure, angles, electronegativity, polarity, etc.
But of course those small unstable molecules can further breakdown, react with each other, and react with the parent molecule to form new stuff. So basically the parent molecule is part of some huge “power set” of potential molecules all interacting with each other.
Strictly speaking, the answer is "no", because it's not the formula that matters but the shape. And something like C₆H₁₂O₆ doesn't tell you how all of the bonds hook up--are we looking at esters, alcohols, ketones, aldehydes, carboxylic acids? There's several distinct molecules that have that formula, and those distinctions matter for chemistry.
But given the actual structure of the compound? Yeah, we can compute a lot of stuff. That list of words that probably meant nothing to you--that's different kinds of functional groups, and functional groups tend to react in very similar ways when given several compounds. And organic compound is basically all about identifying these groups and the ways in which they react.
> Is there a way to figure this out from first principles?
"First principles" in this case would basically be a large dose of molecular orbital theory, derived from quantum mechanics. And yes, we can develop a good deal of explanations by recourse to molecular orbital--for example, why aromatic and antiaromatic compounds exist, despite the fact they superficially look like the same structure.
> Can you look at a reaction and tell whether it will "compile" and what it does, the same way you can look at a program and figure out if it will compile and what it does? If you can't, why not?
Typically, the difficulty is in figuring out how selective reactions are. If you've got a molecule with a couple different C=C bonds in it, and you're doing an addition reaction across those bonds, predicting how many of those bonds, and which ones specifically, change in the reaction is more of a crapshoot. So it's not foolproof, but it is generally reliable enough at this point that organic synthesis has moved from "here's a Nobel Prize for figuring how to synthesize vitamin B12" to "congratulations on being hired; why don't you synthesize this molecule while we ramp you up on the job."
Either they do something clever to exclude real molecules, my understanding of chemistry is too limited (100% possible), or it's more like "this molecule might not exist"...
https://www.thischemicaldoesnotexist.com/molecule.pdb
There appear to be several repositories of this format. Maybe they just randomly generate until they find one with a hash that doesn't exist? (Though it's not clear to me how much order of the lines in the format matters).
What these structures remind me most of is what you would find in a sour heavy crude oil. In fact, I can guarantee the person who named this website has never looked at high resolution mass spectroscopy analysis (like an FTICR-MS) of any type of petroleum, or they would have named it "this chemical is probably being pumped out of the ground right now".
Some of the most powerful explosives are made by attaching nitro groups to stressed ring or cage structures.
When the nitro group breaks apart, the nitrogen finds another nitrogen from another NO2 and forms N2 gas, which is highly stable due to its triple bond, so this part of the reaction releases a lot of energy and produces a lot of gas, and is very fast since it does not depend on any other sub-reactions. And stuffing lots of just nitrogen (without oxygen) into a molecule is in itself a way to make it very explosive - see azidotetrazolate salts.
In NO2 decomposition, the oxygen then goes on to find carbon to make CO2, and hydrogen to make H2O, releasing more energy and producing more gas. But this first requires breaking down the relatively stable bonds in the hydrocarbon, so it actually consumes energy from the NO2 decompostion, before it releases more energy than it consumed.
Then stoichiometrically you want to ensure that you have enough oxygen for all your carbon and (ideally) hydrogen, or you'll end up producing a lot of "unburnt" stuff which is inefficient. Notice that for each carbon in a linear hydrocarbon chain (-CH2-) you need 1.5 NO2 groups to get a complete reaction into CO2 and H2O. If you only have 1 NO2 group, CO2 will be formed and you will have excess H2 which is not combusted.
Now as you say there are some stressed rings or cages that are hideously sensitive explosives, precisely because the hydrocarbon bonds have also had their stability reduced. But these typically are not practical explosives. For that you want the stuff to be a solid at a wide range of temperatures, you want it to be non-sensitive to friction and impact, and you want it to have a low vapor pressure. These are all details which depend strongly on the internal structure of the molecule.
Also as an aside I believe there's a current trend to generate chemical compounds by creating SMILES strings using BERT which is a cool way to incorporate language and chemistry (An example of a team doing that https://www.cell.com/iscience/fulltext/S2589-0042(21)00237-6)
Putting that rant aside, most of these chemicals are small enough to have rather tame IUPAC names. As an example, these are the IUPAC names of the first five molecules I loaded (disregarding stereochemistry):
N-{1-[(2-fluorophenyl)methyl]-1H-pyrazol-4-yl}-2-[(pyridin-3-yl)oxy]benzamide
N-ethyl-1-{5-[(2-methylcyclopentyl)methoxy]pyridin-2-yl}ethan-1-amine
2-fluoro-N-{4-[3-(hydroxymethyl)azetidine-1-carbonyl]phenyl}benzamide
N-(8-methoxy-2H-[1,3]dioxolo[4,5-c]quinolin-7-yl)-2-methylcyclopropane-1-carboxamide
1-ethyl-N-(2,3,5-trimethylcyclohexyl)piperidin-3-amine
If you want an example of an absolutely ludicrous IUPAC name, I'd suggest the one I put onto https://en.wikipedia.org/wiki/Maitotoxin earlier this month.Or more aptly, what constitutes a single molecule, rather than a complex structure built from a repetitive pattern of almost-equal components?
The bonds. Not a chemist, but if I remember enough chem 101, all molecules are bonded together by either covalent or ionic bonds between the individual atoms. There are names for some of the common "building blocks" of molecules; i.e. a methyl group is a single carbon with 3 hydrogens bound onto it, and an arbitrary atom on the 4th side.
A "complex structure built from a repetitive pattern of almost-equal components" sounds more like a crystal. There are still individual molecules in a crystal, but the molecules are arranged in a precise and repeating pattern.
But why is this interesting? Minus the animation, isn't this something any smart high-schooler could do with pen and paper?
What's most interesting about that is somebody actually bothered to put it together.
Though it did also occasionally spit out a "number" with a letter in it, like "q29199.951301068788".
Our 8 year old does this with pen and paper, and sometimes also with some website that lets you draw molecules (there's a few, forget which one(s) he uses). That said, his molecules aren't always possible (he understands valence but sometimes he makes mistakes with it or just stops caring about it).
This is also something a talented high-schooler could do well before now. The interesting part there is teaching a computer to generate plausibly-human faces, not the act of generating fake people in general. This is the same way.
This seems more interesting because it has practical implications. Generating chemicals that haven't been investigated is the first step towards a programmatic pipeline that can identify and investigate novel chemicals. These seem of particular interest because they're of the "doesn't currently exist" variety, rather than the "cannot exist under the laws of physics" variety.
If by "chemical" you mean "substance (as represented by this molecule)" and if by "exist" you mean "hasn't been made yet," then it might make sense.
"This substance (as represented by this molecule) has not been made yet" doesn't have the same ring to it, though.
Molecules are abstractions. Leaky ones at that.
Whoever can find such an algorithm will put Corey and Woodward out of a job and the entire field of organic chemistry will study your name and life in future.
I hate to burst your bubble, but that is almost certainly not literally true. Any device with those capabilities will be the end result of heck of a lot of engineering (much of it incremental), but whatever scientific discoveries are necessary will be so far removed that they will hardly seem relevant.
Theoretically, the 2016 Nobel for chemistry might prove to be the relevant foundation for the necessary engineering, but time will tell.
Anyway, discoveries like CRISPR, that can be immediately (relatively speaking) deployed as a revolutionary tool, are by far the exception.
As an exercise, feel free to identify the Nobel prizes that were awarded for the invention of xerography (aka photocopying), digital laser printing, and thermoplastic 3D printing (or any other extant additive manufacturing method).
Many of the images seem reasonable. They can have odd asymmetries that may give an unnatural vibe, though most don't seem to have majorly overt issues.
Most of the more overt issues seem to be melding facial wear (like glasses and ear-rings) into skin.
The most overt oddity was a woman with "stuff" splattered on her face.. I'd be curious how/why that'd be something that could be generated..
A lesser oddity was a man who had a mustache that appeared to be shaven on one half, but not the other.
Well explained here: https://www.youtube.com/watch?v=SWoravHhsUU
> Are reposts ok?
> If a story has not had significant attention in the last year or so, a small number of reposts is ok. Otherwise we bury reposts as duplicates.
It's intentional unclear what significant attention means, but the last submission has (3 points | 64 days ago | 3 comments) that is not very significant.