Be careful with this one, the premise is sound (coming from a DL background), but I skimmed the paper and didn't see them comparing it to existing simple compounds and don't show if their compounds are actually any better.
Am I missing something or did they also not use a held-out validation dataset to assess performance? It seems to me they ran their autoencoder, got some suggested compounds, listed anecdotal evidence about those compounds, and called it a day.