Imagine you have a string of length 3 billion made by randomly choosing from 4 characters. Like this
dna = ''.join(random.choices('atgc', weights=[30.9, 29.4, 19.9, 19.8], k=3_234_830_000))
you get to randomly sample 1 billion[3, page 7] overlapping substrings of length 200[3, page 7] with .1% of the characters randomly changed[3, page 8]. Trying to find the original string from this is technically an undecidable problem. If there's a sequence 400 characters long that repeats multiple times, how could you know if it repeats 5 times or 50 times? (this would be unlikely to happen with random.choices() but DNA isn't random). This is called sequence alignment and it's one of the hard problems in bioinformatics[4].[0] https://docs.python.org/3/library/random.html random.choices() was added in Python 3.6
[1] http://www.biology-pages.info/B/BasePairing.html source for `weights`
[2] https://en.wikipedia.org/wiki/Human_genome source for 3_234_830_000 (python ignores underscores in numbers)
[3] http://sci-hub.io/10.1111/j.1755-0998.2011.03024.x illumina is the most popular producer of genome sequencers