I can't find the original paper, but with an appropriate amount of pseudorandomness to avoid dead ends, this primitive algorithm would generate the occasional sentence that almost made sense and that bore little resemblance to the original data.
Because of the state of computer technology it was a massive effort and a source of general astonishment. I suspect we're now recreating that minimal environment, this time with better ways to curate the data for small size and maximum drama.
Let's remember that a modern GPT isn't far removed from that scheme -- not really.