ParentFull threadsudoscript·It's a dataset of questions marked as duplicate, so it could be used to train on semantic understanding. There are 400,000 pairs -- that's not huge, but it's not bad either.View on HN