Isn't that just for data collection though? I imagine the fully trained model they're expecting to have will be able to map thoughts -> text. The data collection will naturally need to have the 2 sets in order to do self-supervised training.
And, my point was, there may be zero words to collect at all, because they'll tend not to be actively thinking about any words.