Wit: Wikipedia-Based Image Text Dataset
github.com
github.com
But.. why post it at all if you're not going to give a download link? Arg, such a tease. (This is also technically an "announcement of an announcement" since there's no substance here, aside from the paper at https://arxiv.org/abs/2103.01913.)
(For 47m maybe you're thinking of a cleaned subset from YFCC100M or something that one of them used, but they all use total images an order or two larger than Wit.)