But Google is probably a close second, with an even bigger archive of pre-LLM text (including the world's largest collection of digitized books), and supposedly a good LLM and lots of smart ML engineers.
But no matter what method google finds, whether it's some complex AI trained on massive amounts of training data or a simple n-gram analysis that ends up being telling, they aren't going to tell us. That would just accelerate the arms-race between LLM spam and LLM detection. If they have something that works (big if) the details will stay hidden, the best we are going to get is some kind of press release aimed at dissuading people from using AI on their websites.