We benchmark on pre-2023 datasets of O(10M) documents not in our training set. Other detectors seem to have between 1-3% false positive rate and ours is around 1 in 10,000 as of our latest model update. We do a lot of active learning + core set selection to keep FPR low and improve recall on larger LLMs. Our white paper with some methodology is here: https://arxiv.org/abs/2402.14873