1. Free rlhf
2. They cookie the hell out of you to breadcrumb your journey around the web.
They don't need you to login to get what they need, much like Google
They don't need you to login to get what they need, much like Google
Of course numbers are pretty random, but it's just to give an idea of how these things scale. This is my experience from my company's own internal -deep learning but not LLM- models to train which we had to buy data instead of collecting it. If you can't tap into data "from the wild" -in our case, for legal reason- you can still get enough data (if measured in GB), but it's depressingly more repetitive, and that's not quite the same thing when you want to generalize.
Modern captchas are self driving object labelers; you just need a few to "agree" to know what the right answer is.