Anyway, I think your point here is interesting and was kind of the idea behind a lot of the "gather lots of data" startups. A lot of those failed in part because the frontier of AI is moving pretty quickly. You need a lot less data to do interesting thing today than you did not that long ago. Because we've thrown more and more data at more and more compute, I think people don't appreciate how much we've truly progressed algorithmically. You need an order of magnitude less data to do the same thing for each "generation" of AI.
That frontier cuts against the ability to build a moat on user-generated data, so long as it's readily available or somewhat replicable. Your competitor is naturally going to have a cheaper time getting into market than you if they wait longer to do so.
However, this definitely does stand if your area truly is obscure (e.g. specific industry), annoying to gather data in (e.g. certain healthcare applications), or actually proprietary (e.g. your own device data with a different modality).
Not putting words into your mouth that you aren't saying the latter here—just making a distinction since it's easy to imagine any data being a moat, which is a common mistake I see.