It doesn't work: even for the tiny slice of human work that is so well defined and easily assessed that it is sent out to freelancers on sites like Fiverr, AI mostly can't do it. We've had years to try this now, the lack of any compelling AI work is proof that it can't be done with current technology.
You can't build on top of it: unlike foundational technologies like the internet, AI can only be used to build one product, a chatbot. The output of an AI is natural language and it's not reliable. How are you going to meaningfully process that output? The only computer system that can process natural language is an AI, so all you can do is feed one AI into another. And how do you assess accuracy? Again, your only tool is an AI, so your only option is to ask AI 2 if AI 1 is hallucinating, and AI 2 will happily hallucinate its own answer. It's like The Cat in the Hat Comes Back, Cat E trying to clean up the mess Cat D made trying to clean up the mess Cat C made and so on.
And it won't get any better. LLMs can't meaningfully assess their training data, they are statistical constructions. We've already squeezed about all we can from the training corpora we have, more GPUs and parameters won't make a meaningful difference. We've succeeded at creating a near-perfect statistical model of wikipedia and reddit and so on, it's just not very useful even if it is endlessly amusing for some people.