There are other datasets being developed that use high quality images that are manually labeled by humans, such as by Unstable Diffusion who's having a Kickstarter right now [0]. They say they will be able to get a much more high quality model due to such high quality images and captioning, so we'll see. They also want to make the model and code entirely open source rather than the license that Stable Diffusion has which is not open source (it has many restrictions, enforceable or not, on the images made).
[0] https://www.kickstarter.com/projects/unstablediffusion/unsta...