The idea maze for AI startups (2015)
cdixon.org
cdixon.org
Source: https://cdixon.org
So ... after he wrote this blog post?
https://cdixon.org/2010/01/03/the-next-big-thing-will-start-...
I was working on an typing autocorrect project and needed a corpus of "text messages". Most of the traditional NLP corpuses like those available through NLTK [0] aren't suitable. But it was easy to script ChatGPT to generate thousands of believable text messages by throwing random topics at it.
Similarly, you can synthesize a training dataset by giving GPT the outputs/labels and asking it to generate a variety of inputs. For sentiment analysis... "Give me 1000 negative movie reviews" and "Now give me 1000 positive movie reviews".
The Alpaca folks used GPT-3 to generate high-quality instruction-following datasets [1] based on a small set of human samples.
Etc.
The answer probably depends a lot on your specific problem domain and constraints, but a non-trivial amount of the time the answer will be that your task could be solved by a wrapper around the ChatGPT API.
And hope OpenAI forever provides the service, and at a reasonable price, latency, and volume?
Better quality training data might enable you to build a leaner more efficient model than is far cheaper to implement and run than the expensive model used to generate the data to train it.
See for example: https://twitter.com/SebastienBubeck/status/16713263696268533...
They are enjoying the be the market leader for now, but OpenAI will soon be facing real competition, and LLM services will become a commodity product. That must be partly why they seeked Microsoft backing: to be a part of the "big tech".
Oh you certainly could.
See here: GPT-3.5 outperforming elite crowdworkers on MTurk for Text annotation https://arxiv.org/abs/2303.15056
GPT-4 going toe to toe with expertrs (and significantly outperforming crowdworkers) on NLP tasks
https://www.artisana.ai/articles/gpt-4-outperforms-elite-cro...
I guess it will tke some time before the reality really sinks but the days of the artificial sota being obviously behind human efforts for NLP has come and gone.
At least that's true around the mean. If your application needs to handle long-tail cases, an LLM won't easily give you that. But depending on the application, that may not be necessary. So yeah, sometimes this is a bad idea, but for many applications it may be just fine.
This may apply to text too.
Partial or fully synthetic data is OK when finetuning existing LLMs. I personally discovered its not OK for finetuning ESRGAN. Not sure about diffusion models.
Diffusion models are still approximate density estimators, not explicit. They lose information because you don't have an unique mapping to the subsequent step. Got to think about the relationships of your image and preimage.
So while they have better distribution that GANs, they still aren't reliable for dataset synthesis. But they are better than GANs for that (GANs will be very mean focused, which is why we had such high quality images from them but we also see huge diversity issues and amplification of biases).
Human-curated synthetic data is commonly used in finetuning (or LoRa-training) for SD. I doubt that uncurated synthetic data would be very usable. There might be use cases where curating synthetic data with some kind of vision model would be valuable, but my intuition would be that it would be largely hit-or-miss and hard to predict.
No. Just no. Dear god, no.
This isn't too different from GPT-4 grading itself (looking at you MIT math problems)!
Current models don't accurately estimate the probability distribution of data, so they can't be reliable for dataset synthesis. Yes, synthesis can help, but you also have to specifically remember that typically they don't because they generate the highest likelihood data, which is already abundant. Getting non-mean data is the difficult part and without good density estimation you can't reliably do this. The density estimation networks are rather unpopular and haven't received nearly as much funding or research. Though I highly suggest it, but I'm biased because this is what I work in (explicit density estimation and generative modeling).
Only with heavy curation [0], otherwise your new models will be trained on progressively worse data than earlier models.
I would argue that the first step of the maze makes a ton of sense for the voice recognition/image classification/driving use-cases of 2015 that had binary outcomes, but now-a-days, what would it even mean for an LLM to be right 80% of the time? 8/10 words are predicted correctly? It can speak correctly on 80% of topics?
The reason people are so jazzed about generative AI is that it's not autonomously doing a task - it's helping a human operator by making (sometimes very useful) guesses on their behalf. It's much more of a tool than a solution (even if a lot of people want it to be a solution).
https://spark-public.s3.amazonaws.com/startup/lecture_slides...
Better to ignore folks who have no experience in building cutting edge product - he's just a average philosopher turned VC because it pays more.
Wondering if you're talking about that new one or a previous one
Then, the question becomes: how to create a great fault-tolerant UX?
There are some nice recent cases... Github Copilot is one...