Would this not be a better "middle" ground? Thoughts?
Would this not be a better "middle" ground? Thoughts?
If people really want to be careless they can get AI generated code from copilot or chatgpt on their own already, I don't think this would be worse than that.
It's a pity, because you've also posted good comments and I think the proportion of good comments has been getting better over time, which is great, but that doesn't make things like the above ok. Also, you have a history of using this site for ideological battle and we don't want that here—it's not what this site is for, and destroys what it is for.
If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html.
* Edit: on closer look, I can't tell if it might have been a bad joke instead.
[And don’t believe ChatGPT that claims it is an only a Language Model. It is not. It is a RL Agent trained with PPO.]
Does it apply to things that aren't programming? e.g. People are already using these AIs for legal work.
But, perhaps considering AI as an adversary is a bad idea. And alignment along the “Love is all you need” lines is the actual solution. Tricky problem…
So then the asker gets the wrong answer and no one else can correct it? Seems even worse.
If they stick to their guns, Stack Exchange won't be around in a decade. That's how fast the world is about to change.
Human evaluators can filter out the (for now) intermediate bad results. All of our AI models and products will climb a quality gradient year by year to the point where this will be increasingly less necessary.
To your other point, you can use machine data to bootstrap a "real" model. My team has done this to incredible effect.
ChatGPT (and presumably soon CoPilot) learn via * Reinforcement Learning from Human Feedback*[1] on top of the raw language model. This takes low numbers of high quality examples to improve the quality, and avoids the "junk training data" entropy problem you identify.
The OpenAI Codex text-davinci-002 and text-davinci-003 models have been trained for code generation using this human feedback process[2].
[1] https://huggingface.co/blog/rlhf
[2] https://beta.openai.com/docs/model-index-for-researchers
Technically, they aren't trained that way, their opposite recognition AI is trained that way. That's less than ideal when their original set of training data aren't drawn from their own output, only from a pool of humans influenced by their output. It becomes suicidal to the AI when it begins consuming its own output through what it considers to be trusted channels.
[edit] The ancient coder phrase "garbage in, garbage out" has never been more significant than now in the context of what's fed into neural nets and ML algos. I think it's basically silly hubris sprouting from a lack of underpinning technical understanding when people assert these models will train themselves without much more extreme means of assessing their output than can be provided by a few minimum wage workers somewhere.
No, they are trained that way. These aren't GANs or anything resembling that. There is a reinforcement model that is build from expert human guidance which is then used to fine-tune the LM.
From the HuggingFace explainer page I linked above: [there are] three core steps:
1. Pretraining a language model (LM),
2. gathering data and training a reward model, and
3. fine-tuning the LM with reinforcement learning.
Read https://arxiv.org/abs/2203.02155 if you want all the details (although the HF explainer is easier to follow). From the abstract of that paper:
> Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback.
There is no "opposite recognition AI".