I will say GPT-4 is mostly better than referencing stack-overflow or library documentation. Maybe once that gets fast/cheap enough to use as co-pilot we'll see some of these mythical productivity gains.
I will say GPT-4 is mostly better than referencing stack-overflow or library documentation. Maybe once that gets fast/cheap enough to use as co-pilot we'll see some of these mythical productivity gains.
The reason being that the source of material for GPT-4 is rotting underneath. As the internet gets ruined by AI garbage, the quality of output of AI will get worse, not better.
edit: This is a modern day tower of Babel event. We had a magnificent resource and then polluted it with garbage and now it's quickly becoming useless.
1) The data scraped from the web is filtered, deduped, ranked, and cleaned in a variety of other ways before being used for training. While quantity is necessary for a model to learn the structure of language, quality is even more important once you have a model that can produce coherent output so a lot of work goes into grooming the data.
2) There have been a bunch of papers and training runs that show synthetic data created specifically for training a model is as good or better than scraped human produced data. The importance of scraped web content is quickly declining because you can now generate infinite higher quality examples using existing trained models. The only relevance it has now is for knowledge about new developments and that is a much easier stream to filter since most of the important stuff comes from official sources and you don't need as many variations since you can just generate your own using one or more examples.
It's really good at spitting out code in common, familiar patterns, but the failure rate is very high for anything a little unusual.