That's not all of what we are doing for at least a year, possibly few. LLMs are trained increasingly on generated inputs. Soon human sourced material is going to be rounding error in the process of training.
That's not all of what we are doing for at least a year, possibly few. LLMs are trained increasingly on generated inputs. Soon human sourced material is going to be rounding error in the process of training.
If you add two random numbers and calculate the result and those happened to be numbers noone else ever had idea to add you created a new piece of information. Template is not new, but the piece of information is. And sure, this template might be very simple, too simple, but you can come up with more complex one. And metadata is data. You can create templates in similar manner to how you create new pieces of information using them.
And labs training AI are doing it for years at this point. And it is ever increasing fraction of all training.