> Labs spend billions hiring experts to generate new data
I thought this is mostly RL data. In my previous comment i was referring to pertaining data.
I thought this is mostly RL data. In my previous comment i was referring to pertaining data.
I would also be very shocked if they weren't filtering or prioritising existing pre-training data as well, for example to do curriculum learning or to avoid data that degrades performance.