based it off this 2017 dataset from UCSD (https://cseweb.ucsd.edu/~jmcauley/datasets/goodreads.html)
some fun writing CUDA accelerated code to cluster the data and then find similarity scores
21 karma · joined May 5, 2020
some fun writing CUDA accelerated code to cluster the data and then find similarity scores
That is literally how openAI gets data for fine-tuning it's models, by testing it on real users and letting them supply data and use cases. (tool calling, computer use, thinking, all of these were championed by people outside and they had the data)