>> You put that information into the dataset. In supervised learning it's called labels.
I was assuming that the "general purpose computing" training dataset would be so unfathomably large that unsupervised learning would be a necessity.
Who is going to label the "general purpose computing" training dataset and how would it be persisted?