Ask HN: Experience with weak labelling (e.g. Snorkel) for data annotation?
Weak labelling: writing heuristic rules to approximately label a dataset, and using those as a form of 'weak' supervision for training a machine learning model. Snorkel.org were pioneers of this approach.
Would love to hear any real world tales of trying it! What worked well and what didn't? How easy was it to get domain experts to write the rules? How did you mix ground truth data with probabilistic labels? that sort of thing.
context: we're building tools in this space.