1) Mis-identifying a lead as a non-lead.
2) Mis-identifying a non-lead as a lead.
I would guess that (1) would be more costly if non-leads go to somewhere where they aren't followed-up on. But I'd appreciate the insight.
Thanks again!
1) Mis-identifying a lead as a non-lead.
2) Mis-identifying a non-lead as a lead.
I would guess that (1) would be more costly if non-leads go to somewhere where they aren't followed-up on. But I'd appreciate the insight.
Thanks again!
Mis-identifying a lead as a non-lead is potentially loosing out on a big deal that can make or break your company. You never know what email will lead to a quickly closed $30k ARR sale, which are golden for any SaaS startup.
The reverse has almost no consequences unless you're really going to town with the emailing and end up being flagged for spam. Usually people just ignore you (not so smart) or write back that it's not relevant.
I don't think the scikit learn algorithm differentiates between the two types of errors, in terms of cost.
Though it seems to give less false negatives than false positives overall, when testing on new datasets.
I'll put in the f1 score in the article when I have time.