Boosting Sales with Machine Learning
medium.com
medium.com
Is it typical that ~2000 samples (1000 leads, 1000 non-leads) is "enough"? I see that accuracy on the training set was ~86% and that the results are "starting to become useful for our sales team", but I would have thought more samples would have been needed. (I guess that since you're always collecting more real-world data you can continue to train the model and so it should get better?)
Getting more data is one of our top priorities going forward with this project.
1) Mis-identifying a lead as a non-lead.
2) Mis-identifying a non-lead as a lead.
I would guess that (1) would be more costly if non-leads go to somewhere where they aren't followed-up on. But I'd appreciate the insight.
Thanks again!
Mis-identifying a lead as a non-lead is potentially loosing out on a big deal that can make or break your company. You never know what email will lead to a quickly closed $30k ARR sale, which are golden for any SaaS startup.
The reverse has almost no consequences unless you're really going to town with the emailing and end up being flagged for spam. Usually people just ignore you (not so smart) or write back that it's not relevant.
I don't think the scikit learn algorithm differentiates between the two types of errors, in terms of cost.
Though it seems to give less false negatives than false positives overall, when testing on new datasets.
I'll put in the f1 score in the article when I have time.
If they were trying to detect pedestrians from a self-driving car then they'd need a lot more accuracy and so a lot more training data.
In terms of practical machine learning, plotting learning curves are a great way to know if samples are "enough" [0, 1]. If your algorithm under fits (high bias) both errors (training and validation) will be high and in this case, adding training samples would not help.
[0] http://www.astroml.org/sklearn_tutorial/practical.html
[1] http://blog.fliptop.com/blog/2015/03/02/bias-variance-and-ov...
In the EU this is not public information, and as such there's no scalable way other than experience to find out if a company ships 500+ TEU.
After nine months I can easily spot what stands out about companies that ship 500+ TEU from a description, but this is by far faster.
Some sort of little financial coach that asks you about purchase patterns, your state of mind before and after, how you feel about the whole thing in retrospect, and why it is you think you 'need' this thing right now.
If it turns out you compulsively buy baseball hats for the wrong team when you're drunk, maybe your phone should ping you if you walk into a sporting goods shop after you've been in a bar for three hours. Then it shows you a picture of your daughter and reminds you that you PROMISED that this summer you'd take her to Disney World.
Also, I can't help but smirk at the thought of mentioning Disney a couple sentences after talking about an "escalating arms race of advertising tools and trickery". Perhaps the phone should have pinged before you made that promise too.
>We're talking about enterprise sales in the logistics and transportation industry here. I doubt the final decision whether or not to buy this particular freight rate benchmarking tool is being done on an impulse. There are whole teams responsible for enterprise purchases who have already ripped apart this offer and know every common sales trick in the book.
Within the software industry this is a cliche. Someone who hasn't touched code in 15-infinity years makes a buy order for a demonstrably inferior product, and we waste $500k in labor and overhead costs so that he doesn't look bad for making a $200k order for solutions nobody wanted. Whether they intend to or not, tools like these are going to pick up on patterns of weak judgement and exploit them. Really the same problem with A/B testing.
> Also, I can't help but smirk at the thought of mentioning Disney a couple sentences after talking about an "escalating arms race of advertising tools and trickery". Perhaps the phone should have pinged before you made that promise too.
Haha. Touche. It was the first thing that popped into my head when trying to think of a common social obligation that is difficult to fulfill if you can't manage your finances.
Because I read an article about a team using publicly available data about public companies, then training an algorithm to comb that data to determine which ones were likely customers (which would save their salespeople time and, probably, save time for unlikely customers who are no longer receiving an unwanted solicitation).
It's not that I don't think it's useful, I just wonder about the ROI for cash constrained businesses.
The most important thing will be the training data. You need a good number of samples, and the data also needs to be reasonably "clean".
To put it another way, the business case feels week most [i.e. small] profit seeking enterprises.
Scikit-learn implements things like grid search natively [0], and tools like SigOpt [1] (YC W15, disclosure: I'm a founder) do this automatically as well.
[0]: http://scikit-learn.org/stable/modules/generated/sklearn.gri...
A popular alternative is to use a distributed word embedding such as word2vec[1], where similar words are grouped together in the vectorspace.
Edit: If there are few observations, like in this case, we don't need to train the word2vec model on the dataset itself. We can use pre-trained word embeddings such as the one publicly released by Google which was trained on the Google News dataset.
Random Forest won in the samples tried, but I wager a support vector machine with a histogram kernel would do fantastic.
edit: im not saying dont try it, i certainly would! lets look at the data on github? maybe we could have a wack at the data
I haven't been very involved in using random forest at work (yet, I hope to), but I've done various mathematical programming work in the past to generate business insights (mainly through regressions and linear programming/optimization).
One thing that you make very clear through this blog post is how much value comes from the "non-technical" aspects of mathematical decision analysis. You have to see the application, find the data, clean the data, figure out what to actually put into the model, and get results in a way that can lead to an actual outcome with value.
Here's the thing, the reason I put "non-technical" in quotations is that it's actually a mix of technical and non-technical. You need to be aware of how these algorithms work and how they are implemented in order to have that insight. There's the old statement that everything looks like a nail to someone with a hammer, but knowing what tools are and what they can do can help frame issues in a way that you can approach them from new angles. This is why I do think it's worth learning various ML and other algorithms (like LP, NLP, etc) through contrived examples - once you understand them, you'll start to see the opportunity to apply them.
One last thing - kaggle. Kaggle is super fun, and I highly recommend it for people looking for an opportunity to try this out and learn it. However, good real world data science probably has less to do with making exceptional refinements to models. You know that data set you get when you are doing a kaggle competition? That's a huge amount of the actual work, right there.
You can do so much with basic RF and KNN (and with LP for that matter). This post is a pretty good illustration of this.
Anyway, pretty cool, thanks for sharing.
If anyone has any questions about our sales process and how we use this day to day, fire away!
We're constantly trying to improve our company data, social especially, stay tuned for that. That said, I'd love to hear any feedback you have at the same address.
Any plans to open-source this?
The concept would involve processing millions of companies names found on the "Bill TO" field of sales records. Then using these records to populate a ElasticSearch index for use with Graph Query API to help further normalize/dedup the company names that share similar string semantics. The next stage of the process would be to scan the normalized, dedupped, list of company names and attempt to locate the company website URL by crawling the first page of Google search results. This would need to be metered because I assume Google would block me if I performed rapid attempts. After gathering a list of company URLs the plan would then shift gears into attempting to identify if any of the companies websites contain the typical components that make up an eCommerce website. Think searching the HTML for all variations of "add to cart", "shopping cart", "my account", etc.
This is arguably the toughest problem in the industry, and solutions by Google, Adobe, etc. are just starting to make headway, but are still very expensive and very custom with few exceptions (see the new data-driven attribution release for AdWords for example).
nlp.tokenize(text, { dont_combine: true }).reduce(function(sentences, nlp, key){
var words = nlp.tokens.reduce(function(tokens, token){
// ...
});
// ...
});Just fyi, you can usually buy lists of buyers in target companies. They exist for most markets, though it's possible such a list may not exist for your market. These lists will give you actual names and contact information, and are probably a more efficient way to contact potential buyers.
Edit: Added URL
I just think the title was a bit sensationalist.
Comments that are only dismissive, or are supercilious about knowing more than others, are deprecated here. It would be a good idea to read the following which describe what we're looking for on the site:
https://news.ycombinator.com/newsguidelines.html
https://news.ycombinator.com/newswelcome.html
Oh, and welcome to Hacker News! (I'm a moderator here.)
Having worked in business-related mathematics (did an MS in industrial engineering/ops research), I have definitely noticed how critical the "non-technical" aspects are, and how much mileage you can get out of relatively basic stuff if you do those steps well (and how little mileage you get out of sophisticated stuff if you haven't).