Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
bloomberg.com
bloomberg.com
I just wanted to chime in that we're a YC company as well (S16), and I'm thankful to the HN community for having been supportive through our whole journey.
The problem I see with wading into other subfields (like my own) that need high quality training datasets, is that the datasets may be proprietary, and may not really overlap that much between companies in the same industry. For example, assembly line datasets for companies making almost the same product may be vastly different. I'm really struggling to see how you can possibly achieve the same scale in other industries.
Is it weird sharing the same name as a fashion icon ;)
And I'm curious about your ML "stack". Particularly the chicken and egg problem. Are you using something like Tensorflow with pre-trained binaries, perhaps from a vendor? Or is it 100% proprietary. Thanks!
Re 2—As with most companies working on ML these days, our stack is not fully proprietary. We don't take too strong an opinion on ML framework and use both Tensorflow and Pytorch currently. We generally use neural network architectures from the literature and then iterate on top of them to suit our unique problem requirements.
If I may, can you please tell us:
As your business has grown, what has changed the most in terms of how you run it?
What were some of the biggest challenges you've overcome and any major obstacles you see in the near future for the business?
Who are your mentors?
Thanks.
Overcome many challenges, but per my last answer, building a team of the best people has been the most important and most challenging. That, and learning how to do sales ;)
Too many mentors. People in Silicon Valley are incredibly helpful. To name a few: Dan Levine, Mike Volpi, Nat Friedman, Adam D’Angelo, Ilya Sukhar, Jonathan Swanson, Albert Ni, Jeff Arnold, Charlie Cheever, and Drew Houston to name a few. I’m very very lucky.
What principles/rules did you stick to when growing your company that you thought helped improve the culture/profits?
Thanks again for acknowledging the Hacker Network community!
It’s a small thing, but it’s surprising easy to spot once you look for it. And it really matters—startups are the business of building something from nothing. You need people who believe they can bend the earth.
I'm really looking forward to more of what Scale will do in the future!
In the meantime, check out our open source datasets:
https://scale.com/open-datasets/nuscenes https://scale.com/open-datasets/pandaset https://level5.lyft.com/dataset/
Second, self-driving as a problem space will need labels for a very long time. In an application where (1) verifiable model performance is paramount, and (2) the models need to be extremely robust for cars to be safe, the need for labeled data is only magnified.
There’s some social commentary in there somewhere.
One difference from "Not Hotdog" is that our data is used to power the algorithms of other AI/ML companies like OpenAI, Waymo, Lyft, etc., so it's imperative that we have impeccable quality. That necessitates humans to ensure accuracy, particularly in safety-critical applications like self-driving cars.
The $125M is around 12 large customers with contracts of $10M each, which buys them services of 2500 labeling contractors for 2000 hrs/year at $2/hr ($4K/yr).
At some point they will stop being a services company which carry a low multiple and switch to automated labeling without contractors (ala self driving cars) or develop some unique IP that they sell as a service?
Another risk is that the well-funded self-driving customers go belly-up. However, one important facet is that dead players don’t release much data. MobilEye has a vast dataset (including images from not just Tesla but other automakers) but that data isn’t going anywhere. Neither is Nvidia’s 180PB of HD recordings. (Release or transfer in part requires dealing with PII of the people in the recordings. Now if only the offshore labelers weren’t handed PII for free...).
The valuation is likely a forward-looking bet on AI as whole versus the current suite of contracts. Anybody using an off-the-shelf model will want some labels after their first proof of concept. I wouldn’t argue that the math makes sense but rather that demand does look underserved.
Machine learning indeed.
You can see some videos of what this looks like in this Twitter thread: https://twitter.com/BW/status/1158407524216909826
Once a sufficiently large corpus of human labeled data is available (across clients and datasets probably), that labeled data is used to train a 'first pass' labeling system.
It then becomes a virtuous cycle. Now the labeling is done in two phases. The first pass system makes its best guess, which is then reviewed by the existing human work force. Over time the first pass labeler gets better and better, till only very tricky/borderline cases need human intervention.
The end game is anyone's guess. Pretty clever biz model hack.
Note that this creates a positive feedback loop. I.e. as they get more results they can improve the initial stage.
If I had to guess, the long term plan probably is to move up the stack and sell the models to their clients.
From a technical perspective, can someone just post labelling task to mechanical Turk? what is the difference here?
Tomorrow (Maybe): Huge corpus of previously solved cases by humans, that can be used to train a custom model to replace the humans
What does it really matter how "old" the founder is, does the business have a workable business plan? Can it be profitable? Do people pay enough money for its goods and services to return a net income? Those are interesting questions. That it was started by a teenager is not, to my way of thinking, particularly relevant.
I'd much prefer that the article focus on these things which helps us understand the value that they bring to the market and what makes them unique.
[3 points] Scale AI (YC S16) raises $100M at $1B+ valuation to go beyond AI data labeling: https://news.ycombinator.com/item?id=20615657
[11 points] Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation: https://news.ycombinator.com/item?id=20614672
I understand it makes for a more click-attracting title, but imagine if the title was "Silicon Valley's Latest Unicorn Is Run By an Asian" - it sounds offensive, at least to me. Preferably the articles would focus on something other than a protected class.
Note that age is not technically a protected class, only "advanced age", defined as 40+, but I personally think that making unbounded age a protected class would be beneficial for everyone. Refusing to hire someone because they're "too young" is equally as offensive as refusing because they're "too old".
Customers flock to fame, press and investors do as well - they “know” where the customers will be so they go there, increasing the chances of success in a positive feedback loop. That’s how show business works.
San Francisco is the new Hollywood, basically.
This is in general not a good sign for the company or for the VC.
Me too. The era of lightweight clickbait headlines can't end soon enough.
Harshness is in order---regarding the business plan. But the age is a total red herring.
Don't fret, it is not likely to last much longer.
I like how the investors rationalized this devil's deal and the usurpation of the poor: "If you could be pulling a rickshaw or labeling data in an air-conditioned internet café, the latter is a better job."
This worked in part because my income was portable, so I was able to take a train to a more affordable area to get a place within my limited budget. This ability to move at will and take my income with me was historically largely limited to the Jet Set and comfortably well-off retirees.
For many people, doing gig work is a tremendous opportunity with a very big upside. It can be a huge improvement in both their standard of living and quality of life.
Most people decrying such labor arrangements aren't doing anything whatsoever to offer a better alternative. Color me unimpressed.
and this isn't wrong either
"If you could be pulling a rickshaw or labeling data in an air-conditioned internet café, the latter is a better job."