HNHacker News
TopNewBestAskShowJobs

ayw

438 karma · joined February 4, 2016

scale.com
submissionscomments
ayw··on AWS Data Exchange
At Scale (scale.com), we strongly believe that the “open-source” alternative to this is pretty critical.

We’ve built this index for autonomous driving datasets (https://scale.com/open-datasets) and are building that out for other domains right now.

Open source data has been a pillar to progress in ML (starting with ImageNet). It should continue to be the case that data that enables researches is sufficiently democratized.

ayw··on Ask HN: Why do so many startups claim machine learning is their long game?
Strongly agree with this.

One thing I’ll mention is that this is true both at the very early stages of a ML project, and even when an ML project is scaled up and in production. Oftentimes, the data pipeline is the true way in which a model will improve versus anything else, so it’s pretty critical that these data pipelines are setup to get an initial dataset but also to scale properly.

It’s one reason I started Scale (scale.com). It was viscerally clear that the real bottleneck to ML was getting the needed data, and in our case, annotating that data appropriately. It is very heartening to hear it echoed in this whole thread that data is very clearly what “matters” for ML.

ayw··on Fine-Tuning GPT-2 from Human Preferences
It's a good idea! They didn't demonstrate a lot of the inputs as the models were training, but that was very entertaining of course.
ayw··on Fine-Tuning GPT-2 from Human Preferences
founder of Scale (scale.com) here! We worked with OpenAI to produce the human preferences to power this research, and are generally very excited about it :)
ayw··on Is Elon Musk Wrong about Lidar? A Quantitative Study
1. While stereo depth estimation would work in theory, none of the self-driving cars actually have camera configurations that allow for stereo depth estimation (see here: https://electrek.co/wp-content/uploads/sites/3/2016/10/tesla...)

2. Stereo depth estimation is quite unreliable in practice because it requires you to match up pixels between the two images very precisely (1-2px difference can be a large disparity in distance), so it is not reliably used.

ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
We have many clients who have switched from Hive. There’s usually a step change improvement in quality and scalability—up to 10x improvement in error rates.
ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
We have a rule when hiring people—we look for people with an internal locus of control. Roughly speaking, this means people who believe they have control over outcomes in their life, as opposed to external forces beyond their control.

It’s a small thing, but it’s surprising easy to spot once you look for it. And it really matters—startups are the business of building something from nothing. You need people who believe they can bend the earth.

ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
The biggest change is your jobs goes from doing things (which makes sense) to building an incredible team that can do things (which is a more unintuitive job). In the limit, it’s always a people business.

Overcome many challenges, but per my last answer, building a team of the best people has been the most important and most challenging. That, and learning how to do sales ;)

Too many mentors. People in Silicon Valley are incredibly helpful. To name a few: Dan Levine, Mike Volpi, Nat Friedman, Adam D’Angelo, Ilya Sukhar, Jonathan Swanson, Albert Ni, Jeff Arnold, Charlie Cheever, and Drew Houston to name a few. I’m very very lucky.

ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
Self-driving is one of many applications of AI/ML to the real world, each of which likely requires high-quality labeled data to truly be production-ready. This includes other robotics, self-checkout like Amazon Go, natural language understanding, and more.

Second, self-driving as a problem space will need labels for a very long time. In an application where (1) verifiable model performance is paramount, and (2) the models need to be extremely robust for cars to be safe, the need for labeled data is only magnified.

ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
Thank you! We have some exciting stuff cooking that we can’t wait to share with everyone.

In the meantime, check out our open source datasets:

https://scale.com/open-datasets/nuscenes https://scale.com/open-datasets/pandaset https://level5.lyft.com/dataset/

ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
Re 1—It has been a bit of annoyance growing up (for example, Google autocorrects "Alexandr Wang" to "Alexander Wang"), but we run different circles ;)

Re 2—As with most companies working on ML these days, our stack is not fully proprietary. We don't take too strong an opinion on ML framework and use both Tensorflow and Pytorch currently. We generally use neural network architectures from the literature and then iterate on top of them to suit our unique problem requirements.

ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
To be clear, there is real machine learning that makes the labeling more efficient.

You can see some videos of what this looks like in this Twitter thread: https://twitter.com/BW/status/1158407524216909826

ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
I genuinely am not sure who you're talking about, but good to know!
ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
We do use AI and ML to help making the labeling process more efficient, but you are correct we do have scaled human insight that ensures very high quality.

One difference from "Not Hotdog" is that our data is used to power the algorithms of other AI/ML companies like OpenAI, Waymo, Lyft, etc., so it's imperative that we have impeccable quality. That necessitates humans to ensure accuracy, particularly in safety-critical applications like self-driving cars.

ayw··on Scale (YC S16) Raises $100M from Accel and Founders Fund at $1B Valuation
Hey everyone! I'm Alex, CEO/founder of Scale!

I just wanted to chime in that we're a YC company as well (S16), and I'm thankful to the HN community for having been supportive through our whole journey.

ayw··on Lyft releases self-driving research dataset
Hi, I'm the CEO of Scale.ai.

This comment does not represent the company's viewpoint, and cardigan is not speaking on behalf of Scale.

We are very excited to have been able to work with Lyft in open-sourcing this dataset and advancing the research community. We are also very grateful to Lyft for choosing to leverage our point cloud viewer and have credited the annotations to us on their launch page.

ayw··on Show HN: Training Data for Robot Dogs
:) we care about our labels!
ayw··on Show HN: Training Data for Robot Dogs
You never know what gets funded these days...
ayw··on How to make MongoDB not suck for analytics
You just need to read the oplog, so it only needs to track your saves.

In general, you probably should have at least something in your stack which reads all changes from your DB, at the very least for backup reasons.

ayw··on How to make MongoDB not suck for analytics
For better or for worse, MongoDB tends to be easier for developers move quickly, so it ends up getting adopted quite a bit. This is more about how to deal with it after it's already in your stack.
ayw··on How to make MongoDB not suck for analytics
Great idea!! :)
ayw··on Amazon’s Mechanical Turk Has Reinvented Research
Alex from Scale (www.scaleapi.com) here! We've taken an extremely quality-first approach and build out large workforces for datasets with high quality requirements and complexity. For example, we do a bunch of LIDAR / 3D labeling (https://www.scaleapi.com/sensor-fusion-annotation) which is very complex and labor intensive, and provide extremely high quality that would not be possible otherwise.
ayw··on Show HN: Scale 3D, API for 3D labeling of LIDAR, camera, and radar data
It's amazing what humans can do :)
ayw··on Show HN: Scale 3D, API for 3D labeling of LIDAR, camera, and radar data
Thanks for the tip. We'll definitely explore that data type!
ayw··on Show HN: Scale 3D, API for 3D labeling of LIDAR, camera, and radar data
That's part of the secret sauce :) But we work very hard to find ways to allow our Scalers to label incredibly complex data.
ayw··on Show HN: Scale 3D, API for 3D labeling of LIDAR, camera, and radar data
Thanks for the suggestion! If you try using it on mobile, we actually have joysticks for you to use which might work reasonably well. I'd try it out.
ayw··on Show HN: Scale 3D, API for 3D labeling of LIDAR, camera, and radar data
Hey — what browser / device are you using?
ayw··on Show HN: Scale 3D, API for 3D labeling of LIDAR, camera, and radar data
Hey everyone! I'm Alex, CEO and co-founder of Scale. One of the biggest bottlenecks to development in perception and vision for robotics and self-driving companies has been the ability to label 3D data. The ability to label LIDAR, camera, and radar data together has been able to massively accelerate our customers' timelines.

We've worked with a number of self-driving companies like GM Cruise, nuTonomy, Voyage, Embark, and more to build high-quality training datasets quickly leveraging our API. Scale is the perfect platform for this work—we're focused on really high quality data produced by humans via API.

ayw··on Expensify sent images with personal data to Mechanical Turkers
For anyone interested in solving this problem in their own business—we take privacy extremely seriously at Scale API (http://www.scaleapi.com) and implement numerous safeguards operationally and technically to ensure this doesn't happen.
ayw··on Creating a Modern OCR Pipeline Using Computer Vision and Deep Learning
Super cool!

Solving problems like DropTurk with stronger guarantees around turnaround/quality/confidentiality is exactly what we do at Scale. Would love to hear more about DropTurk’s experiences around building this internally and how it compares to our product:

https://www.scaleapi.com/ocr-transcription

Page 1 of 3Next →