Lyft releases self-driving research dataset
medium.com
medium.com
This is true of pretty much any AI research. Look at Puffer[0], which was just on HN a couple of days ago. They're running a free streaming service just to get enough data to train their algorithms, and in fact mention in their FAQ that they would love to use commercial data if they could get it.
Unfortunately, academic and commercial incentives don't really align here. Most commercial entities don't want to share their data because it's valuable to them, and if they let researchers in, they want the output of the research to remain proprietary to their commercial enterprise.
I wonder if there isn't some sort of governance solution to this. Like give companies big tax breaks for sharing their data with researchers, or something like that. Essentially subsidize academia indirectly.
I think that's a really great idea. Not sure how many would take advantage of it but if it could be made to work then it would be really awesome.
It would also be extremely prone to abuse, though. Patenting is already an art of pretending to explain in clear terms what you are doing, while actually describing something as broadly and vaguely as possible. It would be pretty easy for a TON of things to leave out some key things that make it impossible or unhelpful to have the information.
You could form industry-specific regulations or even an active agency to prosecute abuses like that, but it would be immediately overwhelmed. The patent office is already heavily gamed by patent trolls, who bank on long odds for small judgements. Now imagine if millions or billions of dollars of taxes were on the line, and major companies were investing significant resources to open source while protecting their IP.
Even if that were all figured out, how would you value open sourcing stuff, even something as simple as data? Do you give breaks by size, importance, proportion of profit or future profit? Cost of the research? How do you guard against overvaluations and abuse of accounting? Even if you had perfectly accurate, annually-updated solutions for all that, companies can still game the system. Lyft has decided this dataset is what they need; if they could get a bigger break by collecting more data, they'd do that. Plus- facebook and google release tons of open source stuff. Do they deserve more than say, pharmaceutical research?[1]
Similar (IIRC Nixon) tax breaks already exist for R&D, and they are a notoriously abused loophole. Simplified but illustrative example: you build your R&D lab in the shape of a factory, do your research for a while and then suddenly scale back and replace it with machinery- well, the original building was still deducted from taxes.
Pharma is actually a perfect example. It's a well known fact that R&D only accounts for 22% of pharma industry revenue (almost equal to advertising at 19%), but only ~30% of that actually goes to new drugs. The rest takes advantage of marketing and the patent system to re-release drugs that are essentially the same. Two thirds of their research is obvious changes that are only protected because they owned the original patent- those shouldn't be getting the benefit of incentives.
You're commenting on an article in which a commercial entity is sharing their data despite it being valuable to them. Maybe they are the outlier but I've seen plenty of companies share data, especially in the ML space. Here are some datasets[0]. Maybe you would prefer more, but compared to other fields there is a lot of sharing. A "governance solution" could make things worse. If there was some mandate that companies that collect this data have to share it in a costly way, then it would discourage collection.
[0] https://blog.cambridgespark.com/50-free-machine-learning-dat...
What's their incentive to share?
1. If people use your dataset, they can do research into things relevant to you
2. Some people like working for companies that share data/code back with the community, helping hire & retain staff
3. Bits of publicity, either to potential engineers/researchers or others seeing lyft in a better light
4. Improvements in the domain, no matter where they come from, may be beneficial to your business
"A classic pattern in technology economics, identified by Joel Spolsky, is layers of the stack attempting to become monopolies while turning other layers into perfectly-competitive markets which are commoditized, in order to harvest most of the consumer surplus."
Even if you create an algorithm five times better and faster, you still lack the data to feed it..
* The raw data in nuscenes ( https://www.nuscenes.org/ ) is about 5x larger than this dataset from Lyft. 300GB train vs 60GB train. Argoverse ( https://www.argoverse.org/ ) is is about 3.x larger at 200GB. The Waymo dataset will (allegedly) be an order of magnitude larger than nuscenes ( https://i2.wp.com/syncedreview.com/wp-content/uploads/2019/0... ). BDD100k ( https://bair.berkeley.edu/blog/2018/05/30/bdd/ ) is the "largest" public dataset to date, but lacks lidar, and labels are inconsistent; most of the 100,000 scenes only have one labeled frame.
* The Lyft sensor suite has bumper-mounted lidar, which is absent from other existing datasets. Point cloud data in these areas is critical for pedestrians, bikes, and various road hazards. So this dataset alone is useful for validating work trained through other means.
* The current Lyft Level 5 release has no explicit test / validation set, which is crucial for properly measuring performance of any experiment one might do with the data. In nuscenes and Argoverse, there's a small snippet dataset that helps you prepare your pipeline. Feels like Lyft might have rushed things a little here-- they could have posted a "teaser" and then the full train and test/validation set a couple weeks later.
Great to see more public data (especially from a more modern sensor suite), plus investment into a contest with prizes.
Hmm this blog post and the website doesn't mention that this dataset was mostly annotated by Scale (scale.ai), as part of a partnership with Lyft ... We're going to publish a blog post about this soon, but if anyone at Lyft is reading this, please figure out how to reasonably credit Scale since I doubt leaving out Scale completely from the announcement is in the spirit of the agreement. Scale should probably also be added to the bibliography and website in some form
Contrast this with the nuScenes website, which was also annotated by Scale, and whose data format set the standard for this dataset: they credit Scale pretty reasonably
Dear Lyft marketing person who wrote this: we are a data labeling company, and you may think that means we have a bunch of useless bozos working here like most other data labeling companies, but that's not true - e.g, Steven is one of the smartest people in the world - https://stats.ioinformatics.org/people/3113 - he learns ridiculously quickly - e.g, gets to number one on random video games in a few weeks and learned to boulder L10 in a few months from scratch (normally takes years/decades and most climbers never get there)
Without some kind of gymnast background I honestly don't think it's possible. Even 0 to v10 in a year is hard to believe. I've heard of phenoms doing it in 2 or 3 years, and don't doubt a year is possible... but a few months? Kind of like going from couch to sub 5-minute mile in a few months.
Also has a fan club on HN. :D
But then this comment took it into a weird turn with how fast steven can learn rock climbing (seriously, i am still kind of unsure if we're talking about rock climbing because its so random and unrelated).
Genius
Wow dude (or dudess), I was on your side but you are losing me there.
You seem to be pretty pissed, and I hope you are anonymous and not Steven Hao... There is little room for emotion in business. Grow up.
FYI, you’re unlikely to solve any of your problems airing grievances on a public forum instead of just directly emailing the people involved.
Not if they leave attribution.
This comment does not represent the company's viewpoint, and cardigan is not speaking on behalf of Scale.
We are very excited to have been able to work with Lyft in open-sourcing this dataset and advancing the research community. We are also very grateful to Lyft for choosing to leverage our point cloud viewer and have credited the annotations to us on their launch page.
Good people come to comment on a minute notice!
Thank you.
(I don't really)
CEO Narrator: it was.
MULTI-KILL
BBBBUYOUT
PPPPAAAAYDAY
ACQUI-FFFFIRED
Also hopefully Scale will use this opportunity to educate team members about situations like this.
I wasn't involved in our communications with Lyft, so I was talking about something I didn't know much about. My audience was just the anonymous commenteriat: turns out a lot of people whose opinion makes a material difference to Lyft/Scale read these comments too. Sorry for not realizing that; I probably wouldn't have posted an uninformed personal opinion had I realized that.
I was being way too aggressive - genuinely sorry to anyone at Lyft who felt maligned by these comments. I woke up to 20 messages from coworkers who told me I was being an ass - genuinely sorry :(
Also I really should've clarified I was not speaking on behalf of the company: this was just a personal, uninformed opinion.
I cannot go into the details I learned about Scale's agreement with Lyft since it's confidential
Also, and I really mean this in a friendliest way possible, take a pause commenting here. You're not doing yourself a favor. I suggest you talk to your PR/marketing department before exposing more internal details, or run comments by them before posting.
I've had it with you Scale.AI people always trying to take credit for Lyft's work. We've been working weekends and nights for years, even the hourly workers have had to take unpaid overtime. All that time, I've never seen Scale.AI do extra work to help us before a big deadline.
But do not forget there will be 10s (if not 100s) of people working on this for 30 days. The man-hour this competition will use is highly disproportionate to the amount offered overall.
I guess it’s a decent opportunity if you’re trying to break into DL?
This was pretentious AF. People who win such competitions _allow companies_ to interview them sometimes, not the other way around. It's not like working at Lyft is some amazing privilege.
You might not think that working at Lyft is an amazing privilege, but it could be a good gig for someone that needs a paycheck and is looking for a job in the self-driving field.
I unofficially got first place after finding a bug in their test set (allowing me to blow the competition away). I reported the problem directly - they decided not to fix it and asked me to take my submission down. They said they'd still offer an interview.
However - this interview wasn't even for their DL team. They offered an interview with the web tools support team because they felt I didn't have enough experience...
Reference - https://github.com/bearpelican/lyft-perception-challenge
Imagine if it was a bug in their app leaking millions of users data and they go "We do not fix it. You can only use Uber now" "Also, come for an interview if you want to fix our web tools instead"
I would make an analogy where the training data is like a textbook. If I read in a textbook about how to design/build a bridge, I don't have to give royalties to the textbook author when my civil engineering and construction firm gets paid to build a bridge. The copyright/license of the textbook can't prevent me from using the knowledge gained from the book to do a commercial job. In a similar vein, the knowledge gained from a public data set is probably fair game for whatever you want, just as long as you aren't repackaging the data itself. There's probably a good boundary requiring someone to stop at a point where the model weights can't be used to reconstruct the original data.
Of course, other people can disagree. I would look forward to an actual legal opinion to clear this up.
You're confusing this CC license with open source licenses that do not restrict use but require derived works to be created/distributed under certain conditions. This CC license restricts use, in that you are not allowed to use the data for any commercial purpose - this is totally different from open source, and more like the "academic use only" licenses which used to be more common.
I doubt that any company would be willing to take the risk of creating a commercial self-driving car using this data to get a head start. I don't think I'll be seeing this particular license fought over in court.
In this case, imagine a student working on a class assignment. They use this data for purely academic purposes with no commercial intent in mind. After they train their system, they realize yow I could use this trained system and get rich. There was arguably no commercial use during the training. The use of the data was purely academic, like a person learning math or French. What you do after running the learning is a separate matter, just as using a CC licensed textbook to learn math doesn't prevent you from getting a job as a statistician.
Again, the tl;dr is instead of trying to divine how a court will deal with a poorly specified problem, it's much better to just not license your stuff using a CC license. There are almost always much better licenses to choose from.
(With, I suppose, hints about what restrictions you might be wanting to grant or prohibit by particular license options?)
If that sounds like a crappy legal footing for the next 20 years of software development, yeah, it is. It's also why Mitch Kapor and a few others founded and funded the EFF to try to encourage better case law decisions around the early days of electronic privacy law. We definitely need an EFF like effort around ML/DNN/etc., but I'm not holding my breath.
I wish I had a better answer, I'm mostly hoping someone else here does.
Car driving is in decline http://www.washingtonpost.com/blogs/wonkblog/files/2013/04/m...
A more concrete goal for transportation would be to reduce driving times even more by adopting remote work. That is reachable within the decade. In the meanwhile, car safety features should be ramped up, but autonomous driving so far doesn't seem very safe.
Imagine how much better things could be if everyone working on maps felt the same.
They are the only ones with a broad real world data source and seem to have wisely taken the right path by not adopting LIDAR, focusing purely on passive vision.
tesla decided to not include lidar because they couldn't find a manufacturer that would make one cheap enough for them/fell out over terms. Its not a statement of vision. Its exactly the same decision that Apple dropped Flash support for the iPhone, The processor and ram were too limited to support it, adobe refused to make compromises, and it was too late to change before launch.
Firstly, Tesla is not focussing purely on passive vision, they are using radar as well. But because radar is nowhere near high resolution enough, they need vision to provide categorization.
Now Musk makes a lot of noise about avoiding lidar, thats mostly because he knows its a massive gamble. Yes, he bleats on about its power budget and cost, but using pure AI cost a whole more in RnD, plus a boat load of latency. Not to mention the massive power budget needed to run the custom silicon.
_eventually_ vision + radar will be more than enough to provide life critical level 5 autonomy. However Tesla barely provide more than level 2.
They have a number of problems to overcome, rain/bug occlusion of vision sensor, low light performance, sunrise/sunset, fog, reliable realtime depth estimation, etc.
I suspect that CCD based time of flight depth sensors will become cheaper, low power, and small (They are almost certainly going to end up in mobile phones soon) before pure vision realtime life critical depth estimation is a thing.
Tesla has never been shy of using expensive components for their products. Their entire strategy relies on bringing down battery costs.
Musk makes a very good argument for his choice to not use LIDAR - nothing alive on this planet uses it. Every creature navigates our world on vision, even dolphins and bats use sonar as a secondary system. The proof of concept is everywhere, Tesla is just adopting it.
And yeah, they're purely level 2 at the moment. I don't agree with their marketing as it does lead some to use it as something more before its ready.
Nothing alive uses wheels, apart from us. Its a null argument.
I don't mind a gamble, what I mind is patent bullshit. If he'd been honest and said "Lidar is great, but its too expensive, and the manufactures are not willing to compromise" and left it at that, it'd be ok. But he didn't He dressed it up in semi prophetic _wank_. Now we have legions of "experts" blindly parroting that ToF sensors are dead and will be never used into the future.
I'm willing to be that a Lidar-like Time if Flight sensor will be shoved into a smartphone in the next 4 years. why? because it makes SLAM so much more easy/immersive/better. Which makes AR better, and more useful.
Once that happens, all bets are off.