AI is supercharging the creation of maps around the world
tech.fb.com
tech.fb.com
Facebook wasn't the only company adding low-quality data: https://forum.openstreetmap.org/viewtopic.php?id=64430 (Previous discussion on HN: https://news.ycombinator.com/item?id=18723138 )
This article's quotes of what appear to be OpenStreetMap representatives are generally positive, so maybe that means they fixed all the problems they caused.
https://wiki.openstreetmap.org/wiki/TfL_Cycling_Infrastructu...
That gets me wondering - is the future of AI really just a semi-autonomous twilight zone where cheap / free labor augments an already faulty system? If not, what possible application is there for an expensive and closed automated system which works 100% requiring no human input, when other options are cheaper and leave clear directions for improvement?
In practice, almost nobody is even thinking about building a fully automated process for every case. The reason is simple: automating the first 60% of work takes x effort, automating the next 30% takes 10x, and automating the next 9% takes 100x and operating the final 1% is essentially impossible. So if you came to the table with the goal of 100% automation right out the gate you'd spent 10 years developing something with little to show.
I think full automation of some systems is possible, but is actually blocked by generational norms. By and large the systems that "Ops Plus" startups are attempting to automate were designed by people who are not digital natives. They're not illiterate, but things like instant messaging, async communication and and structured data are not natural primitives for them. I'm not saying everyone in the Fortnite generation is a master data modeller, but I think that when they join the workforce they'll set up systems that are much more feasible to automate.
This is the idea behind GANs, as it stands.
Hopefully they eventually release the ML pipeline itself as well.
Their RapiD editor has some similarities to a research project I was involved with: https://mapster.csail.mit.edu/maid.html
(https://www.youtube.com/watch?v=i-6nbuuX6NY vs https://tech.fb.com/wp-content/uploads/2019/07/add_ML_road.g...)
My understanding, from following the Australian OSM mailing list, is that it takes an individual to pursue this with a government agency, which is a ton of work, and often you'll just get a 'no'.
I suggest posting to the OSM talk@ mailing list, or if there's a local estonian one too
OP links to https://ai.facebook.com/blog/mapping-roads-through-deep-lear... which says that it's D-LinkNet specifically: http://openaccess.thecvf.com/content_cvpr_2018_workshops/pap... (more or less a Unet).
A huge amount of landcover segmentation in remote sensing still relies on simple models - either linear regression (thresholds) or classical machine learning like random forests or SVMs. For a lot of cases, these techniques will get you 90% of the way and it's very rare to have ground truth data that is accurate enough that you can measure the difference with any real degree of confidence.
A big problem in the field is the lack of good (public) ground truth. There's so little hand labelled data to work with that without humans in the loop it's extremely difficult to validate the results meaningfully (unless you have an army of staff to do it). With something like roads you could also have heuristics about what a road looks like and where it goes (e.g. it's a continuous thin line), which can help condition things.
I've seen a lot of papers which are applying deep learning for semantic segmentation for satellite mapping, but they evaluate on very limited datasets, they attempt to regress to simpler models without realising it (e.g. trying to predict a linear model), or they leak train and test data and report amazing results because they randomly split data from the same region.
I'm not saying that convnets aren't better than simpler models, but particularly for satellite imaging I'd take them with a pinch of salt and see what the improvement from a baseline method is. If you look at a random sampling of papers from the DeepGlobe competition, almost none of them provide the results from e.g. a cheap linear SVM.
Fun side note - several existing "famous" datasets generalise poorly to the developing world because most of the imagery is from the developed world (and even more specifically the West) and infrastructure looks totally different.
Have a look at mnist classification using a linear SVM, for example.
[1] https://www.csail.mit.edu/news/mitqcri-system-uses-machine-l...
Whenever Facebook releases something we should all directly worry about what their real motives are.
I worry that they'll try to embrace and extend in some way.
It's a reaction to Google Maps: a monopoly on high-quality up-to-date global maps with business location is dangerous to everyone else, as a chokepoint on mobile applications. It's less about 'acquiring data' and more about not being extorted by GM. Classic 'commoditize your complement' dynamics: https://www.gwern.net/Complement
Figure 3 demonstrates the scale of corporate OSM edits.