Why Book Corners won't sync contributions back to OpenStreetMap
andreagrandi.it
andreagrandi.it
There needs to be a Single Point of Truth, and it should not be your little website.
BigCentralThing should be the official database that knows about where all the Whatevers are. Your little Whatever Locator site should only ever _pull_ from them. Need a new Whatever in the database? Get BigCentralThing to add it, then pull it in.
As an example, I run https://bettybeta.com/, a select guide to bouldering in Fontainebleau tailored for Women and short folk. We have a list of boulder problems in our database, but no way for anybody (even an admin) to add a new one. Every boulder problem record includes a link to its page at https://bleau.info/, who are dedicated to cataloging every climb in the forest.
Bleau.info is the Single Point of Truth. If one of our users does a First Ascent and wants to log it with us, we'll point them at that other site with instructions on how to get their climb in there, and tell them how to get us to import it. Any other path would lead to our database and the "official" database eventually diverging.
There are half a dozen other websites doing similar things to us. Some do what we do, but the ones that don't quickly find themselves out of sync and end up with a lot more work on their hands.
It's tough because, as the author says, writing the flow to import at our end is really easy (I've written it, before coming to this realization and hiding it again). But doing it is a trap.
I run valideparkeren.nl, a crowd sourced overview of all accessible parking spots in the Netherlands. Ideally, there should be a government-owned single source of truth, but the whole reason for starting this is because that source doesn't exist yet. I'm starting to work with the government to eventually hand this project over to them, which was the goal from the start, but the only way to make that happen is by first becoming the source of truth.
Edit: I was also planning to add my data to OSM at some point, but I'm reconsidering that after reading this. I completely understand that they only want high quality data, but in my case I'm convinced that high coverage and slightly lower quality is better than low coverage and high quality. Of course, transparency of how sure we are about the data is vital, so users can decide the risk they want to take.
Being The List of anything is a big, thankless job filled with nonstop toil to keep it up to date. It's not something you want to do unless it's truly your passion or somebody is paying you (a lot) to do it.
The problem is that building to tools to Be The List is pretty straightforward and fun. You only learn about the toil you've signed yourself up after the fact, once people start relying on it.
Good luck!
I worked on "openopeningstijden", where my aim was to improve the data for businesses in OSM and to disclose that in apis and uis.
There were other reasons why I had to pull the plug, but a big difficulty was the understandable conservatism of OSM. I didn't even want any tagging schemes changed, or the insane dsl for "opening hours" improved, at that time.
All I wanted was to allow obviously "wrong" data to be flagged with a note. Where "wrong" could easily be proven with links to other resources. "The shop is permanently closed. See this url at archive org about their announcement" or "contact details wrong. See their website at url and cross check with BigFoodDeliveryPlatform entry at ..."
Again. Reluctance for such automation is understandable wrt intelectual property rights. It becomes too easy for editors to copy in data that's not allowed.
But the result is that businesses on OSM are poorly represented, often many are missing. That data is hopelessly outdated. And that consumers use proprietary sources and apps to find e.g. a vegetarian restaurant in an unknown city, or to find out if the hairdresser is open at noon.
The Dutch OSM-community indeed is very conservative, a bit too much IMHO. We Belgians are more relaxed :)
Such a dataset would still be useful, but perhaps in a quality assurance tool, not as map notes that would clog the map. Can you link more information?
If "Map Notes clog the map" then there still is something needed in OSM to allow data that's not directly geographical, to be kept up to date, the POI data.
Aside from technicalities: if we want a free, open, etc etc alternative to Google Maps, we need to ensure POI data is of high quality, up to date, and comprehensive.
While it's very useful for many to have all traffic-lights, bike-lanes or waterways mapped, in practice, osmand, maps.me, organic-maps etc etc aren't used when Google maps shows accurate list of Bakeries (or campsites, or anything really) but the OSM-sourced ones don't.
Users, like me, only accept driving to e.g. a campsite that's been closed for over a year a very few times before grabbing Google Maps again. I'll edit OSM in these cases. Open Openingstijden was one of a few experiments where I researched "improving and maintaining POI data".
If we consider the amount of "alternatives to Google Maps" that now don't contribute back to OSM, because of technicalities, legal things etc etc, we miss a lot of good POI data that makes these alternatives viable in areas where Google Maps is the one and only king atm.
It's just a matter of dropping a link to that route's page into an importer to scrape it in and have it be usable at our end.
But yeah, for a space where the Big Database didn't often have a record for the things your users want to add, I could see it being more of an issue.
Does the submitter add to spot and then shoot you a message to pull by some unique id (author, submitter)? If you’re curating from their database, then spot doesn’t have the schema features to resolve to your subset by filtering alone, right?
It still lives in both places, but (assuming other websites also use the same Single Point Of Truth) if the title changes or whatever, everything stays synced up everywhere.
In my concrete example, bleau.info doesn't need to know that a given line has never had an ascent by anybody shorter than 165cm. So I write that down on my copy. But I don't need to track anything else about that line to use it and to link to Boolder, 27 Crags and 8a.nu's pages about it.
I was looking into building a canoeing map based on OSM data. An important thing to map is “is a particular stream between two lakes navigable by canoe?” OSM has a defined canoe=yes/no tag, but it’s supposed to be used for legal access, not navigability.
In the place I wanted to map, canoeing is basally legal everywhere, but there are many streams that no one ever canoes down, because it would be an enormous pain. There isn’t a way to represent this is OSM today, and community tends to be resistant to adding tags that map “subjective” data.
The solution would probably be for me to keep my own database of navigable streams with references to the OSM features, except that OSM features don’t really have stable ids. If another mapper comes along and decides to map an existing stream in more detail, they might choose to delete an existing stream feature and replace it, split it up into multiple features, or join several existing features together — all of which will result in id changes.
I have to deal with this in my little bouldering database as well, since the source I mentioned does in fact change or reuse its IDs every so often. There's already a periodic update that has to happen on any given boulder problem, to pull down changes from the other end. That also needs a way to self-heal if a record at the other end splits off into two pieces, and my ID now points to the wrong half of it.
It's all stuff you have to keep on top of, but your idea of keeping your own table with each record pointing back to OSM seems like the only path that avoids madness. Even it it does take some work to keep those references up to date.
At the same time, a point might be repurposed as something else entirely.
Though I kinda doubt they want to support maintaining these for every single project that wants one.
[1]: https://wiki.openstreetmap.org/wiki/Key:wikidata
[2]: https://wiki.openstreetmap.org/wiki/Key:gnis:feature_id
If you work this out, you'll get there
Had you considered `canoe=discouraged`?
https://wiki.openstreetmap.org/wiki/Key:canoe#Access_values
> paddling is technically allowed but may often be unsafe or impractical (e.g. open water or shallow creeks)
Making a workflow central to your business dependent on the goodwill and efficiency of an external entity is a very bad idea!
It's good that you've found a partner who is willing to put in the work, but that is luck rather than what normally happens.
I used YouTube as an example earlier. OSM is a good one too. Neither needs to care about you. They only need to care about their own job, which is knowing about their list of things (videos, roads, electronic part numbers, etc.). So long as they continue to do that, and have an easy way to tell them about the things they do care about, you're good.
They're already "putting in the work". Because it's their work to keep the list. That's all they do.
A project like OSM would be bombarded with spam and junk submissions if it didn’t have these barriers to submission. Understandable.
If their apps instead submitted directly to OSM (letting submission trickle back down), then the flow would fit into "user submission". Book Corners simply becomes an app though which the user performed the submission, rather than an intermediary exporting data.
(an import is when there isn't complete review...)
I'm not sure about this. The barriers to submissions are honestly more likely to cause object and amenity level data to become out of date. Google is able to have such an up to date place database because it relies on a large quantity of crowdsourced submissions, and employs consensus to figure out the truth.
As it is today you can largely expect OSM POI data to be incomplete or out of date unless you have someone particularly keen on keeping your local area up to date. OSM is best used as a background map layer for your own custom applications, like embedded maps in apps showing your own data overlayed. Rather than an alternative to the google maps app.
OSM doesn’t have that option. If it didn’t have these guardrails in place, a million CS students will see some open-looking data (or not even that, maybe something scraped) and throw it into OSM with a cursory Python script, resulting in an almighty mess. We need the process because we don’t have Google’s leverage of “we pay you”.
Eg. Draw me a map, but include only data points tagged with 'osm-license-verified-and-spam-filtered'.
That way users of the data get to decide their own tradeoff between legal risks, data freshness, spam, etc.
and people say hacker news has no sense of humour
Google denied our address existed[0] for years (which caused several problems with online shops that used Google Maps as the source of truth for address resolution) - took a good few tries to get them to actually add it.
Currently Google Maps is showing an "ALDI charging station" directly across the road when it is, in fact, about a mile down the road.
It's also showing a Chipotle just around the corner when, you'll not be surprised, there is no Chipotle there and, hilariously, the photo it has is of the Chipotle at the O2 Dome a good 10km away.
Absolute clusterfuck of nonsense is Google Maps.
[0] A 1960s London estate which was adequately represented on Google Maps and Streetview, mind.
I work in a warehouse, and if you type our address into Google Maps, it puts the maker right on our building. But they also label the building as a pub, two buildings down the street.
I've submitted corrections, and had a reply saying they were accepted. You can look on street view and the true location of both pub and warehouse are clearly visible, so there'd be no problem verifying my correction. Their map view includes the outline of the pub, as does their satellite view. The pub's been there 10 years, the warehouse since the 1990s. The Street View car has driven through the pub's car park three times.
And yet, the pub marker stays in the wrong place.
To contribute a bookcase, ask them to create an OSM account, log in with OAuth, and send the edit straight to OSM.
Book corners becomes just a custom view and editor over OSM data, instead of being its own database.
Users are now responsible for the data they submit (OSM already has tools to deal with vandalism etc...).
> Contact the relevant local communities affected by the contributions
Do I literally have to do new local community outreach every little town one of my bookcases shows up in? Ridiculous if so
if you're talking about the export restrictions of map data, that only applies to official, government-managed geodata. OSM explicitly doesn't rely on any other maps, so the law doesn't apply.
If every contribution required community outreach, we wouldn’t see any contributions.
These restrictions are unique to projects doing automated submissions of mass data.
If you're an individual just trying to make the world a better place, you can make an OSM account, go to the website, click, add a thing, hit submit, done. Someone will review it.
Apps like CoMaps and Organic Maps make it super easy to contribute data for businesses, landmarks, and such.
I will share an unpopular opinion: thil will only make a freeworkforce for those projects who profit from using OSM. And the OSM has surprisingly unreliable data quality for walking around even in the popular tourist places like Tokyo or Kyoto, or New York. I've tried, and it's so worse than GMaps (or Bing) to being completely unusable. No work hours of Starbucks, no local dining cafe, no menu contributions.
It's great for the trail and hike outdiors, sure. But that's very different from the city map, as the nature featutes rarely change, not like city POIs.
This moves the goal posts of OSM
- opening_hours=*
- amenity=cafe
- website:menu=*
But they can easily become outdated again. Still, local QA via StreetComplete, MapComplete, EveryDoor and a multitude of other apps is always needed.I don’t think you’ll find even most the die hard OSM fan disagree with you on the POI data.
On the other hand, the road network in OSM is very high quality. As you point out, it’s used and contributed to by companies like Lyft and Amazon.
The fact that anyone can download an accurate detailed global road network for free is pretty crazy.
Throughout nearly all of human history, one of the most difficult questions to answer was "where the hell am I?" and an even harder question was "how do I get to where I want to go?" Anyone can now answer both of those questions for free, within seconds, with accuracy within 3 or 4 yards. And, when it's wrong, they can contribute back to it instantly. OSM is the logical conclusion to one of the most difficult technical endeavours ever taken on by humans. It's astonishing when you think about it like that.
I would assume that tourist places don't have many locals going by and updating them and tourists don't want to fill in info on that kind of app, but that's mostly an adoption question. If more people contributed more info would be mapped.
I don't think this is an opinion, this is just an empirical fact. But it's good! OSM is a public service! This kind of thing should be public, individual citizens should have the agency to improve it, and businesses should benefit from its existence. It's the adopt-a-highway of cartography.
It's available from every users dashboard (it's a ~2 MB compressed GeoJSON file). It includes both libraries orignally imported from OSM and additional libraries added by users.
While I still think the whole OSM process to submit back is a bit too much for the time I have available (but I still understand and respect), at least I want to give users an additional option to retrieve the data (in addition to the public API).
If you have any other advices, I will be glad to hear them and, compatibly with my spare time, I will try to implement them (if they make sense and are useful for users) :)
Thanks
> There are also important licensing questions. A user’s permission to send a library to OSM is not automatically the same as having a sufficiently clear right to release that factual information under terms compatible with OSM
which would imply that you could not make available your database available for download under ODbL. Yet you have made it available under that license.
Maybe that part of the article needs further editing?
Also, if the data of your tool is now available under ODbL, then the tasks related to importing it into OSM can be effectively performed by anyone, not necessarily you.
Yes, I should edit it. I had not though about the bulk download until another user on Mastodon suggested it to me.
> Also, if the data of your tool is now available under ODbL, then the tasks related to importing it into OSM can be effectively performed by anyone, not necessarily you.
Definitely. I would be more than glad if libraries which are not on OSM were added back, I just don't have the required spare time to go through all the process and also maintain it.
https://wiki.openstreetmap.org/wiki/Key:image
The only difficulty is data than can not be easily represented in key=value pairs, such as timetables for buses. But even this can be resolved by adding external references.
This one sends data straight to OSM. I don't have my own database to "moderate", pictures are saved into a Panoramax-instance.
As users authenticate by logging in with OSM, the boring, risky legal stuff is handled by OpenStreetMap. Uploading images has a small print that says that their images will be reshared under CC-BY-SA.
The part where I suck is in creating a nice, polished website with extra pages, SEO, ...
Or set up a maproulette challenge with much the same detail.
Or just a geojson dump under an ODBL licence and share with with relevant communities. The third one or something very like it is required if you have mixed ODBL with non ODBL content.
I suppose if you want something badly enough you have to do it yourself, but it's not that the thing doesn't exist though it doesn't seem to. It's that I want others to want it too :)
Getting an official "charter sign" will cost you 50 bucks.
It could use GPS for rough geolocation, and 3D models of all the scenery could be generated. New contribution traces would contain changes compared to the past. Volunteers could request new paths it would like to see explored, and OSM could propose paths that cut through or ride along segments of recently submitted recordings of other users, to check if those changes are real, without checking the whole suspected recording.
Separate ground truth recording from its interpretation into mappable concepts, going out to make a recording or observation is a different task from deciding how to canonicalize the content.
I vaguely remember a blog post about doing photogrammetry out of their images, but I'm not entirely sure about it.
Make it work with the recording capabilities of a good modern phone and handle the compute in non-real-time.
They're already covering the blurring, as well as detecting useful objects in imagery (traffic signs for a start).
This sort of feels like a generational thing but I would have flown over.
Which is probably what might happen if you would have the option to "report them to osm". Someone should check things, because maybe you are a new person to them, maybe the spam filter might not be as robust as you might think, etc. Even if you are perfect, other submission might not be, so someone needs to check/review/etc.
I'd rather no data than bad data. There is no seperate baby & bath water. If a bulk source of data contains an unknown mix of good and bad data, that is all one big single item of bad data that is of no use to anyone.
Yes, you'll always need a user account
> Example of a valid and useful note can be "a new road was constructed here" or "this shop is closed and does not exist anymore".
Perhaps it's not that OSM that needs a signal, but perhaps it's that there's an opportunity for an open map that has many signals, of which OSM and Book Corners are examples.
Is it possible that we (myself included) have become conditioned to think content is llm generated as the default opinion?
> The workflow I had in mind was deliberately cautious
> The code was not the difficult part.
> These requirements are not a one-time form to complete and forget. They create an ongoing responsibility around the account, the documented process, community feedback, failures, and potential reversions.
Just a few examples. And yes, Pangram 4 also flags it as 100% LLM written. I don't mind being downvoted or flagged, but I think more people should be aware of the LLM style, even if they're ok with it being used without any disclaimer. It's honestly sad that nowadays people on HN cannot recognize this style.
I don't trust pangram 4 much. There was a post here earlier confirming it fails for many others.
Why would you want to have more people from anywhere else possibly raid them for anything useful? This way nobody is inclined to put anything worthwhile in there if it just ends up outside of their communities. Not that I'm in general pessimistic, just would like to know what is gained by putting the exact location into a map and not just have a list of neigborhoods with a book corner and optional photo on that neigborhood level, if people want to network around them.
I sometimes edit OSM, so I'm wondering whether any public bookshelf omissions are intentional, and adding the shelves would be unwelcome by the people running them.
I would not include the wishes of the people into it, because there are various groups that might have strange ideas (dunno, like only churches of some type should be shown; or no shops of a certain kind because it hurts my feelings, etc.).
I personally would not like any edits that advances the wishes of "any" group of people (like removing features). I do not like all things that exists on maps, but I will not remove them because of my feelings.