Show HN: Free database of geographic place names and geospatial data
github.com
github.com
The actual source along with the complete build process is described in "SOURCE.md". Sorry for possibility of confusion!
Not sure if we should remove "RESOURCES.md" altogether or just rename it to make clear what these links actually are.
A big difference from a CC-SA licence is that it is possible to make a produced work from OSM data and all you have to do is attribute OSM, there is no share-alike requirement. In this way, the OSM ODbL is less restrictive than a standard share-alike licence.
The main example of that is making a map image. You can make a map image from 100% OSM data, and that image doesn't have to be share-alike.
If you create a database as is the case of FreeGeoDB (and perhaps if you use OSM to geocode another database), then share-alike applies.
In practice, no-one's really done that or likely to do it. Either you'd do something silly like make the map a SVG with all data encoded as textual attributes and your "Computer Vision Algorithm" is basically grep (in which case it would probably be seen as a Derivative Database), or do real CV on a real image, which is very hard to do and will result in bad results. It's sufficiently hard that no-one's worried about it.
If you really don't like the OSM licence, you are free to go to another map data provider, pay them what they charge and agree to whatever they want, and get something else. If you want OSM, agree to OSM's terms.
Looking at the SQL file, "Zürich" is listed as "Zdrich", Munich is only shown as "Munich" without the German name "München" anywhere, and Cologne is listed as "Cologne" with the wrong "Koln" as alternative name.
Doesn't look like a reliable data source.
It's strange, I can't really understand how these accents can get mangled as these particular single letters.
I think I also have seen such character set corruptions before, but I can't remember how exactly you can get these particular corruptions.
We'd love to fix all those issues -- maybe with some help from the community :)
The goal is definitely to turn this into a more and more reliable data source every day.
The complete build process is described in "SOURCE.md". That's where we got these encoding issues from -- right from the source material. Maybe we didn't open or parse the source material correctly. But we were not able to get them without the encoding errors.
We'll love to fix all these issues :)
It's a nice start. I like it. Would it be possible to add the programm /instructions you used for converting the data from their original shape?
I would love to see an additional SQL version for PostGIS and the likes so that I can use their a large number of spatial functions to work with this data.
Not a big issues merely it's more convenient otherwise ;)
We'll love to extend the data to make it ready for PostGIS etc. Will be useful!
at least the points are in WKT, so should be able to be converted into something useful.
We know that Apache, MIT, etc. are usually for code and the Creative Commons licenses are for content (e.g. writing, images).
We thought that the data sets (CSV, JSON, SQL) are rather in the intersection of code and content, so the Apache license would be okay.
Is there anything specifically wrong with the Apache license for this type of project? We couldn't find any tangible downsides but we'd love to hear about any pros and cons.
your code would be the scripts you write to process it, and your content would be your data in my opinion
http://atlas.sollo.io/atlas/api
it's possible to easily navigate through resources (places data) and, a helpful feature is its capacity to have synonyms.
Numerous website use the Geonames dataset foe there work. (I also partisipated in this madness: http://www.wemakemaps.com/)
And, yeah, although I don't really care for licenses usually, here it makes me feel uneasy.
What makes you feel uneasy about the license? We'll love to fix this!
Regarding the airport lists, we're definitely open to merging in more complete and accurate data.
Three things we definitely wanted were (1) complete boundaries, (2) easiest programmatic access and (3) efficient collaboration.
Regarding the reference frame, that's definitely an issue. Right now, it follows the souce's guidelines (see "SOURCE.md"), which means "boundaries of sovereign states according to defacto status. We show who actually controls the situation on the ground. For instance, we show China and Taiwan as two separate states. But we show Palestine as part of Israel."
As outlined in "SOURCE.md", one part of the update process should be to continuously compare to the latest Natural Earth data.
Apart from that, we're totally open to contributions and ideas from the community!