Overture Maps Foundation releases open map dataset
overturemaps.org
overturemaps.org
The consortium intends to enable a framework for enhancing geospatial data based in open data sets (e.g. OSM) with their own proprietary processes, and re-release it with a permissive license (the Community Database License Agreement - CDLAv2), but keep the data and processes required to create that dataset proprietary.
The project has created a lot of conversation in the OpenStreetMap community, but in general I think it's good to see so many resources put into the OSM-adjacent world.
I am not a lawyer, but I would assume someone at Overture has deemed their license (the CDLAv2 [1]) to be "compatible" with ODbL. (edit: From the link in OP, the OSM-based data is being released with ODbL. I think I'm wrong about license compatibility. Again, not a lawyer.).
[0]: https://opendatacommons.org/licenses/odbl/1-0/ [1]: https://cdla.dev/permissive-2-0/
From a brief skim of the landing page, it looks like the OSM-derived content is not being offered permissively. See https://overturemaps.org/download/.
Are they releasing OSM data or OSM derived data on CDLAv2 anywhere, right now?
There was issue like this in past but it was resolved.
That's a thing? This must explain why Google Maps turned into shit over the last few years. Google seems obsessed with replacing humans with half-assed automation.
Imagine how useful tools like that can be if they just had a human review process. That added step in the process not only improves the quality of the final data but provides important feedback to improve the tool itself
IMHO because now I have twice as much work: knowing the problem domain and reviewing someone else's contribution. It's like code review 24/7 from untrusted contributors: it requires more focus than just trying to understand and author the original contribution
Well, I guess one can just continue to prompt any such AI over and over "have you done it correctly? No? Then try harder"
I recognize that my experiences and expectations from AI differ from seemingly the vast majority of users who get benefit from it. I'm glad for them, but I don't ever want to opt-in to having "AI assistance" because with the current state of the art it generates more work for me than value in my life
The places layer would be invaluable for (for example, stuff I'm working on now) geospatial healthcare access analytics if I could rely on the places being a reasonably accurate source of provider locations.
It's a lot of work to tie all the CMS data with health plan data etc.. and then to geocode it all with OSM tiger geocoders. If I could rely on that data already being present and scrubbed, it would be a godsend.
Totally understand if you can’t discuss (or don’t want to), but it’s in my domain and I’ve spent a fair amount of time thinking about how to do it without coming up with much, so I’m super curious about your thoughts.
So you can setup something like the following:
S3 (host the myfile.pmtiles) <- Lambda (takes the x/y/z from the path and requests the correct range) <- Cloudfront cache tile response
Then you can setup tiles.mydomain.com (or use the cloudfront domain directly) and then use Leaflet or similar on the frontend to fetch/render the tiles. For Leaflet you use the protomaps plugin/lib and give it a url like "https://tiles.yourdomain.com/20230408/{z}/{x}/{y}.mvt" where "20230408" maps to "20230408.pmtiles" in your S3 bucket. Now I can drop new pmtiles files into that bucket and update my clients to use the new source. And since the tiles are in vector format you can theme them however you want in the client which is neat. Lastly you don't have to use the 100+GB whole-earth tileset. You can use a tool [1] (provided by the same guy) to download a dataset for just a given geographical region.The .pmtiles file is a little over 100GB but the whole setup took me only an hour or two max to get running and will cost way less than Google Maps to run.
digging into their GH repos surfaces https://protomaps.github.io/PMTiles/?url=https%3A%2F%2Fr2-pu... which is at least more detailed if not quite a GMaps killer
Edit: And I missed parents link to https://app.protomaps.com/downloads/small_map which is actually better.
It's important to me that I showcase the system working as advertised - you are free to download the 110GB file from that static storage with no restrictions. Unfortunately, the demo is slow because I'm hosting it on Cloudflare R2 which has no outgoing bandwidth fees. Before, it was on DigitalOcean Spaces, which is low latency but would cost me $1-2 out of my own pocket each time someone clicked the link.
The Lambda/Workers integration as described in the other comment is the preferred solution to this for developers operating this at planet scale. Of course, the static storage solution is perfect for small to medium-sized areas, even on GitHub Pages.
Can you expand on this a little? Cloudflare is supposed to be pretty competent as CDN and with caching.
https://protomaps.com/docs/cdn/cloudflare
However, it is an optional acceleration layer on top of the PMTiles access pattern; and it doesn't make a compelling demo, it's just a Z/X/Y layer which is how map tiles have worked for the past 20 years.
To expand a bit -- There are adapters for mapbox(gl/libre) and leaflet, and so on that will allow you to use a .pmtiles directly from the client with a range request supporting server, like s3 or nginx.
I've used protomaps basemaps to generate an Ireland and EU basemaps for a project from OSM data. I still haven't quite figured out the tuning for which features should be shown, combined, or excluded yet, and the generated layer names are sort of, but not quite compatible with similar base layers + styles from mapbox (which uses the same base data, but a different, proprietary conversion). This part takes a reasonable skill as a GIS person and some taste in the design of the feature symbology.
The downside of doing this on S3 directly is that you might wind up publishing a link to a 100 gig file that would cost $10 if someone just downloaded it. Lambda invocations seem cheap compared to that risk. OTOH, throwing it on a hetzner box is quite reasonable.
This is the focus of most development for the next few months - like you said, it's also a matter of aesthetics and context, like if the map is underlying another PMTiles data layer it should be lower contrast then if it's overlaid with pin markers. The end goal is to have a flexible basemap solution adaptable to most apps but you are free to roll your own proprietary style on top of the open source tiles if you'd like (there are already companies doing this)
I've got a couple of goals, or at least evaluations for goals --
1) Have a stripped down basemap like Carto Light, but with different emphasis for the "bicycling friendly" roads, and selectively higher contrast/emphasis for those minor through roads that discourage car traffic but are perfect for a bike. (e.g. Irish boreens, those 1 lane or double track "paved" L roads that are everywhere outside of the major cities.) I find that I simply can't see the minor roads on apple/google maps when I'm out unless I'm zoomed in to the point that I can barely see a network. (fading eyesight and super low contrast).
2) Another lighter base map, but with transit focused features to compliment some of the open data in Ireland around transit -- the routes and timetables and so on. I'm aiming for a site that can be 100% statically hosted but will show the routes in a more friendly manner than Bus Eireann's site.
3) If it goes well, I might be doing this in a more commercial context for some clients who are currently using mapbox for their basemaps but due to political concerns need to insert in different names/boundaries for disputed areas/features. We'd love to be able to self host, but quality is a major concern there.
The bugs that I'm seeing in the 1 and 2 cases are things like handling the look of freeway interchanges/flyovers at close zooms, consistently getting rivers to show at appropriate zooms because they're into two separate feature types depending on the width, with the narrow bits getting dropped. And the general mismatch between the styling I'm used to from the mapbox converted layer names/feature types and what's coming out of pmtiles/basemaps. The feature bits look like tippecanoe coalesce/drop densest features tuning, but I'm also looking at just dropping out entire feature sets to drop the size of the tiles. That should help the coalesce/dropping behavior as well.
OSM is a freeform dataset and not a cartographic product, so most of the basemap work now and in the future is on getting good results of this transformation for 200+ countries. The mismatch vs. existing maps must exist because we need to ensure to downstream commercial users that all end products of the map generation are openly licensed - that's why this is being pursued as an independent, self-funded project with support from GitHub Sponsors: http://github.com/sponsors/protomaps
With 512MB RAM these generally complete in under 100ms so the unit costs of Lambda are within a few multiples of Cloudflare Workers (.5USD / million invocations)
Do you have any tips for doing so without renting memory in the cloud?
I have my own tiny personal project for maps renderer using Leaflet. For now I use 3 sources for tiles (Google, OSM and Esri) but having more sources would be very cool.
Oh that is going to be fun. If I recall correctly Google Maps alters the boundaries of places based on the views of the location the map is being requested from to avoid getting in the middle of disputes.
Not "correctly" showing boundaries is a crime in many countries.
Edit: here is a source https://qz.com/224821/see-how-borders-change-on-google-maps-...
And unless you have millions and millions of users, no one will care even in those countries.
There's nothing stopping either AWS or Azure from granting `s3:Get` and `s3:List` on those buckets to enable unauthenticated reads, if they were thus interested
aws --no-sign-request s3 ls s3://overturemaps-us-west-2/release/2023-07-26-alpha.0/
works fine for me.
Overture Maps appears to be quite a closed and proprietary project, with claims of openness limited to being able to download a data set and accompanying schema specification. Some issues that immediately come to mind:
1. There is no published description for how the data was generated. End users thus are given no assurance of how accurate and complete the data is.
a. As an example, administrative boundaries are frightfully complex and include disputed boundaries, significant ambiguity in definition of boundaries, and trade-off between precision of boundaries versus performance of algorithms using administrative boundary data. Which definition of a boundary does Overture Maps adhere to, or can it support multiple definitions?
b. It's probable that Microsoft have contributed ld+json/microdata geographic data from BingBot crawls of the Internet. This data is notoriously incorrect, including fields mixed up and invalidly repurposed, "CLOSED" in field names to denote closure of a place 5 years ago but the web page remains online, and much ambiguity in opening hours specifications. For AllThePlaces, many of the spiders developed require human consideration, sometimes of considerable complexity, to piece together horribly messy data that is published by shop and restaurant franchises, and other organisations providing location data via their websites.
c. For location information where +/- 1-5m accuracy and precision may be required (e.g. individual shops within a shopping centre[3]), source data is typically provided by the authoritative sources with 1mm precision and +/- 10-100m accuracy. AllThePlaces, Overture Maps, Google Maps and similar still need human editors (OpenStreetMap editors) to do on-the-ground surveys to pinpoint precise locations and to standardise the definition of a location (e.g. for a point, should it be the centroid of the largest regular polygon which could be placed in the overall irregular polygon, the center of mass of a planar lamina, the location of the main entrance, or some other definition?).
d. If Overture Maps is dependent on BingBot for place data, they'll miss an enormous number of points of interest that BingBot would never be able to find. For example, an undocumented REST/JSON/GraphQL API call or modification to parameters to an observed store locator API call may be necessary to return all locations and relevant fields of data. Website developers routinely do stupid things with robots.txt such as instruct a bot to crawl 10k pages (1GB+) from a sitemap last updated 5 years ago rather than make 10 fast API calls for up-to-date data (5MB). Overture Maps would be free to consume data from AllThePlaces as it is CC-0 licensed, and possibly correlate it with other data sources such as BingBot crawl data, a government database of licensed commercial premises or postal address geocoding data. However the messiness of data in various sources would be approaching impossible to reconcile, even for humans, and Overture Maps would possibly have to decide whether to err on the side of having duplicates, or lack completeness.
2. There is no published tooling for how someone else can reproduce the same data.
a. AllThePlaces users fairly frequently experience the wrath of Cloudflare, Imperva and other Internet-breaking third parties, as well as custom geographic blocking schemes and more rarely, overzealous rate limiting mechanisms. If Overture Maps is dependent on BingBot crawls, they'll have a slight advantage over AllThePlaces due to deliberate whitelisting of BingBot from the likes of Cloudflare, Imperva, customer firewalls, etc. However, no matter whether you're AllThePlaces or Overture Maps or anyone else, if you want to capture as many points of interest as possible across the world, use of residential ISP subnets and anti-bot-detection software is increasingly required. They'll need people in dozens of countries each crawling websites targeted to the same country, using residential ISP address space. Otherwise they end up with an American view of the world, or a European view of the world, or something else that isn't the full picture.
b. If Overture Maps has locations incorrect for a franchise/brand due to a data cleansing problem or sourcing data from a bad source (perhaps non-authoritative), there are no software repositories for the franchise/brand to raise an issue or submit a patch against.
[1] https://www.alltheplaces.xyz/
[3] Example Australian shopping centre as captured by AllThePlaces: https://www.alltheplaces.xyz/map/#18.07/-33.834646/150.98952...
This is true, and is actually the whole point of Overture.
Overture was developed to enable private companies to leverage open data (like OpenStreetMap) but also combine it with their proprietary data and processes.
The intention is to share the result with a relatively permissive license (a new thing called the Community Database License Agreement) but keep the process and underlying data proprietary.
I'm big on OpenStreetMap, but I can't deny its a bit of a liability for Facebook and other companies that display maps to their users. There is the occasional vandalism edit that simply can't be shown to an end-user. Facebook put significant effort into maintaining a moderated version of the OSM database that lags behind the real-time edits. Facebook, Microsoft, TomTom, etc. know this is a ton of work and want to pool their resources. Making it open also helps to openly compete with Google, the other big map data provider.
If you want to contribute to Overture as an end-user, AFAICT your best option is to edit OpenStreetMap and see if your changes eventually get pulled in. Overture has promised the OSM community that they'll make much of their data available to be contributed back, we'll see if that pans out.
When it comes to AllThePlaces -- as an OSM nerd it seems like there is an opportunity to build a better bridge between this and OpenStreetMap, to make it easier to quickly update businesses in an area. Recently there has been a pretty successful push to link OSM data with WikiData, using tools like the OSM ↔ Wikidata matcher [0]. For POIs, it's a lot of work to add a bunch of local businesses, even with tools like EveryDoor [1]. It would be so cool to see AllThePlaces integration into RapID for example, if there isn't already(?)
I remember the same sort of arguments being made about how web sites could not possibly ever accept user comments or submissions, or could not ever risk having users sending links to one another. Those all proved to be false.
The heyday of comments sections on news websites is now in the past. Not long ago, Lonely Planet took down its renowned Thorn Tree forums, which had been a big part of the travel internet since the 1990s. Friends who run a major website for a particular hobby told me that they canceled their plans to launch a forum, since their site is advertising-supported and a forum could damage their relationships with advertisers.
Reliable moderation costs money, and if you don’t moderate heavily enough, you’re going to get user comments that tarnish your brand (or at least scare execs into thinking that the brand will be tarnished).
The OSM basemap is used in many official publications, in many social media applications, etc. I'd actually recommend using the basemap for most simple mapping cases, as long as it's being continuously updated from upstream (or, it is the upstream basemap). If you take a snapshot of that data, however, you risk capturing some bad stuff. That is a real risk for a company like Meta.
Once in 19 years, I think?
It’s not great but let’s not pretend it’s more of a problem than it is.
https://en.wikipedia.org/w/index.php?title=The_Rescuers&oldi...
Interestingly, they’re also a lot harder to catch than the slur naming example
What this is taking about is closer to Wikipedia. And vandalism is a real thing there, so portion of topics is moderated.
Funny thing is that this part is most often the cause of a data breach when looking at majority of pentesting reports.
One thing is to expose something to few ppl you know and another thing is a possibility to send things to millions of ppl that are constantly abused like Twitter or Youtube.
Part of the problem is not entirely clear copyright/copyright-like status of this dataset. Thanks to https://en.wikipedia.org/wiki/Database_right and similar things (in general OSM is really careful with legal status of datasets being imported).
As always, I will be vigilant about checking licenses before adding data to OSM!
[0] https://github.com/alltheplaces/alltheplaces/issues/5133
Things like it could come from alexa, pc or any devices devices with forced opt ins that keep scanning all your neighbours wifi networks and mac addresses.
Believe there was also an initiative where amazon devices will provide adhoc internet connectivity by piggy backing on other amazon devices on different networks with connectivity.
So all the openness but without any controls. There should already be a better term for things like this.
Perhaps Overture Maps has used impressively accurate satellite imagery tracing to detect the demolition and rebuild of a structure somewhere in Sudan, and can output a new polygon. No OSM mapper is setting foot in Sudan, and recent satellite imagery for the area is not available through companies that share such data for OSM use.
The issue for an OSM mapper who sees the conflict between OSM (with the old building) and Overture Maps (with the new building) is they don't have any information to know which result is accurate. Is OSM just out of date? Has Overture Maps produced the result from outdated satellite imagery and OSM is more up-to-date? Is the result form Overture Maps the result of a mistake in an automated tracing algorithm?
Seems like a play at the old Microsoft "Embrace, Extend" approach. Whether or not there's an Extinguish after that is yet to be determined.
Curiously, two of the four layers use the ODbL.
Probably this data is reformatted OpenStreetMap data or derived from it (so is licensed as ODBL and requires attributing OpenStreetMap)
> Transportation: The OMF’s Transportation layer represents a worldwide road network derived from data in the OpenStreetMap project. This community-built data has been recast into the Overture data format which provides consistent segmentation of the data and a linear reference system to support additions of data such as speed limits or real-time traffic.
The OSM ODbL is crystal clear that OSM contributors have to be credited. I don't believe that CDLA Permissive v2.0 magically allows Overture to bypass it.
--EDIT: I missed that they're using different licenses per dataset, the transport theme is OBDL, which I'm sure will trip up users who are not careful.
https://cdla.dev/permissive-2-0/
My read is that they're not going to avoid credited OSM, but rather they're going to credit OSM / maintain the license for the parts they use from OSM and then the rest will be CDLA 2.0 licensed for anyone to use.
I wish there was a way for people to fund satellite imagery that got pushed into these systems after purchase. Sunnyvale, for example, paid for a lot of imagery of the city that they use/used in staff discussions about traffic, zoning, etc. It would be nice if they could then push those images into the open data set.
Though there are a lot of sources for satellite imagery right now; Google may not be that hard-up for new stuff. I suspect the commercial vendors they work with probably image the entire CONUS area every 12 months or so.
The imagery that would seem to be more in-demand would be the aerial photos used at very high zoom levels. I'm not a geospatial expert but I think these images are combined with some sort of LIDAR or multi-spectrum imagery to have height maps in tandem with the visuals. That strikes me as pretty expensive to obtain.
Overture is a data-centric map project, not a community of individual map editors. Therefore, Overture is intended to be complementary to OSM. We combine OSM with other sources to produce new open map data sets. Overture data will be available for use by the OpenStreetMap community under compatible open data licenses. Overture members are encouraged to contribute to OSM directly.
https://docs.overturemaps.org/
OSM focuses on open map editing, but due to it's flexible schema, it can be hard to extract needed information from OSM.
Overture seems to focus on providing map data (from multiple sources) that can be used more easily.
EDIT: OSM is also trying to improve the OSM data model for easier processing. https://github.com/osmlab/osm-data-model
Overture is a new group of companies releasing some datasets on open licenses, but methods used to create them remain proprietary. Some of released data is theirs, many datasets are repackaged OpenStreetMap data.
There’s a gravel path near my house that maybe sees 20 people using it daily. Due to some work done nearby, the path was partially moved a few meters to the side. OSM reflected this new reality the day after.
Could also be that OSM leverages an open data set.
If my French doesn’t deceive me, https://opendata.paris.fr/explore/dataset/les-arbres/informa... has data on 207,688 trees, tagged with species, height, and circumference.
I know this is "booo-google", but I just want to write some joins with other tables I happen to have in bigquery. I'm wondering if there is some "community" BQ rather than maintaining import of my own.
Looks like it will be a while before that can be done, seeing as it uses a custom schema.