It uses the Google dataset, which is under the public domain. Might make sense to compare formats?
Btw, I love how your worldwide.yaml, both the deduplication and the redirection for subterritories (such as Vatican City).
It uses the Google dataset, which is under the public domain. Might make sense to compare formats?
Btw, I love how your worldwide.yaml, both the deduplication and the redirection for subterritories (such as Vatican City).
And thanks @bojanz for asking Google about their data's license! :)
Haven't looked in details at your addressing PHP module or even Libpostal, but I feel like there should be some ways to deduplicate efforts and converge all datasets. Both for testing and i18n/l10n.
In the mean time, OpenCageData's address-formatting language-neutral YAML structure seems quite nice.
The big challenge I think is the conflicting use cases between "official" postal format of a country and trying to represent an address in a way that makes sense to users - especially when you only have limited data available (for example when using a datasource like OpenStreetMap where you are at the whim of what the mapper decided to add). Our project isn't about forming perfect postal addresses for things like printing labels and such, it's about taking the real world data in OSM and making it look reasonable. As an example one of the next things I want to add is basic rules about postal codes so we can catch garbage that comes in when mappers mistakenly put the town name in the post code tag and such.
Will definitely take you up on the offer of comparing formats. Any further feedback you have would be really useful, you've obviously spent a lot of time thinking about this space.