Open-source geo is really something right now
trackchanges.postlight.com
trackchanges.postlight.com
Clicking around you'll find repeated references to Leibniz and Descartes dreaming of a way to capture all of human knowledge via axiomatic statements, lemmas, and conclusions; then easily dismissed as impractical.
This article talks about the Semantic Web, which is then immediately made fun of with derisive references to Rudolf Carnap and the urge to "Pokemonize" all of human knowledge.
What I want to know is, even if it is impossible to capture all of human wisdom/conclusions, why does that by definition mean the effort is worthless? After all, when looking at information rather than conclusions - wikipedia isn't complete, will never be complete, but it is still useful for what is there.
Data is messy. The only thing you can do with it is preserve the original format that was used to record it, transform a copy to a format that's useful for the analysis you want to do, and join it to other, similarly messy data.
Wikipedia is a perfect example. It's messy as hell. It's blobs of data with connections between them. Most of the blobs are unstructured text (or at best semi-structured) but there are a bunch of images in various formats there, and links to outside datasources in myriad formats and locations. There are some standards and attempts to maintain order, but the whole thing is damn difficult to process by machine. It's a very, very long way from the Semantic Web.
Storing the information isn't the real challenge, structuring the connections between different pieces of knowledge is what seems to be the very difficult part.
Without the latter, you can have all the data in the world, but no way to make any use or sense of it.
Let's start with your statement: "If it can store and retrieve data, it's a database." One of the notable characteristics of human memory is that it is not reliable. It does NOT reliably retrieve the data that it stored -- and I don't just mean that we forget some things, I mean that many of our memories are factually incorrect.
Human memory is extraordinarily USEFUL, but in order for the term "database" to have a useful meaning, I have to categorize human memory as a form of information storage that is NOT a "database".
That matches what happened, the OpenAddresses project mostly put all the messy data in one place, which made it possible to extract the parcels from the mess. Tracking down the data provided a lot more value than pontificating about how parcel data should best be stored.
Paul Ford himself doesn't claim it's a "worthless" exercise in this piece. In fact, he says "It’s pretty exciting to imagine that one day we’ll stumble into the one true universal database of all human knowledge."
As he portrays it, the vision of the semantic web is a kind of receding horizon that drives forward a lot of worthwhile efforts. That's why he brings it up in the first place. Sure, he pokes fun at it a bit, but that's just Paul Ford's (highly effective and enjoyable, IMO) writing style.