119 karma · joined September 27, 2017
Strictly we could have titled this: the markdown database pattern for small personal databases.
The point was that for personal use e.g. keeping a list of projects, or books etc this is perfect. Or even for maintaining this list on a website.
It isn't for multiple simultaneous writes
Point on comparison with other formats is really well taken. I think the main argument i would have is ubiquity of markdown and markdown tooling.
We have this overall sense that "markdown is eating the world" in the sense that it is becoming a more and more ubiquitous and powerful format. What i would term "markdown plus" i.e. markdown plus extensions like frontmatter, wikilinks, typed backtick blocks you have something really pretty expressive - and add MDX and you have full JS(X) components.
At least, for reasonably small use cases (think under a 1000 items), this makes a collection of markdown files a great basis for a whole bunch of use cases from a document database to a PKM etc.
What we're trying to do here is set out the claims clearly (and atomically) and evaluate them in as transparent and honest a manner as possible. So far I've not seen this a lot on the internet despite the huge amount of material on this topic.
If you want to know more about motivation and approach there is quite a bit of detail on https://web3.lifeitself.us/about
Massive +1. And i agree that perhaps "reusing" the book would be useful. E.g. take the key factual approach and update analyses for particular technologies and then plug those together - which to some extent is what the pathway calculator did which is why it is worth trying to get the source for that https://github.com/life-itself/climate/issues/2
At Open Knowledge we built a really early one called opendatasearch.org in 2011/2012 - now defunct - and were involved in the first version of the pan EU open data portal. We also had the original https://ckan.net/ (and subsites) which is now https://datahub.io/ and has become much more focused on quality data and data deployment. [Disclosure: I was/am involved in many of these projects]
The challenge, as others have mentioned, is that data quality is very variable and searching for datasets is complicated (think of software as an analogy - searching for good code libraries is a bit of an art).
I imagine Google are trying this out before making datasets another "special type" of search result -- after all you can already search google for datasets. In addition, Google are already Google so including datasets will have a level of comprehensiveness and exposure you struggle with elsewhere (part of the power of monopoly in a sense!).
PS: for those looking for data gov sites https://dataportals.org/ has most of these.
The ideas there are now getting realised in Frictionless Data https://frictionlessdata.io/
This an initiative providing a simple way of "packaging" data like software plus an ecosystem of tools including a package manager etc - https://frictionlessdata.io/data-packages/.
Aims to be minimal, easy to adopt etc (e.g. based on CSV). It has got significant traction with integration and adoption into Pandas, OpenRefine etc.
https://datahub.io/ itself is entirely rebuilt around Data Packages and includes a package manager tool "data".
If you're interested to talk more please come chat on http://gitter.im/datahubio/chat
There's a wide set of tooling for Tabular Data Packages, plus underlying data is CSV which anyone can use.
If you want this done automatically you can just publish your CSV to https://datahub.io/ and your Tabular Data Package is made automatically for you.