Stephen Wolfram on a .data TLD
blog.stephenwolfram.com
blog.stephenwolfram.com
I think the right way is to put things under the domain. So data.google.com, or google.com/data or even a META tag on a web page that tells the browser the URL for the data relevant to a particular page.
Yep. Just call it a New Kind of Internet.
Edit: Relevant:
Edit2: Amazon removed the review! Search for "A new kind of review" on this page
As for that notion, maybe we should switch to naming it such as we do for java packages :D
Google.com would be com.google.search, com.google.mail, etc :P
"I have to say that now I regret that the syntax is so clumsy. I would like http://www.example.com/foo/bar/baz to be just written http:com/example/foo/bar/baz where the client would figure out that www.example.com existed and was the server to contact. But it is too late now."
This sounds to me like a high-level description of how the web is supposed to work today, only implemented using a new TLD instead of HTTP headers.
It sounds odd to me, coming from someone whose major web service sends all results -- even text and tables -- as GIF.
Allow me to explain.
RDF[1] was created towards solving the "data web" problem. However, the challenge has been representation and modeling "things" such that we can cross-link "data" on "web". The language to create such shared representations (Web Ontology Language[2]) is difficult to use and standardize. Nevertheless, this approach has been hugely successful in knowledge-intensive domains such as biology and health care.
On the Wild Wild Web, the microformats[3] have got wide support from Search engines and web publishers.
2. http://www.w3.org/TR/owl-features/
3. http://support.google.com/webmasters/bin/answer.py?hl=en&...
The index page will give you all the discoverability and from there you can go to google.com/employees or bestbuy.com/products etc showing whatever data is public (or private provided oauth mechanisms) and what can be created,modified and deleted according to roles and security levels.
This has been tried before but the well was poisoned when they dropped SOAP in it.
I certainly hope so; linked data (of which RDF is the main implementation nowadays) is much more useful than having disconnected silos. Of course, we won't transfer HTML, but that's just one implementation of hypertext.
Besides, even if we weren't, why would we replace HTTP by something that accomplishes the same? Doesn't make much sense to me.
This seems more intended for bulk data, which is likely going to be some pregenerated chunk in the MB, GB, TB range and so less suited to the JSON API call paradigm, and more likely to involve a simple lookup to disk rather than being computed on the fly from some database.
I'd say thats a pretty good reason for using the new TLD, technicalities aside.
Who would be the standards body for defining and regulating such a uniform mechanism?
It’s just a namespace, one of many possible choices. But I wouldn’t discount its importance as a protocol, or an expectation. “.com” has a very important non-technical meaning.
A big problem is how to ETL these datasets between organizations, and I think Hadoop is a key technology there. It provides the integration point for both slurping the data out of internal databases, and transforming it into consumable form. It also allows for bringing the computations to the data, which is the only practical thing to do with truly big data.
Currently there are no solutions for transferring data between different organizations’ hadoop installations. So some publishing technology that would connect hadoop’s HDFS to the .data domain would be a powerful way for forward-thinking organizations to participate.
Another path towards making things easier is to focus on the cloud aspect. Transferring terabytes of data is non-trivial. But if the data is published to a cloud provider, others can access it without having to create their own copy, and it can be computed upon within the high-speed internal network of the provider. Again, bringing the computation to the data.
> [Hadoop] provides the integration point for both slurping the data out of internal databases, and transforming it into consumable form
Hadoop does no such thing. It doesn't "slurp data out of internal databases". It's just a DFS coupled with a MapReduce implementation. Perhaps you're thinking of Hive?
> Currently there are no solutions for transferring data between different organizations’ hadoop installations.
All data isn't "big data". By being myopically hadoop-focused, you're ignoring the real problem, which is data interchange. XML was supposed to be the golden standard; it's debatable how far it's achieved its initial goal.
> So some publishing technology that would connect hadoop’s HDFS to the .data domain
So basically, forsake all internal business logic, access control, and just pipe your database to the net? When you have a hammer...
> Transferring terabytes of data is non-trivial. But if the data is published to a cloud provider, others can access it without having to create their own copy, and it can be computed upon within the high-speed internal network of the provider
See AWS public datasets for exactly this, but it's still a long shot. It also ignores the problem of data freshness (i.e., once a provider uploads a dataset, they also need to keep updating it). http://aws.amazon.com/publicdatasets/
There is a reason XML, the semantic web, linked data failed to really change the data world, whereas hadoop did. The reason is computation.
The problem isn't data interchange formats and ideal representations, the problem is being able to compute with data. Distributed computation can then be used to solve all the other problems.
Case in point: Slurping data out of databases. Apache Sqoop leverages the primitives provided by Hadoop, in terms of partitioning and fault tolerance, to make it easier to do massive data transfers out of existing databases.
Another example of a solution coming from the hadoop perspective: Avro. It beats the pants of off XML as a data interchange format, precisely because it makes computing with the data (which is the ultimate point) easier.
Now, there is a reason I called Hadoop the integration point. It is becoming a general purpose computation system, which at the same time is also the datawarehouse for organizational data. So rather than dealing with the details of proprietary commercial systems, programmers can target applications to the open-source hadoop ecosystem, and have those solutions be reusable and customizable on a large scale.
The "publishing solution" would of course deal with access control, business logic, freshness, etc. That is exactly what I'm advocating be built.
Individual pieces of data may not be big data, but the aggregate problem still is. In fact this is exactly the Wolfram Alpha case: tons and tons of little datasets that add up to a lot of headache.
RDF and ontologies are just more data. Without computation, that data is not useful, and all the things one "could do" with it will not come to pass without a credible computational platform that people actually want to use.
So IMHO I would like to see that community focus less on standards and ontologies and RDF-as-panacea, and and more on the infrastructure needed to put the data to work.
Are people willing to make micropayments for access?