Introducing schema.org: Search engines come together for a richer web
googleblog.blogspot.com
googleblog.blogspot.com
So, at the end of the day, Google, Microsoft, and Yahoo just made web development more expensive. They probably also just made the web a better place, too.
I'm guessing this is like salt... a little dash'll do it.
If your website employs an SEO or webstandardista, you should already have your sites marked up with metadata. Reviews and rich breadcrumbs etc. have been around and supported for years now.
I suppose that since Google wants to solve these problems algorithmically first and foremost, that many of these structures already get recognized. Right now you don't _have_ to mark-up your breadcrumbs, for them to still appear as rich snippets on your search result listing. Google recognized the structure without added mark-up.
For now, I will build the new types in my CMS like a good web developer. That won't cost me any time in the future, and now I'll have a way to separate myself from those that won't add schema's or metadata to their mark-up. So, at the end of the day, I just got more expensive :)
"Google currently supports rich snippets for people, events, reviews, products, recipes, and breadcrumb navigation, and you can use the new schema.org markup for these types, just as with our regular markup formats. Because we’re always working to expand our functionality and improve the relevance and presentation of our search results, schema.org contains many new types that Google may use in future applications."
"Google doesn’t use markup for ranking purposes at this time—but rich snippets can make your web pages appear more prominently in search results, so you may see an increase in traffic."
So, if the metadata doesn't directly increase rankings (which I'm pretty positive it won't), it can indirectly do so by grabbing the users eye and improving CTR, which I am most certain it will.
https://gist.github.com/1005688
I like this quite a bit better.
On the other hand, the RDFa way is more painful when the opening hours differ between the days of the week.
The tools aren't there yet so it's not as useful for linked data as RDFa right now, but hopefully it will be more useful soon.
Schema.org is doing it with the SEO angle: mark up your pages like this, and they'll be presented better in search engines.
With VIE (https://github.com/bergie/VIE) we take a different angle: mark up your pages with RDFa, and they'll become editable.
Most mainstream users are not tech savvy. Imagine a traveling person arrives in a city and spontaneously decides to see a movie. They enter the search 'good movie playing in nowheresville'. As it stands now their query will likely be matched by the keywords 'movie', 'playing', and 'nowheresville'. The returned results might include a news article about local theatres, with no actual focus on reviews. The searcher might get frustrated and just decide to rent a movie instead. However, with schemas in wide use search engines will know exactly what web sites are talking about movies and whether it's in the context of reviews. The searcher can then be passed on to the relevant site.
In other words, do you think it's better to tell search engines this is sort of what I have or this is exactly what I have?
This announcement is not the proposal of a new technique, but rather the extension of one which is already working and is a good thing for the web.
I am not sure if I want to provide all my hard work in a format which will maybe help the search engines a bit, but mainly the spammers a lot as they will be able to automate the creation of content farms even more.
Mixed feelings... all the world data in a well structured format is a wonder but at the same time, what will be the incentives to create such an easy to digest content if the world at the end do not even know you are the one how produced it?
Kind of the old media against new media dilemma but applied to the new media. Interesting.
I guess it could actually hurt them as well. If users aren't providing information back to the algorithm in the form of a click through related to a search term, don't the search engines also risk losing a key signal of relevance?
If I see immediately relevant data for restaurant hours, movie times, a person's bio, etc, I'm far more likely to click-through and start looking at a menu, making a reservation, or getting more background.
They may well expect increased click-throughs leading to more site traffic for those who adopt.
It takes too long to render sites with loads of images and adverts on my phone, so I always dread clicking a search result.
Information extraction, just got that much easier. Hello, baby semantic web.
if I picked 10 web devs off sitepoint and instructed them to add 10 assertions to an HTML page and didn't give them a validator, I'd be amazed if more than 3 got 80% of them right.
i like the taxonomy though, but honestly i think instances are much more interesting than types... rather than saying "George Washington" a :US_President, can't we say "George Washington" is :George_Washington where :George_Washington is his identifier in Freebase?
I guess when you are big G, you can do anything you want.
But you don't have to use microdata or microformats; you can still use this schema with RDFa on an XHTML page.
You should be able to express everything in Microdata _and_ RDFa. I don't think RDFa is semantically richer than these other formats.
The differences as I understand them are three: * schema.org has an implicit vocabulary, if you want to use more than one you can stil use RDFa and use the schema.org vocab explicitly * some syntactic hacks are missing (curies, chaining) but these do not remove expressiveness. Again, implicit schema. * typed literals are missing. And once more, not really needed when the schema is only one
I still would have preferred if they had used straight RDFa 1.1, but I think their main motivation is that the way the web is going (HTML5) does not seem to be the same it was when RDFa was initially invented (xhtml).
This solves concrete a finite set of problems now, while in the semweb world people still have to agree on how to express a person's name :/
It seems that they chose Microdata over RDFa because the latter's syntax was deemed to be unwieldy.
It's not really true that RDFa is more extensible than microdata, there are a small number of missing features related to XML data, but nothing too significant for these use cases; see, for example, [1]
[1] http://bnode.org/blog/2010/01/26/microdata-semantic-markup-f...
2. As for tooling, here are two super-easy ways of adding rich GoodRelations data to your site: - http://www.ebusiness-unibw.org/tools/grsnippetgen/ This creates a snippet of a few additional divs/spans based on your data; simply paste it before (!) the respective visible content. You are done ;) For products, see the effect in http://www.google.com/webmasters/tools/richsnippets - if you are using a standard shop package, e.g. Magento, osCommerce, Joomla/Virtuemart, or WordPress/WPEC, there are free extension modules that add GoodRelations:
http://wiki.goodrelations-vocabulary.org/Shop_extensions
A similar module for Prestashop and Oxid eSales is in the making.
Best Martin Hepp
Disclaimer: I am the inventor and lead developer of GoodRelations. GoodRelations is free to use, remix, or adapt under a Creative Commons license.
I am still not really convinced that it is possible to integrate handcoded schemas for a wide range of use cases into search results in a meaningful way.
The solution Google proposes here will also restrain the content of websites in a lot of ways if it becomes widely adopted. Look at the recipes-example, it defines markup for including nutrition information for recipes:
"Can contain the following child elements: servingSize, calories, fat, saturatedFat, unsaturatedFat, carbohydrates, sugar, fiber, protein, cholesterol"
Every company that serves recipes on the web and decides not to offer this information because it deems other properties of recipes more important is now at a disadvantage. Google will show more information about the recipes of their competitors and presumably also rank them higher because they have included 'valuable' markup information in their recipe.
This approach favors shallow information ressources over complex ones as the former can be more easily parsed by metadata-crawlers.
Chicken recipe where calories<=350 and carbs<=20g
ranking>=4 and reviews>=50
Site also looks a bit like spam. Needs more Firefox-esk awesome graphics, imo.
Let us say we both link to http://schema.org/docs/search_results.html#q=test and http://schema.org/docs/search_results.html#q=product .
Since it is a hashtag we link to, all the link juice will get consolidated in the search page, which passes it on to the rest of the site.
If we had linked to http://schema.org/docs/search_results.html?q=test and http://schema.org/docs/search_results.html?q=product we would have created two (low-quality and near-duplicate) pages in the google index.
The same principle applies to pagination. If you can do javascript pagination #page=2 vs. dynamic pagination ?page=2 you are nearly always better of with the hash pagination. If you do it right, you get the benefit of a single page, with the added bonus of being able to bookmark a certain page and having browser history working.
Now if they'd only add some schema targeted towards downloadable public data sets. I'm dying for a good global public dataset search beyond competing data markets and data.gov.* sites.
There's a lot on schema.org for social media websites/business lookup but should be more for open data. I was looking for a linked data schema to represent financial transactions (X paid Y $999 for Z) but schema.org only goes so far as Sales. XBRL explictly states it is not for "A transaction level activity".
No problem, I'll add a few kbytes to every single page of my sites, so I can replicate the information I've already stated in a number of sitemaps, video sitemaps, headers and XML files.
P.S: I'm not agaisnt standardization at all, I'm just saying, this comes a bit late.
Microformats? I'll pass.
<link rel="data" type="data/json" href="http://example.com/recipes/chicken.js" />
<link rel="data" type="data/xml" href="http://example.com/recipes/chicken.xml" />
The resource can be cached, served static or even included in the page inside a <script data> tagHere is the html:
<div itemid="1234">Chicken marsala</div>
<div itemid="1235">Fried chicken</div>
<div itemid="1236">Chicken curry</div>
Here is the data island in json: {
head:{
title:'',
source:'',
version:''
},
items:[
{
id:'1234',
type:'recipe',
title:'Chicken marsala',
ingredients:'here...'
},
{
id:'1235',
type:'recipe',
title:'Fried chicken',
ingredients:'here...'
}
]
}
Here is the data island in xml: <data>
<head>
<title>here</title>
<source></source>
<version></version>
</head>
<items>
<item id="1234" type="recipe">
<title>chicken marsala</title>
<ingredients>here...</ingredients>
</item>
<item id="1235" type="recipe">
<title>fried chicken</title>
<ingredients>here...</ingredients>
</item>
</items>
</data> <link rel="alternative" type="application/event+json" href="http://example.com/events/2010/06/03/schweet.json" />http://dev.w3.org/html5/md/#application-microdata-json
So you'd have an alternative resource
<link rel="alternative" type="application/microdata+json" href="http://dev.w3.org/html5/md/#application-microdata-json" />
That file would look like this: {
"items": [
{
"id": "http://example.com/events/2010/06/03/schweet",
"type": "http://schema.org/Event",
"properties": {
"startDate": ["2010-06-03"],
"location": [{
"id": "http://example.com/places/my-crib",
"type": "http://schmea.org/Place",
"properties": {
"url": ["http://example.com/places/my-crib"],
"address": [{
"type": "http://schema.org/PostalAddress",
"properites": {
"addressLocality": "Knoxville",
"addressRegion": "TN"
}
}]
}
}]
}
}
]
}
Who knows if Google will actually use that file though.kb/us == mb/ms == gb/s
On the other hand, I agree that a even a kilbotye of extra data to have fundamentally better search result pages is a big win for everyone.
Microformats should never invalidate any doctype as they are just class-names.
Microdata can be used in valid HTML5 doctype pages.
RDFa can be used in valid xHTML+RDFa doctype pages.
Why would I care? What matters to me is that Google gets people to my site.