Tech giants let the Web's metadata schemas and infrastructure languish
threadreaderapp.com
threadreaderapp.com
I know that hating on Google is fashionable, but that's a bit too much editorializing. Especially considering the content of the post, and Google just being a small side note.
---
On-topic: I recently looked into using schema.org types as the basis for a information capturing system, but many of the types are somewhat outdated, of questionable quality or just missing. Development indeed seems slow, while changes that are needed by one of the larger involved companies get pushed through quickly.
I think a big part of that stagnation is a lack of interest though. The whole semantic web domain has been pretty much inactive.
It's a real shame: having canonical types for most things in existence, and have those actually be supported as import/export formats or for cross-app integrations, would be immensely valuable! But there is absolutely no business incentive there - rather the opposite. Easy portability of data is not something most companies would want.
That depends on what kind of data it is. For example, your home address is not part of your bank's primary business model, but keeping it up-to-date is important for it. If data portability in and out of the bank makes it more likely that you'll keep it up-to-date, that's useful for your bank as well.
Legislation and customer demand is also making it more and more palatable. If some data is not critical to your business model, but being the sole guardian of it is a legal/reputational liability is, then actually handing control over that data over to someone else and re-using that is very useful.
If the bank was the one owning the information they would not want it to be shared with others as that would allow their client to easily migrate to another bank which they definitely do not want.
But as the one receiving the data,sure, it would be nice to have others share it with me, they'd say.
I'm afraid without legislation data sharing is never going to be a thing.
(Note that the bank is an example - it could be another party.)
Unfortunately it's not allowed by law that a consumer gives "push access" to e.g. banks, health insurance or employers.
The typical solutions to collective action problems are (1) benefactors who subsidise production (either privately or through taxation), or (2) direct command and control. Google was apparently filling the role of benefactor.
I could certainly imagine that for many companies the disadvantages of using a model that's not simply a copy of their specific view of that problem domain are larger than the hypothetical benefits of interoperability, so even if such a shared ontology would exist, many would intentionally choose to use their own ontology instead of adapting to that standard.
Even inside a single company I get scared now when someone says "if only we had just one standard way to handle this sort of thing"... if it's rarely simple for just one company, how would it work globally?
Without substantial benefits of a large universal ontology and without the ability to painlessly diverge from an ontology to not compromise on accurately modeling a particular domain, it’s hard to see the net benefits. Everyone will want to customize things for their domain, or point of view. An ontology should be easy to fork, like a repo.
I would certainly want to extend, modify and replace the data models for my core business as I see fit. But beyond those there are still going to be a whole lot of models in need in order to run the company but want to keep low maintenance. E.g. for hiring I might not have strong opinions what a job post, a candidate or an application should look like, so I'd be happy sticking to the standard in those cases and benefit of it being easier to mix and match tooling and pass around the data.
Also I reckon to me that a partially customized ontology, which is inevitable, is still easier to map between orgs than if they build it from scratch completely
Or maybe see it less as a standardized ontology but as a standardized way to create ontologies
To the extent that other uses can basically piggyback on data that sites added to target Google, it does provide some value, but I don't see it as really even attempting to be a generally useful "semantic web" or linked data vocabulary in the sense of interoperating with other things.
Whichever company did that would be accused of trying to "take over" the web.
Ideally large companies should be sponsoring open efforts to define things that affect how the web works rather than doing the work themselves. Smaller open teams that move fast to define structures that work for as many people as possible, even if they're not perfect for Google, Microsoft, etc, would be more useful to the internet industry as a whole.
i also went through microformats, which seems to be much smaller, and more tightly-focused around blogs and structuring data shared among federated sites.
Submitted title was 'Google is happy to control core Web schemas, but they neglect project'
It’s the latter I think is clearly valuable, in order for us to have competition for the likes of google and Facebook. It lowers the barrier for creating competing search engines, modern rss readers, and even things like distributed social networks.
It grew out of the semantic web community so this was roughly what I expected. That space just seems cursed to have these lofty ideals which are never realized because it’s hard to justify spending time on something which has no known consumer. Schema.org seemed poised to change that but they only use a couple of types and then only for a few types of searches.
[1] triply.cc
Starting a new project that garners widespread attention looks good in a package, but replacing lightbulbs and scrubbing floors doesn't. Folks create a splash, get promoted, then move on and are not replaced.
I've never worked at Google, so I do not know if this dynamic is real. I would be interested to hear from Googlers about incentives to work or not work on something.
But that leaves me wondering how the linked situation occurs. It's a cliché that Google shutters projects or loses interest. I don't know whether that's particular to google or if it's due to an availability heuristic (Google is well-known, so its wanderings-off are widely publicised).
But if there are such dynamics and incentives, they are worthy of attention. Google exerts enormous gravity on the fabric of the technology industry, it would be helpful to avoid hurtful externalities arising from internal incentives.
The rest of us have to make do with leaky canoes, going up a certain creek, often sans a paddle.
It was clearly designed by bureaucrats who enjoy making rules and sub-rules and sub-sub-rules. It doesn't matter if it works, or is useful, as long as there are plenty of rules.
There's a reason nobody wants to play with the jerk dungeonmaster.
Maybe we can finally stop using ontologies for the semantic web and start solving the hard problem of language pragmatics.
Ideally it'd be nice (imo) if schema.org had more domain specific extensions, similar to the bib[0] one which allows for things like comic book properties to be described.
They don't care to address any of the issues or "fix the infrastructure" because this isn't a "organize all the information in the world!" project at all. The guys that take Google visitor retention stats into their next performance meeting are probably poking fun at all the ontology nerds that have descended on their metric-driven scheme.
> Schema.org is a collaborative, community activity with a mission to create, maintain, and promote schemas for structured data on the Internet, on web pages, in email messages, and beyond.
> A shared vocabulary makes it easier for webmasters and developers to decide on a schema and get the maximum benefit for their efforts. It is in this spirit that the founders, together with the larger community have come together - to provide a shared collection of schemas.
If this isn't an "organize all the information in the world" project, then Google and the other companies involved are branding it in a horribly dishonest way. In which case, they should be criticized for presenting a company-specific visitor retention strategy like it's some kind of altruistic gift to the world.
Sites like Facebook and Twitter have their own 'lite' metadata schemas that they use to help identify and render links. Hardly anyone criticizes them over it, because they haven't registered a generic domain like 'schema.org' and presented their work like it's some kind of community-driven collaboration. They're upfront that it's just a simple API for their website.
Care to elaborate further as to which of the schema that your kept/found most useful?
And you can generate it dynamically: https://developers.google.com/search/docs/guides/generate-st...
JSON-LD is far easier to work with. Having to mix metadata with markup was a pain. Half the required elements would have to be hidden in CSS because they didn't make sense in context.
The only way around that is for somebody to do the processing of the real data to validate that it isn't just bullshit for a nefarious purpose. From what I've heard about the Semantic web conceptually seems a bit skeumorphic as a concept.
It's missing the T, as-in: infrastrucTure
Would forcing the proposer to quantify costs and benefits help?
Wow! Nobody else does anything to collaboratively, inclusively develop schema and the problem is that search engines aren't just doing it for us?
1) Search engines do not owe us anything. They are not obligated to dominate us or the schema that we may voluntarily decide to include on our pages.
We've paid them nothing. They have no contract for service or agreement with us which compels them to please us or contribute greater resources to an open standard that hundreds of people are contributing to.
2) You people don't know anything about linked data and structured data.
Here's a list of schema: https://lov.linkeddata.es/dataset/lov/ .
Here's the Linked Open Data Cloud: https://lod-cloud.net/
Does your or this publisher's domain include any linked data?
Does this article include any linked data?
Do data quality issues pervade promising, comparatively-expensive, redundant approaches to natural-language comprehension, reasoning, and summarization?
Here, in contributing this example PR adding RDFa to the codeforantarctica web page, I probably made a mistake. https://github.com/CodeForAntarctica/codeforantarctica.githu... . Can you spot the mistake?
There should have been review.
https://schema.org/ClaimReview, W3C Verifiable Claims / Credentials, ld-signatures, and lds-merkleproof2017.
Which brings us to reification, truth values, property graphs, and the new RDF* and SPARQL* and JSON-LD* (which don't yet have repos with ongoing issues to tend to).
3) Get to work. This article does nothing to teach people how to contribute to slow, collaborative schema standards work.
Here's the link to the GitHub Issues so that you can contribute to schema.org: https://github.com/schemaorg/schemaorg
...
"Standards should be better and they should pay for it"
Who are the major contributors to the (W3C) open standard in question?
Is telling them to put up more money or step down going to result in getting what we want? Why or why not?
Who would merge PRs and close issues?
Have you misunderstood the scope of the project? What do the editors of the schema feel in regards to more specific domain vocabularies? Is it feasible or even advisable to attempt to out-schema domain experts who know how to develop and revise an ontology or even just a vocabulary with Protegé?
To give you a sense of how much work goes into creating a few classes and properties defined with RDFS in RDFa in HTML: here's the https://schema.org/Course , https://schema.org/CourseInstance , and https://schema.org/EducationEvent issue: https://github.com/schemaorg/schemaorg/issues/195
Can you find the link to the Use Cases wiki (which was the real work)? What strategy did you use to find it?
...
"Well, Google just does what's good for Google."
Are you arguing that Google.org should make charitable contributions to this project? Is that an advisable or effective way to influence a W3C open standard (where conflicts of interest by people just donating time are disclosed)?
Anyone can use something like extruct or OSDS to extract RDFa, Microdata, and/or JSON-LD from a page.
Everyone can include structured data and linked data in their pages.
There are surveys quantifying how many people have included which types in their pages. Some of that data is included on schema.org types pages.
...
Some written interview questions:
> Which issues have you contributed to? Which issues have you seen all the way to closed? Have you contributed a pull request to the project? Have you published linked data? What is the URL to the docs which explain how to contribute resources? How would you improve them?
https://twitter.com/westurner/status/1291903926007209984
...
After all that's happened here, I think Dan (who built FOAF, which all profitable companies could use instead of https://schema.org/Person ) deserves a week off to add more linked data to the internet now please.
If you or your organization can justify contributing one or more people at full or part time due to ROI or goodwill, by all means start sending Pull Requests and/or commenting on Issues.
"Give us more for free or step down". Wow. What PRs have you contributed to justify such demands?
https://schema.org/docs/documents.html links to the releases.
Isn't this SOP for Google?