Python Tools for the Semantic Web, an Overview
github.com
github.com
Semantic Web was going to be a mechanism to distribute information in a reusable way, but its rise was poorly timed with the emergence of the platforms. If it had come into full swing five to ten years earlier - basically the start of the web - we'd be using the tech right now. It's still the right mindset to get away from Facebook, Google, Reddit, Twitter, etc. It has the potential to give us disproportionate power against centralization.
HTML is an unstructured soup, and extracting information from it requires human scraping or sophisticated models. It's also a sort of shifting sands subject to change. Semantic web presented a methodology for sharing data and distributing it in a variety of serializations that you could also embed into anything. It was an API for data. Not just loosely structured bespoke JSON, either.
There's no reason this comment should live on HN or Reddit or whatever. It could be shared over a form of HTTP federated aggregation or directly via P2P with lots of semantic metadata (title, author, subject, date, contents, etc.) and interest graph markup. It could also be further enriched with algorithms that summarize, extract sentiment, keywords, and so forth - but used directly for our own time saving purposes instead of fueling an algorithmic news feed that optimizes for advertisers.
People you like would have signed profiles that tell you where they publish so that you can subscribe. You can follow their interest graph but use your own algorithm to filter and rank content.
Articles, photos, calendars, recipes -- everything would be semantic, sharable, and interoperable. You could even save the content you like to your personal knowledge base since it's right at your fingertips, richly annotated with no cruft.
Facebook just took the wind out of the sails for a decade and a half. It looks like a lot of people are starting to think about this problem again.
[1] https://news.ycombinator.com/item?id=29898810
[2] https://news.ycombinator.com/item?id=24584488
Solving custody and permissions issues in a scalable way hasn't really been doable here, not in a way that isn't massively out-performed by The Platforms. I believe the semantic web makes sense for this, but people have demonstrated that privacy, and positive vs negative externalities concerning data, aren't as important as core functionality.
Another thing to consider is that content is fairly commodity compared to being apart of the reaction to that content. Configuring your own web to allow for this is a lot more work - the Facebooks curate this for you, as a matter of engagement.
I love your comment though, this is where my thoughts have been for a while now. If people could serve, or have served, their own interest graph and activity, that would be the basis for a lot of interesting things, including qualifying organic activity, and real sentiment.
Particularly people saw vocabularies such as FOAF and thought this was an effort to force everybody to publish things in a defined format and they felt the W3C was trying to put them in a straightjacket.
If you read not just the OWL and RDF specs but also the literature that contextualizes them you see the goal of the semantic web is just the opposite. Hypothetically you should be able to take data from multiple sources and various vocabularies, throw them all into one big graph (really, with namespaces this 'just works'!) and then do OWL and RDF inference to infer new facts in a desired vocabulary.
What I learned from 10 years of work around the semweb is how to build a very sloppy system and apply 'data lake' techniques to full scan it and refine a process of creating data extracts that completely fit the requirements for things I want to do with data
https://ontology2.com/the-book/data-lakes-ponds-and-droplets...
The trouble is that the standards like OWL, RDF and SPARQL are nowhere near sufficient for this task. OWL and RDF, for instance, have no answer for the problem that some people might publish temperatures in Fahrenheit and others in Centigrade and you need rules that can do a little bit of math. SPARQL really needs path variables (some predicate path matched you need to know which one.) SPARQL really needs fluent handling of ordered collections the way some object-relational SQL dialects (PostgreSQL) do.
Unfortunately most of the people who have been working on the semantic web are sick and tired of the standards process and too burned out to move these things forward.
In the end you have to process data to use it in some particular way and you have to conform to requirements. You can't count on it being clean at the beginning you need a repeatable and scalable process to do it. SHACL is much closer to what it is needed. If people had some insight they would have developed it first and then built something like RDFS as a data integration language better. Unfortunately they just adopted some ideas from the description logic world (forms of logic that are highly scalable) but were seduced by the implementability of the logic as opposed to "is it sufficient?"
Absolutely right.
> throw them all into one big graph [...] and then do OWL and RDF inference to infer new facts in a desired vocabulary. [...] What I learned from 10 years of work around the semweb is how to build a very sloppy system and apply 'data lake' techniques to full scan it and refine a process of creating data extracts that completely fit the requirements for things I want to do with data
> You can't count on it being clean at the beginning you need a repeatable and scalable process to do it.
Just to paint the picture for others, this is the "hard" vision for Semantic Web, where all information arises from disparate sources and goes into a giant queryable knowledge graph. This vision has a ton of value to certain users (it's very enterprisey), but 90% of us would benefit tremendously from only the most simple tools and ontologies.
If we focus on ontologies around blogging, threaded conversations, identity, and chat, we'd see immediate use and potential. Grow from simple beginnings.
I don't think most people understood it at the time, but schema.org, as it is deployed on the web, is adding some predicate and type structure to text that isn't useful at all unless it is paired with an entity matching system which was beyond state of the art at this time.
Now neural networks for text are nowhere near as successful at they are for images and audio but it's clear to me today that such a system can be built and that you could even use schema.org metadata to bootstrap the training set for information extraction systems that don't need the metadata.
As it was I saw schema.org as much more an effort for "big web" to take as opposed to give.
You can obviously provide entity disambiguation as part of schema.org-formatted data - it's very much part of the spec, for any "entity" type. In most cases, just linking to a well-known description of the entity (such as a wikidata.org/entity/Qnnnn URI if it's generally notable) is enough. The same works to disambiguate named entities as part of running text.
I mean, bully for you - but from interviewing data engineers for the past 5 years what virtually everyone else seems to have learned is "we have to frontload every feature with weeks/months of ontological discussion and then I'm lost because I never actually had to process any data quickly, let alone deal with dirty raw stuff."
As a person who was skeptical from day ~1 of the semweb vision, everything I see around it still feels like an "I told you so" but in a really sad way.
Without canonical names/homes for data, much of this dream feels out of reach. And with it, we are just back to citing data in a textual way. That is, we have it already.
And the canonical step is hampered by the cost of storing information. Publications largely push this to archival services. But nobody is going to want to archive every random page. Especially for the comments, which are often mutable.
I similarly don't see the advantage of requiring so much semantic marking of textual data. Indeed, I'd go so far as to say that grammar is largely superfluous. Exceedingly so if you are into the esoteric rules of the game. Certainly structure is advantageous, but most structure is inferred. And anything that slows the publishing of data is likely as not to just cause it to get lost.
For instance, I like to hang out in math twitter. I wish I was taking classes from these professors, but instead I just think about their tweets and ask myself: "that's a neat claim, could I prove it?"
Twitter isn't the place for getting feedback on proofs, and I don't really want to hijack the thread with an external link, so I just do it on paper and throw it away. But what would be cooler is if I could create a link between my proof and the conversation that inspired it. Maybe some other aspiring student will notice it and start a dialog with me where we can learn from each other.
Sure, I could post a link to overleaf or some other place that supports LaTeX collaboration but now anybody who wants to collaborate must go create an identity on that platform, and now we have the same problem all over again if somebody else, for instance, wants to play around with it via proof assistant software.
You get this proliferation of third parties who have to manage a permissions mess re: who can read/append/edit what and how to render it to users with different preferences--to say nothing of alterior motives on the part of the platform.
Ephemeral is good enough for most and a paradigm shift here might do a lot of good to address an often unspoken problem. Having services keep hold of everything for all time is simply unsustainable. While volumes of data might have future value in creating ML models, the vast majority of it will never be accessed again.
Things that deserve preservation, that have interested stakeholder(s), will be retained.
I worked on ad verification for a while and (in two radically different models!) a fuller set of `<meta>` tags was a weak pro-spam signal.