I have a similar, "One day I will do it" project. The idea is that somebody cares enough to make the semantic web work--it's just not the people with write access to the data.
I think we can use CTPH algorithms to fingerprint the data independently of whatever names are used, and then we can use that to find representations of the "same data" submitted by other users. Probably there would be some reputation stuff involved, a web of trust, etc. The flow would go something like this:
COGNIZE:
0. Encounter messy data in the wild (has pagination, timestamps of access, etc), need other representation (human/computer/whatever)
1. Calculate CTPH fingerprints, use them to search for link: miss
3. Clean data the hard way and publish canonical representation (ipfs?)
4. Generate missing representation the hard way, publish that too
5. Calculate the fingerprints common to both the missing representation and the canonical representation, and publish it as a "link" between the two. Unlike traditional web links, this one is bidirectional.
RECOGNIZE
1. Different user encounters "same" data in the wild
2. Calculate fingerprints, use them to search for canonical representation: hit
3. Find further links to see what other representations of the "same data" are available, download them if desired.
The fingerprint stuff works, but there's a lot of work left to be done re: mapping fuzzy hashes of "in the wild" data to cryptographic ones of "canonical representations" and finding ways to incentivize users to go through the hassle of the "cognize" step so that other users can benefit from the "recognize" step.
Sorry for talking your ear off, it just feels good to know I'm not the only one working on something like this, even if our approaches are quite different (mine works on PDF's only because it works on arbitrary bytes). Good luck with yours.