I submit this link after coming across this site while Googling for info on parsing wikipedia "infoboxes". I plan to check out their "Article Structured Contents (BETA)" API. Improving infoboxes to be machine-readable seems important. And it would be bad if didn't do this because it's a revenue stream for them.
https://m.wikidata.org/wiki/Wikidata:Main_Page
> Wikidata is a free and open knowledge base that can be read and edited by both humans and machines.
> Wikidata acts as central storage for the structured data of its Wikimedia sister projects including Wikipedia, Wikivoyage, Wiktionary, Wikisource, and others.
Its actually quite cool. If you have never played with it i encourage checking out some of the example queries (there is a button labelled examples) on https://query.wikidata.org/
Pulling it into a mediawiki extension (or core) and making it part of the page-level metadata gets suggested pretty frequently, but it's a bit contentious amongst the hardcore editors who'd need to actually adopt such things. The template-based nature of the current infoboxes mean that they're very accessible to the community, and it's easy to spin off new variants or make changes without getting programmers to help you.
There's slow movement towards getting the sort of data that winds up in infoboxes into wikidata, but it's still somewhat spotty.
(If you've never done it, it can be quite edifying to install mediawiki for yourself and seeing how much of the surrounding infrastructure of wikipedia is absent because it's all templates.)
I'm sure we'll be seeing more lawsuits (and likely new regulation and licenses) around this.
Are you saying LLMs are not derivatives of their source material? Why not?
At some point, maybe poorly-trained chatbots that consistently produce what's seen as avoidably/negligantly poor results may become regulated. Like how if a company poorly trains its employees, it is on the hook for their employees' behavior.
It's worth noting that https://creativecommons.org/faq/#artificial-intelligence-and... itself takes the general stance that "as a general matter text and data mining in the United States is considered a fair use and does not require permission under copyright."
But as a practical matter, I wouldn't be surprised if some Wikipedia editors balk at their volunteer work being actively marketed and reformatted for ease of LLM training by the very platform that solicited their volunteer services, regardless of their works' legal status and Wikimedia's technical respect of that legal status.
As someone who avidly edited Wikipedia for 6-8 years, I am happy to see my volunteer work used for LLM training. I also agree some other editors likely aren't.
Redistribution of content is an entirely different matter, and the legal status of copyrighted material in relation to LLM training is an open issue that is currently the subject of litigation.
> "it is important to note that Creative Commons licenses allow for free reproduction and reuse, so AI programs like ChatGPT might copy text from a Wikipedia article or an image from Wikimedia Commons. However, it is not clear yet whether massively copying content from these sources may result in a violation of the Creative Commons license if attribution is not granted. Overall, it is more likely than not if current precedent holds that training systems on copyrighted data will be covered by fair use in the United States, but there is significant uncertainty at time of writing."
The new Wikimedia Enterprise APIs facilitate attribution. For example, the "api.enterprise.wikimedia.com/v2/structured-contents/{name}" response [2] includes an "editor" object in a "version" object. So the Wikipedia editor who most recently edited the article seems quite feasible to attribute. ML apps could incorporate such attribution in their offering, and help satisfy the "BY" clause in the underlying CC-BY-SA 4.0 license for Wikipedia content.
---
1. https://meta.wikimedia.org/wiki/Wikilegal/Copyright_Analysis...
2. https://enterprise.wikimedia.com/docs/on-demand/#article-str...
I think that will heavily depend on just what the money goes to.
A better user experience, tightening up the code behind things, fewer nag screens for donations? Justifiable.
Jimmy Wales going from decently-compensated to a bad case of Founder Syndrome? Not so justifiable.
Founder Syndrome: a psychological condition wherein a person who starts a venture believes the venture should earn them a net worth on the order of billions of dollars, regardless of its actual economic value, and is willing to seriously enshittify the venture's product or service in order to make this delusion come true, especially in preparation for an IPO. See also: /u/spez, SPAC, Facebook, Unity Engine
Jimmy wales is not paid at all (he has a board seat, but that doesn't come with any money)
He of course has leveraged his fame from being "founder" quite extensively. I think most of his money comes from fandom which he is also one of the founders of.
Isn't the point of wikidata to offer machine-parseable formats?
Also: how can those formats be locked into a paywall?
wikipedia infoboxes API is already paywalled, pricing starts from $0.01/req.
And that's the problem with making money off convenience. If the free thing becomes better, then it cannibalizes the paid stuff. So if you take this approach to funding, you will want the free tier to be perpetually bad as much as possible.
MediaWiki
Wikibooks
Wikidata
Wikifunctions
Wikimedia Commons
Wikinews
Wikiquote
Wikisource
Wikispecies
Wikiversity
Wikivoyage
Wiktionary
https://en.wikipedia.org/wiki/Wikimedia_FoundationBut quality does take money.
Curious why you're running a mirror for OSM, is it public or for your private use? If you don't mind, I'd love to learn more about it; sounds super interesting to me.
Don't get me wrong, it is wonderful that the Wikipedia team offers this, and I am grateful they give anything at all for offline usage. It just feels like it's intended more as a side product of their backup process, rather than something you're really supposed to use.