Introducing the Wikimedia Enterprise API
diff.wikimedia.org
diff.wikimedia.org
Wikipedia as the product of a public good foundation should be just that; by the public, for the public, and accessible to the public (including all access methods and API's).
I think this move is great. I'd rather have this money go to the foundation, than ParseAPIco.
Of course anyone can use it commercially and for whichever reason they like. But by creating a walled 'premium' offering you are going against the premise of 'anyone can use it' since not anyone can use the commercial api.
Surely if Google or anyone wants the api so badly they're willing to pay for it, they should fund it. Why does that mean it needs to be 'locked'?
No doubt the enterprise API will add attractive value for smaller companies without the resources to process the raw dumps but I’m skeptical that this will convert Google et al into well-paying customers. Unless they start restricting the free dumps...
They are both users who are paying for the service.
The difference is that funders have political power over the project beyond the scope of their own use.
Customers are always better and more fair to everyone than funders.
Creating an enterprise product with enterprise funding means Wikimedia is directing energy and resources to serve enterprise alone. These APIs were likely demanded by larger tech firms, and I imagine they discussed with Wikimedia the possibility of creating a commercial service. Now they got what they wanted, it won't be a public resource, and it's design won't be guided by the public either.
Now the lines are very blurry. A wiser strategy would have been to create a separate private entity.
Unless I am mistaken, seems like they did do that?
"the foundation’s newly created subsidiary, Wikimedia LLC"
https://www.wired.com/story/wikipedia-finally-asking-big-tec...
Not sure what you have to worry about.
>From op: It's just too often that a commercial offering just disincentives bettering the 'free' product.
Instead of the public foundation focusing on bettering and modernizing the public option, they are -and will continue to be - putting their efforts to create a better api to be used solely for private purposes (albeit with good intentions).
Sure, it's opt-in to use the 'better' api. But that comes at the expense of this better api not being available publicly. So going forward by default the public is always going to be getting the inferior api, saving the 'better' ones for commercial use. That doesn't seem very opt-in-like to me...
https://en.wikipedia.org/wiki/Wikipedia:Wikipedians#Number_o...
https://en.wikipedia.org/wiki/Wikipedia:Fundraising_statisti...
https://en.wikipedia.org/wiki/User:Guy_Macon/Wikipedia_has_C...
I am pretty sure that they would welcome comments from Hacker News users on that page.
Maybe they could alternate their fundraising banners with "contribute to an article" campaigns?
I own multiple sites where I and my users work to produce valuable data (e.g “so so company reviews”, “Is tenet on Disney” and other data of that kind). And what does Google do? Scrap it all and display it on their page. As a result, the page links gets millions of impressions but tens of clicks. Thus, the sites cannot be monetized. Any reasonable person knows this can’t go on for long before the free and open web comes crashing down or Google (and others like it) pays its due.
If Google scraping your sites is a good thing, then why are you complaining?
I hope Google never starts paying for the links. Once there is a precedent, this becomes an effective blocker for the new search engines, visualizers, and other exciting web search startups. A new search engine startup is not going to be able to establish a commercial relationship with every site on the web like Google could.
[0] https://developers.google.com/search/docs/advanced/appearanc...
That being said, of all the sources, Wikipedia actively license their content in such a way that google are well within their rights to slurp it all down and serve it however they want.
Google is already effectively paying for links to news sites as part of the negotiations in Australia. And I agree that this will be a dampener on any competition, I think the era of "ask for forgiveness, rather then permission" needs to stop.
if you want to specifically exclude one entity from accessing information that you've posted for anybody to see, i'm not sure how there's a way that could be "opt-in"
Does this mean that you think there should be less competition for Google?
And when you sign up for netflix or cable tv, there is an agreement you accept that you are not going to pirate.
Remember, the nosnippet does not have to be on every page -- you can put into robots.txt or HTTP header, so it is literally 1 line of configuration for most web servers.
Movie producers can only dream of stopping piracy that easily.
Oh I'm sorry I don't have the ability to look for that, my system is only equipped to look for that specific string.
> And when you sign up for netflix or cable tv, there is an agreement you accept that you are not going to pirate.
Again my system doesn't read the TOS, does Googles?
> Remember, the nosnippet does not have to be on every page -- you can put into robots.txt or HTTP header, so it is literally 1 line of configuration for most web servers.
Remember they just have to add the string "nosteal" to the opening credits. That's a few minutes in final cut pro.
Also, if they forgot to add it or have some other issue I offer no public facing customer service whatsoever.
DVDs have technological protection as well -- the CSS[0] system. So yes, if you don't want your movie to be pirated you need to explicitly enable this. This was probably harder than creating robots.txt too, there were NDAs and stuff involved.
The netflix requires logging in to access the content. If you add the same requirement, then Google is not going to take your snippets.
Unlike the string "nosteal", the robots.txt file is not Google invention, it is as much part of the web standards as all other technologies.
If you want a website, you need a server which can support HTTP, HTML, CSS, links, robots.txt and so on. You can omit parts you don't need, but then you _may_ suffer the consequences -- without CSS your site will be ugly, and without robots.txt your site will be scraped by Google.
Now we are talking specifics! Are you implying that Google is violating the law? Given that the snippet showing has been going for a long time and no one has sued Google for it yet, it does not seem to. Plus, there is the whole Fair Use laws [0].
I personally love that I can take snippets from the random websites on the net, quote them in my posts, and not worry about copyright infringement. And if I can do this, why can't Google?
[0] https://ammori.org/2012/05/08/copyright-misunderstandings-an...
So if I search for e.g. "specific breakdown of something something, in a unique breakdown format that only this website has", then the website owner has worked on, created unique/copyrighted material, and posted it on a page on their site, and Google just extracts that piece, then they might as well have "acquired" the right to host that piece of info on their search results "page".
Google "extracting" that crucial bit of info and essentially "hosting" it on their search results page could definitely be argued to be some sort of abuse of fair-use (and at this point - who is willing or big enough to take on Google on this to set a precedent? The EU, maybe? ). It's not like they're quoting a piece of a large text, they actively find the specific piece of juicy info that relates to your query and host it on their page instead of yours.
How much do you pay your users for the content they generate?
Also, they need to have team of engineers, who support infobox extractor, now this work will be done by wikimedia.
.proto on GitHub is nice, but no pricing, no public docs? This is probably great for Wikimedia coffers, but at the headline I hoped for new/improved Wikidata; instead it's.. different bordering on 'don't care'.
I found this page which has more technical details on what it actually is: https://www.mediawiki.org/wiki/Wikimedia_Enterprise
Also found out that it is open source: https://github.com/wikimedia/OKAPI
I mean, it's fine, I just got momentarily excited for something that the announcement isn't. I wanted to find a pricing page, free tier, API docs, etc. Like Wikidata but.. I don't want to say 'modernised', but made more accessible, and with APIs for higher level content like this rather than just rawer data.
It appears now that they are offering "read-only" access to existing data structured and packaged in a more convenient way.
How long before paying enterprises would like to be able to "update" content on a more efficient basis?
Perhaps Sony would like to add articles about movies that will be released soon, or as they are released? That is a pretty benign example.
Creating alerts that enterprises can subscribe to so that they will be informed if anyone adds any negative content would also be valuable.
These systems already exist in some manner, it would just make it more efficient and more common.
Imagine Wikimedia Enterprise becomes the #1 source of revenue for Wikimedia. Shortly after, people will see that Wikimedia is doing OK and become reluctant to open their wallets and donate.
Then, the top Wikimedia Enterprise customers will acquire leverage over Wikimedia and try to get Wikipedia curated to their convenience. Wikipedia articles will start being indistinguishable from advertisement.
Governments will intervene and want their share of influence too.
Top volunteers will start asking to be paid, many others will leave, some others will become critics of the project.
People will start being skeptical of Wikipedia because of their biased editorial line and then the project will be declared a failure, once everyone is angry and a beautiful project is torn apart by greed.
Also see Wikipedia: Camel's nose: https://en.wikipedia.org/wiki/Camel%27s_nose
Better way for Wikipedia to earn extra revenue are affiliate links. A lot of people when they read and learn about some topic go to Amazon and buy a book about that topic. Wikipedia could embed book affiliate links and earn commission from book sales.
Imagine reading an article from Wikipedia and at the end of the article it says "If you want to learn more about this topic take a look at our recommended books" and link leads you to dedicated page which lists let's say top 20 most popular books about that topic.
It can even replace Goodreads in a sense that Wikipedia community can rate, review and recommend books to each other.
Google's not the only enterprise out there :) I believe Wikipedia's taxonomy is used by lots of people for ML purposes, for example.
Sounds well-intentioned, but it would be immediately gamed by every unscrupulous entity and ruin Wikipedia.
https://wikimediafoundation.org/about/annualreport/2019-annu...
> Google Matching Gifts Program