WikipediaP2P
wikipediap2p.org
wikipediap2p.org
Nevertheless, I have a question: how does it make sure the site you receive is actually the
a) correct one (someone could distribute incorrect pages e.g. about controversial topics like North Korea or Climate Change; or inject some malicious code that e.g. uses your account to submit changes to wikipedia)
b) most recent one (or at least a reasonably new one).
I imagine the first one could maybe be somehow guaranteed through the protocol. Maybe you can just solve the second by invalidating your local copy after it is more than one day old.
The github link on the homepage does sadly not work for me (firefox on linux), so I can't check.
You're making me sad on a Friday afternoon.
To expand a bit on my previous comment, the sad part is that even though we know that the topic is not really controversial scientifically, we somehow, subconsciously, refer to it as one.
Elections and all that.
So North Korea doesn't exist right? :D
I know that for most topics i make uninformed decisions based on the information created/curated by intelligent people. Eg, global warming, vaccinations, etcetc. Sure, some of it may seem obvious, but lets be honest - i have no clue about what a vaccination does in the body, nor how climate change works. They're not school-yard topics, i'm a moron when it comes to them.
So, i play the numbers and base my understanding on a large consensus of educated people. As with the case of climate change. My coworker believes roughly the same, but unfortunately he chooses to believe "scientists who revoke climate change". He has some weird things to say, but ultimately he is doing the same thing as me.
The only thing that bothers me on this subject, is we both believe in green energy, reducing pollution, and etcetc. We have basically the same opinions on what we should be doing for the environment (reducing pollution/etc) and how important it is. Yet, he still believes it is a hoax. That part i don't understand.
Why would climate change be a hoax? Who is making money off of a "climate change hoax"? I don't get that part.. hell, part of the problem with green energy was that it didn't exist because no one was making money off it.
Anyway, /rant. All i mean to say is that people on other sides of the fence likely aren't all that different than you.. well, unless you're a researcher in the field of climate change, of course.
- World may be warming but I don't believe we should act yet
- World may be warming but I don't believe the actions we are taking make a difference
- World is warming but I don't believe we know whether the result will be necessary a "bad thing"
- World is warming but I don't believe we know the main cause
- World is warming but I don't care what happens in 100 years
- World is warming but I feel we should wait longer to see if models are correct before spending money on it
- I feel humans overestimate their ability to model the future
- I feel there are bigger problems we should be tackling..
...and a million other reasons. I don't believe global warming is a hoax, but a combination of reasons like the ones above makes me personally not worry or care about the issue.The opportunity costs towards an individual to willingly shift and adapt to climate change is fairly high -- until Pigovian taxes or some such are implemented the adoption rate will be small for prevention. We're likely in the stage now where policies around adaption need to be crafted and planned.
Government subsidies and mandates can help that. If global warming is a problem and we circumvent the market to funnel money to green energy, a hoax could also circumvent the market to funnel money to green energy.
This is one thing I really don't understand as well. Ok, there's X amount of proof in favor of AGW. Where's the evidence for your claims it's all an elaborate, international hoax? Where's the money?
I've never gotten a good answer. Usually just deflections to how "the Earth always fluctuates" and other dismissive claims that have already been answered a thousand times had they bothered to actually research it.
It's witnessing confirmation bias.
Again, I don't think that's what's going on, but I know a lot of people who question or dismiss the median conclusions in other scientific disciplines, like economics, because of some perception of systemic bias.
Myself, I might discount health studies, but maybe I'm biased by the ones that the media chooses to report, which seem to be mostly contradictory nonsense about how chocolate and wine will make you live forever.
This is just a conjecture. If I'm right, I'm not sure if that's supposed to make you feel better (because your opponents are more rational than you thought) or worse (because complex epistemology is not likely a topic you can discuss with 60% of people).
The entire "Cap and Trade" market?
And that it's still morning here on the east coast :)
Do you believe climate change is a problem and should be addressed but see others making claims beyond what the evidence supports? Or do you think climate change isn't an issue? If the former, what in particular are you thinking of? If the latter, what would it take to convince you that it's an issue?
It doesn't look like WebTorrent is meant to deal with content that is continually changing, and I wasn't able to determine (after about 30 seconds of looking) how either of the two layers on top address that
CacheP2P seems to solve the problem by including a list of sha1 hashes for all files that are distributed through P2P.
Speculation time:
I guess distributing the sha1 hashes for every wikipedia page is a bit too much (especially considering they are always changing). So with every page, you would only distribute the hash for those sites that are directly linked from it.
If you provide the correct sha1 hash for every linked website (which in turn includes the hashes of all linked websites from it...), you would form some kind of merkle-tree of wikipedia pages. Except that the graph of wikipedia pages is certainly highly cyclic (thus no tree), so this isn't feasible. And all the pages change all the time. So this naive approach has already two strikes against it and does certainly not work.
I will try to have a look at the source code at some later point, I guess.
5 Million articles, 512b per Hash, if that's enough, gives 2.5 Gb. Updates can be much smaller. Multiply that by 10 for hashes of the nearest users, who have a copy, and that's still manageable. Just of the top of my head.
edit: inb4 a blockchained approach.
5 million articles times 64 bytes per hash (512 bits) divided by 1e6 bytes per MB is 320.
So, do I understand correctly that each page (that has been distributed through WP2P) includes the hashes for all the linked pages? But I assume these hashes are only for the content, not the metadata of the linked pages. So a malicious actor could still inject a bad page (that includes bad hashes for linked pages) if he wanted.
It is true that signed wikipedia pages would be a really neat idea, since it would enable this as well as other distributed projects to make sure they are displaying the "real" version of the site.
Is this a sign that WebRTC isn't optimised yet, or is it the extension itself that needs some love?
You'll get the benefit of insuring off-line availability as well as only hosting pages you care about. Win-win.
I published the whole Wikidata dataset as separately accessible entities to IPFS and the initial publishing took ~2 weeks. Theoretically updating the dataset with weekly changes should be pretty fast after that, but if there is a change that impacts the JSON structure of every entity you have to start all over (which recently happened).
Right now Wikipedia/Wikidata is only really useful as a stress test for IPFS, but I'm optimistic for the future.
For anyone interested in this there is some (slow) progress on this at https://github.com/ipfs-wikidata , though it's admittedly a low priority for me right now, as I wait for the technical state of IPFS to improve.
One more major improvement is on the way, which is directory sharding using a HAMT data structure: https://github.com/ipfs/go-ipfs/pull/3042
I'm quite excited for directory sharding since it would make the project simpler, but the local filesystem on my server doesn't even seem to be able to handle all the files in a single directory.
Which kinda defeats the point of saving bandwidth on their side.
>js-ipfs node
That could work, though I was under the impressino that the js-ipfs implementation was WIP/Experimental (last I checked)
Traffic doesn't only need to be served by Wikipedia but could also include other volunteers (think archive.org and others) and the setup for helping out would be easy.
> That could work, though I was under the impressino that the js-ipfs implementation was WIP/Experimental (last I checked)
That is absolutely true, which WikipediaP2P is as well, and IPFS for that matter. I was simply pointing out some ways you can make the whole "Wikipedia on IPFS" thing to work in the future.
Simply pointing it at gateway.ipfs.io would achieve nothing but shift the problem around.
Or am I understanding IPFS or the ipfs.io gateway wrong? Does everything go through the ipfs.io server in that case? Or is it still distributed/torrent-like?
IMO the presented solution is a bit better since it only relies on WebRTC to establish P2P, no extra software required.
The problem is Bittorrent isn't the protocol for it, it doesn't allow incremental changes. You also don't want the complete history like git, you want something that passes a diff around.
It would be great to have a p2p network with Wikipedia, a load of academic papers, maps and recipes, which anybody with a computer could contribute drive space to storing.
I worry, though, that this space is fracturing in a way that's hampering adoption. Between Freenet, Zeronet, IFPS, I2P, not to mention things like Retroshare, Matrix, and so forth, there's a lot of redundancy but with important differences between any given two solutions. I'm not sure what can be done about it, but it seems like something in this area needs a network effect to accelerate adoption, but to do that, it needs to be comprehensive in what it offers while also being sufficiently differentiable from other projects. It's great to see so much going on in this area, but it's getting difficult to know what does what, and where one might put some resources.
Also, the “Fork me on GitHub” banner on your website is behind the fancy canvas so we can’t actually click it (Chrome 54).
Great idea otherwise!
Remember diaspora?
>""This is an unofficial extension and not endorsed by Wikipedia in any way, it's just an implementation of latest web-technologies towards sharing of knowelege."" <---knowledge
The two examples here are:
> This color means it's rearching for that article.
I assume that's meant to be 'searching'
And: > This is an unofficial extension and not entitled by Wikipedia in any way
I'm assuming 'entitled' is meant to be 'endorsed'.
Come to think of it, it is interesting that they server all this traffic using Varnish + HHVM/3.3.0.
While the P2P idea is very nice, most residential lines are heavily geared towards download speed. Upload speed is often only 5-10% of the overall speed. Using a P2P network outside of your Lan will probably lead to a slower experience than using a CDN that has PoPs very close to you.
Now I'm pretty sure that Cloudflare would save Wikipedia at least 95% of bandwidth with the free tiers. And I don't think that Cloudflare would complain, the gain of having them as clients would be larger than the costs for bandwidth. Similar for other providers, Wikipedia and their traffic levels would probably get very cheap offers.
But I can also understand that it would somehow question the independence of Wikipedia from corporate interests. That's why they didn't do it.