BitTorrent vs. HTTP
daniel.haxx.se
daniel.haxx.se
It's a standard, we know it works and how.
I dont want to be (potentiall forced to) distribute software.
With recent developments in net neutrality and data plans, p2p could drastically impact my data plan and cost me money.
Sure I can see why you would think the potential is, but for me p2p data distribution is out of the question, at least for now.
Just because you have a logical reason to not want to serve updates, should not prevent IT admins from getting a good default experience when they have 1000 machines on a LAN that all want to update.
Pretty much all internet connections are metered/capped, though prior to recent FCC transparency rules the caps on residential fixed broadband were often undisclosed or affirmatively misrepresented under false "unlimited" labels.
Making the OS appropriately sensitive to the costs incurred by each network-involved transaction relating to updates and the system owner's preferences regarding balancing those costs against the value they provide in update experience is abstractly ideal but decidedly difficult.
Note that, even in that case, the connection to the outside world might still be metered/capped—but the connection to LAN peers obviously isn't. That means that "does this traffic cost anything" is an evaluation the OS would have to make per socket (or requested socket from a higher-level library, like Windows' BITS), rather than per interface.
Disclaimer: Haven't checked in a while and can not do it now, but it was there until not long ago.
What I can't remember and am unable to check at the moment is whether Metered is turned on by default or not.
But regardless of defaults, if you mark a connection as Metered than the background update distribution is turned off.
Done respectfully, it can be a great way to make everybody's file arrive faster while balancing the traffic better.
They already have local streaming for games so they have the infrastructure for this.
The solution is to pressure your government to make net neutrality a law and vote with your dollars and use ISPs and network providers that are not intent on breaking the Internet. All Internet nodes need to have the same access to the network. "I am too lazy to do anything about it and too cheap to pay for data" is not a valid argument on a discussion forum called Hacker News.
Seeding a website would be like paying a membership fee with surplus currency that you have been throwing away monthly.
With the proper incentives and controls I suspect there is content you would be compelled to seed.
It's incredible how Opera was innovative at the time. Nowadays it just another useless chrome.
Too bad it couldn't keep up with the whole Web 2.0 sh*tstorm.
The great thing was that Opera 12 even with IRC, mail, torrent and other goodies (such as HTTP server) was still lighter on resources than Firefox and Chrome.
I really do miss the old Opera, and even the Vivaldi is not really good replacement since it still uses Blink engine.
I switched to Firefox after, I like its extensions but it's not perfect.
[0] http://www.seamonkey-project.org/start/
^ version 4 I think but maybe earlier versions too
HTTP headers can allow to know when to flush cache (just like it's currently done) and provide last known md5/sha1/whatever digest to make sure page is not tempered with (let's say it's checked when the download is complete, and retry a download if the signature does not match: it should not happen often anyway). It obviously won't work for pages which distribute auth related content, but it would be great for assets.
I guess a problem could be that page load will be slower (depends on the ability to parallelize and to contact geographically close peers, I suppose), but it would mean way less heavy load on servers.
The privacy concern makes me realize something else : there's no incentive for users here. They lose privacy, what do they win?
It does not prevent totally the idea of using p2p to distribute content as a low level implementation, but since we (the ones who manage servers) are the only ones to benefit from it, we must first find a way for it to have no impact on users.
The biggest privacy problem is because of how p2p works currently : we have a list of IPs associated to a resource. How can we obfuscate this without going through proxies?
It works if we don't do naive P2P but rather Friend-2-Friend; this is what retroshare does, your downloads can go through your friends so that only they know what you download. But that requires a lot more steps than traditional bittorrent, so I'm not sure it could work in general.
Perfect Dark (a Japanese P2P system) is a direct implementation of that concept, where you automatically "maintain ratio" as with a private torrent tracker by your client just grabbing a bunch of (opaque, encrypted) stuff from the network and then serving it.
A more friendly example, I think—and probably closer to what the parent poster is picturing—is Freenet, which is literally an onion-routed P2P caching CDN/DHT. Peers select you as someone to query for the existence of a chunk; and rather than just referring them to where you think it is (as in Kademlia DHTs), you go get it yourself (by re-querying the DHT in Kademlia's "one step closer" fashion), cache it, and serve it. So a query for a chunk results in each onion-routing step—log2(network size) people—caching the chunk, and then taking responsibility for that chunk themselves, as if they had it all along.
Good ! Now I know you have JQuery 1.2.3, which is vulnerable to exploit XYZ, which I can now use to target you. This is one reason why apt-p2p and things like that can't be deployed in large; it's way too easy to know what version of what packages are installed on your machine.
This makes me think that an other feature could be to not disclose all peers available, but randomly select some. This would force someone wanting to look someone up to download a possible big amount of times the same list to check for an ip, instead of just pulling it once per ressource to know as a fact.
Both ideas are not exactly privacy shields, but steps to mitigate the problem.
Sounds like a cool way to deter people from using it if you ask me.
I'm not sure if the cap is in effect for places with fiber internet (like Verizon FIOS). It wouldn't surprise me, since they've admitted caps are unnecessary.[1]
[0] Charter + Time Warner Cable (after merge) is larger, I think.
[1] https://consumerist.com/2015/11/06/leaked-comcast-doc-admits...
About the time Chairman Wheeler was aiming for net neutrality, he was also looking at data caps and if they were bad. Within a month, Comcast increased theirs to 1 TB. Coincidence?
There's going to be a lot about how 'competitive' ISPs are in the next few years in America now that everything's Republican. But most people have two options: expensive cable ISP that's fast when it wants to be, and slow DSL that's also expensive and not getting better.
Using: 1 terabyte = 8,388,608 megabits
Ultra HD: 25 mbps ~= 93.2 hours HD: 5 mbps ~= 19.5 days SD: Would be good but hey I bought an HD TV for a reason so that would not be fair to skimp out on quality just because they don't want to introduce cheaper bandwidth caps.
side note It is $10 per 50 GB once you go over the cap with Comcast which is pretty huge for overhead if a user wasn't even aware of the cap. Thankfully they cap it at $200 over your bill.
side note2 this number will only increase - so really that 1 TB cap already needs to see a lift to 2 TB just to start to fulfill the requirements of the newer age web.
Wow, that's disgusting.
Back when I was on cheap and cheerful, capped internet - if you went over the (paltry 100GB, although this was the best part of a decade ago now...) cap, I'd pay an extra £6 on my bill and that was it, regardless of how much I went over by.
Ironically, the price to remove the data cap as part of your package was more expensive than the over-cap charge.
In short their approach is that instead of connecting point to point and using addresses of hosts, what if we could address based on the data we want.
This suddenly makes routers aware of what data they are forwarding, that knowledge allows them to start caching and reuse the same data packets when multiple people request the same data.
Things like multicasting or multi path forwarding are simpler. Interestingly, the more people are viewing the same video for example the better, essentially CDN is no longer needed.
Most of use cases how we currently use the Internet, are actually easier to do this way.
Things that are harder though (but not impossible to do) is point to point communication, such as SSH.
(note to others: the description of project can be found here https://named-data.net/project/ )
Actually, what could be done without "rebooting" the whole internet is to use caching relays, which would operate just like DNS servers do, but caching content instead of IPs (maybe it's what they're suggesting, it was unclear from general description).
EDIT : but I wonder if I'm really ready to deal, as a webdev, with content propagation based on TTL, like we do with dns zones :D
EDIT 2 : actually, this is a non issue. TTL is there because when caching something as short as a IP string, you don't want to have to issue a request to parent node. But when caching assets, it won't be that a problem to issue a HEAD http request to asset owner to ask if the cached content should be invalidated.
If you download, for example, a Windows.iso from technet and start sharing it via the integrated p2p system, I can see a shitstorm of lawsuits raging through the web.
You’ll be perfectly fine sharing legal stuff like Linux ISOs or mods for your favorite game using BitTorrent 24/7.
Users should not fear a content-agnostic protocol.
Mass adoption of bittorrent would render this type of discrimination much more difficult.
I think this is a part of the exciting proposition intended by WebTorrent (webtorrent.io) in having an entirely JS/client-side-capable BitTorrent stack that can run directly inside a browser (or in Node or Electron or Cordova apps) with potentially minimal setup.
Why probably? This is exactly how the immensely popular Popcorn Time worked.
"it will not at all distribute the load as evenly among the peers as the "random access" method does."
If people only stay to download the file and then disconnect then there will be a higher concentration of parts from the beginning of the file than to the end. This reduces the resiliency of the swarm.
For instance, if it's a movie being streamed, the majority of connections will finish downloading the movie long before the user finishes watching, making a pretty decent seed time. Those with a slower connection wouldn't have made a particularly great seed anyway.
The only thing that kills swarms is average seed ratio < 1
uTP is a very nice bittorrent over UDP protocol implementing LEDBAT congestion control algorithm [1].
[1] https://opensource.apple.com/source/xnu/xnu-1699.32.7/bsd/ne...
Probably due to the sorry state of affairs that is the consumer router space, running BitTorrent renders even web browsing painfully slow and sometimes completely nonfunctional.
Why BitTorrent does this while normal http transfers do not is not clear to me. Perhaps due to the huge number of connections made.
Either way, when given a choice I'll always take a direct HTTP transfer over a torrent, for no other reason than the fact that I'd like to be able to watch cat videos while the download completes.
On your upstream side (i.e. data you send out to the Internet), you can control this by using a router that supports an AQM like fq_codel or cake. Rate limit your WAN interface outbound to whatever upstream speed your ISP provides. This will move the bottleneck buffer to your router (rather than your DSL modem, cable modem, etc., where it's usually uncontrolled and excessive) and the AQM will control it.
Controlling bufferbloat on the downstream side (i.e. data you receive from the Internet) is more difficult, but still possible. You can't directly control this buffering because it occurs on your ISP's DSLAM, CMTS, etc., but you can indirectly control it by rate limiting your WAN inbound to slightly below your downstream rate and using an AQM. This will cause some inbound packets to be dropped, which will then cause the sender (if TCP or another protocol with congestion control) to back off. The result will be a minor throughput decrease but a latency improvement, since the ISP-side buffer no longer gets saturated.
The best thing to do is avoid home networking equipment entirely. The Ubiquiti EdgeRouter products are cheap and good, as is building your own router and sticking Linux/*BSD/derivative distros on it.
I'd like to mention that both LEDE (www.lede-project.org) and OpenWrt (www.openwrt.org) were the platforms used for developing and testing fq_codel/cake. That means that people may be able upgrade their existing router to eliminate bufferbloat.
My advice: if you're seeing bufferbloat (and a great test is at www.dslreports.com/speedtest) then configuring fq_codel or cake in your router is the first step for all lag/latency problems.
http://lartc.org/wondershaper/
I've been using it for 15 years and it's still working great. Even with multiple P2P clients, stuff like HTTP, SSH, and gaming keep a low latency. Also you learn a lot about networks just by configuring it :-)
Two key reasons, usually, both related to congestion control (or practical lack thereof).
> Perhaps due to the huge number of connections made.
This is one of those reasons: unless the other end of your incoming connection is prioritising interactive traffic somehow packets for each stream will get through at more or less the same rate once the connection is saturated. So if you have a SSH link and are requesting a http(s) stream (for a web page or that cat video) while a torrent process has 98 connections getting data, for every 100 packets down the link only two are for your interactive process. On fast enough link this isn't an issue, but "fast enough" needs to be "very fast" in such circumstances as it is relative to the combined speed of all the hosts sending data. You can mitigate this by telling the torrent client to use minimal incoming connections (limiting incoming bandwidth can have some effect but is generally ineffective as bandwidth limits like that need to be applied on the other side of the link).
The other problem is due to control packets such as those for connection handshakes and so forth fighting for space on the same link as those carrying data. As soon as the connection is saturated in either direction so that there are packets queued for more than an instant, latency in both directions takes a massive hit. This is particularly noticeable on asymmetric links such as many residential arrangements. You can mitigate this by throttling the outgoing traffic either within the torrent client or at other parts of the network (assuming the traffic isn't hidden in a VPN link that means you can't reliably distinguish it from other encrypted packets) and reserving some bandwidth for giving priority to interactive traffic and protocol level control packets but you have very little control (usually practically none) over traffic coming the other way as you the measures have to be taken before the packets hit the choke point and you don't control those hosts your ISP does (they will implement some generic QoS filtering/shaping but more than that requires traffic inspection which we don't want them to do, and they don't want responsibility either legally or in terms of providing/managing relevant computing capacity).
(the above is a significant simplification - network congestion is one of those real world things that quickly gets very complicated/messy!)
Bittorrent very quickly uses 100% of your upload speed, which effectively breaks the internet. The solution is to limit upload speed in your client
My upload speed is awful (~30KB/s), so I limit it to 5KB/s
For higher speeds you can probably get away with limiting it to 50-90%
Actually there is a FF extension call downthemall that multithreads HTTP downloads which will infact max out your inbound speed just like torrents do. As far as router are concerned if you set a particular device with a higher QOS then the one that is torrenting you should not have the problem.
In fortunate circumstances, Bittorrent, it can run fully P2P without needing 'servers' -- depending on whether you have to holepunch a NAT, whether you're okay with peer discovery taking longer in the DHT without a tracker, and whether you're okay with no fallback webseed as a seed-of-last-resort. Magnet links, a distinct concept not only applicable to Bittorrent, are the cherry on top, which enable even the ".torrent" file (or rather, the equivalent information) to be found without having to possess the .torrent file itself, from just a hash.
BT's architecture has a nice effect that popular content consumes fewer of the original host's bandwidth than it would with HTTP, but this works best if peer interest roughly coincides in time. This is why it's a good fit for distributing patches [1][2], or newly released episodic content, and a fairly poor fit for anything else.
The Bittorrent extensions (BEPs) vary widly in quality, clarity, verbosity, and rigor, and aren't always up to the level of HTTP extensions acknowledged by the IETF. Most BEPs were implemented in only one product and then submitted as a BEP retroactively, while throughout its lifetime most of HTTP's enhancements were discussed and developed in a semi-open, but publicly viewable process in the IETF with more emphasis on consensus and early prototyping, rather than final implementation and retroactive standardization. These days this process has partially been subsumed by prominent vendors implementing a behavior, running it for an extended amount of time as a proprietary enhancement, then submitting a slightly altered version to the IETF for discussion, so admittedly the two enhancement processes are a lot more similar now than they've been in the past.
[1] http://arstechnica.com/business/2012/04/exclusive-a-behind-t... [2] http://wow.gamepedia.com/Blizzard_Downloader
> Bittorrent is a protocol designed by the company named Bittorrent and there is a large amount of variations and extensions in implementations and clients.
Extensions are in many cases standarized. There's a lot of variations in HTTP as well.
One important thing that's missed is that a magnet link also contains a hash of the torrent file, which points to a specific file. Over HTTP, everyone across the network connection can change the contents of the file and there's no automated way of detecting that.
https://bugs.chromium.org/p/chromium/issues/detail?id=616212
that have to be worked around by assuming that other devices might not adhere to the specification.
On an unrelated issue, it baffles me how people with uncapped internet refuse to help serve content to other people, like enabling the distribution of windows updates. I would love to have that on linux and would never dare to turn it of unless is was reaching my data cap. It's just liek seeding you torrents, hell i have seed ratios of over 100 for some linux distro images.
Many things can benefit from MPTCP however, including BitTorrent.
>only gets a few
/s/gets/has/
Minimal implementations of HTTP (and I'm strictly talking about the transport protocol, not about HTML, JS, ...) is dead simple and relatively easy to implement.
Of course there's a ton of extensions (gzip compression, keepalive, chunks, websockets, ...), but if you simply need to 'add HTTP' to one of your projects (and for some reason none of the existing libraries can be used) it shouldn't take too many lines of code until you can serve a simple 'hello world' site.
On top of all that, it's dead simple to put any one of the many existing reverse proxies/load balancers in front of your custom HTTP server to add load balancing, authentication, rate limiting (and all of those can be done in a standard way)
Furthermore, HTTP has the huge advantage of being readily available on pretty much every piece of hardware that has even the slightest idea of networking. Any new technology would have to fight a steep uphill battle to convince existing users to switch.
Have I mentioned that it's standardized and open?
(I'm speaking about authenticated connections. For anonymous access - which should be read-only anyway - you're usually better off using HTTP anyway)
So in that case, at least, I was very glad FTP(S) was still around.
FTP does have some advantages, but HTTP has more advanced support for resuming connections, virtual hosting, better compression, and persistent connections, to name a few.
I bet if we had used FTP instead of HTTP for serving HTML right from the start, FTP would today have all of the same extensions and the same people would argue for it being too bloated :) (HTTP started as pretty minimalistic protocol back in the day)
I often find the discrepancy between what HTTP has originally been designed for (serving static HTML pages) and all the different things it's being used for today highly amusing. Yes, some of todays applications for HTTP border on abuse, but its versatility (combined with its simplicity) fascinates me.
The two success factors of http are statelessness and fixed verbs.
http://i3.kym-cdn.com/photos/images/original/000/732/170/796...
HTTP is quite a good protocol. Simple, extensible to a sane extent, but not overly extensible (XMPP i'm thinking about you). HTTP is not accidentally successful. FTP is a bad joke. (stateful. binary mode, 7 bit by default. uses multiple connections (unless in passive mode))
Which are dead simple to construct, send, receive and parse.
Really.
For example, let's curl -L (view everything but the body) for the spec: http://www.ietf.org/rfc/rfc7230.txt
HTTP/1.1 200 OK
Date: Tue, 24 Jan 2017 12:00:55 GMT
Content-Type: text/plain
Transfer-Encoding: chunked
Connection: keep-alive
Set-Cookie: __cfduid=df57c7720b704a40e4c3367bbe248771c1485259254; expires=Wed, 24-Jan-18 12:00:54 GMT; path=/; domain=.ietf.org; HttpOnly
Last-Modified: Sat, 07 Jun 2014 00:41:49 GMT
ETag: W/"3247b-4fb343e4dcd40-gzip"
Vary: Accept-Encoding
Strict-Transport-Security: max-age=31536000
X-Frame-Options: SAMEORIGIN
X-Xss-Protection: 1; mode=block
X-Content-Type-Options: nosniff
CF-Cache-Status: EXPIRED
Expires: Tue, 24 Jan 2017 16:00:54 GMT
Cache-Control: public, max-age=14400
Server: cloudflare-nginx
CF-RAY: 326353e6a6a257a7-IAD
A bunch of newline, (CRLF), seperated key-value mappings. Some with a DSL (Such as Set-Cookie).
It gives you a status message instantly, a date to check against cache, a Content-Type for your parser, acceptable encoding, for your parser, a bunch of other values for your cache. All for free.
As for the body of the content? For a gzipped value like this, it's everything outside the header, until EOF. That's not quite as easy as when the content-length parameter is given, but hardly difficult for parsing.
HTTP is easy.
In fact, HTTP is so easy, that in-complete HTTP servers can still serve up real content, and browsers can still read it.
HTTPS is more complicated, but if you simply rely on certificate stores and CAs, it becomes much easier, but HTTPS is a different protocol.
This is chunked and keep- alive. Things get a little trickier