Instant.io – Streaming file transfer over WebTorrent
instant.io
instant.io
I haven't personally given much thought to solving the problem of streaming, but I am surprised that the WebTorrent FAQ doesn't mention why they didn't take this opportunity to design a protocol that has more suitable trade-offs than BitTorrent. I'm getting mixed messaging; is their goal to connect the BitTorrent network with WebRTC or enable high quality P2P streaming via WebRTC?
The problem is not the protocol but the lack of features like subtitles. And torrent contents not being standard. All this can be easily scripted to cover most cases.
Unlike simple HTTP streaming, the clients can use this spare bandwidth to download some blocks from the end of the file, even when streaming from near the start. So a sensible torrent streamer can still ensure that later blocks are not too rare.
If you download from front and back simultaneously only the middle blocks would be scarce and you still get at least half your original streaming rate. And no weirdo will stop watching midway through a movie.
That being said, though, do streaming clients make much of a contribution to the amount of seeded data?
If you were to combine streaming with a BitTorrent protocol that prioritized the most in-demand pieces you would probably not be able to watch most videos to the end without pauses to buffer or maybe even at all.
By prioritizing the most rare, even if all seeds left the swarm there's a much better chance swarm combined may still have all pieces and the torrent can still be finished.
Clients are meant to prioritize the rarest pieces in the swarm first with some randomization thrown in. Downloading sequentially is bad for the swarms health but as it turns out doesn't seem to be bad enough that the protocol can't handle it, at least as long as there's a mix of streaming and classic clients or streaming clients use a combination of classic and streaming behavior.
Genuinely interested, I know the basics of the torrent protocol, and I don't understand why the torrent protocol wouldn't work for streaming... I mean, you would just need to request the packets in order instead of randomly.
It would be less efficient, sure, but it would work.
Downloaders are only incentivized to give back data until they are done. Seeders are not really incentivized at all, so they can go away at any moment.
So if everyone downloads sequentially and seeders go away, then you can end up getting stuck in a situation where everyone has the beginning of a file but nobody has the last parts.
When you're streaming the protocol is not robust.
With random order you have a robust protocol that just degrades in throughput if people are selfish.
And with a webpage-bsed service being selfish is as easy as closing a tab.
It's proof that seeders seed for the sake of seeding for the entire system to work.
Update: It's also noteworthy that BitTorrent Inc.'s official torrent client (as well as the largest player by marketshare), uTorrent, offers sequential downloading, as well as selective file downloading. And the BitTorrent network remains very healthy.
Seeders are not abundant in all swarms. And nothing in the protocol guarantees their existence, and thus they do not contribute to intrinsic robustness.
Also, torrent clients that run in the background and consume few resources are hardly comparable to things that run in a browser tabs, user behavior will differ.
> [BitTorrent] isn't designed to sequentially stream data
We’re working on improving the algorithm to switch back to a rarest-first strategy when there is not a high-priority need for specific pieces. In other words, when sufficient video is buffered, there’s no need to deviate from the normal piece selection algorithm.
But the fact is that with the speed of today’s internet connections, the user is going to finish fully downloading the torrent in a fraction of the time it takes to view it, so they will still spend more time seeding than downloading.
In practice, the only time that the rarest-first algorithm is important is on poorly-seeded torrents, or in the first few hours of a torrent being published when the ratio of seeders to leechers is really bad. I plan to keep improving the piece selection algorithm so that WebTorrent can be a good citizen.
Also: you should note that not all WebTorrent users stream sequentially. That's just one option for downloading the data.
Also: It's noteworthy that BitTorrent Inc.'s official torrent client (as well as the largest player by marketshare), uTorrent, offers sequential downloading, as well as selective file downloading. And the BitTorrent network remains very healthy.
> why they didn't take this opportunity to design a protocol that has more suitable trade-offs than BitTorrent
BitTorrent is the most successful, most widely-deployed P2P protocol in existence. It works really well. My goal with WebTorrent was to bring BitTorrent to the web in a way that interoperates with the existing torrent network.
Re-inventing the protocol would have made WebTorrent fundamentally incompatible with existing clients and prevented adoption. The way we've done it is better. The wire protocol is exactly the same, but there's now a new way to connect to peers: WebRTC, in addition to the existing TCP and uTP.
Also, re-inventing the protocol is a huge rabbit hole. There was already a lot of risk when I started the project -- will WebRTC get adopted by all the browser vendors? Will data channel stabilize and be performant? Is JavaScript fast enough to re-package MP4 videos on-the-fly for streaming playback with the MediaSource API? My thinking was: Why add inventing a new wire protocol and several algorithms to the table?
Thanks for your thoughtful comment. Hope you'll give WebTorrent and our new desktop app, WebTorrent Desktop a try!
That sounds like you prioritized implementing streaming first over being a good citizen.
> Also: It's noteworthy that BitTorrent Inc.'s official torrent client (as well as the largest player by marketshare), uTorrent, offers sequential downloading, a
To my knowledge that is only available if the swarm condition allows and is not purely sequential. But that is second-hand knowledge, so I may be wrong. But either way, the default is rarest-first.
I read that as "it sounds like you prioritized getting a working proof-of-concept first over working out the long-term details".
And those "long term details" are implemented by all bittorrent clients, so they're hardly something novel that needs figuring out.
These long term details are implemented by all established bittorrent clients. I would bet that version 0.1alpha of many of them did not, but were rather in a state of "holy moly this works! I should go show HN".
Not to mention we're not talking about some optional, nice-to-have feature here, we're talking about a core aspect of bittorrent which gives it robustness.
Also, you forgot to address my other argument.
I just wanted to point out that "streaming over web-torrents" is the feature being demo'd here, which means that (a) it's a new feature (I assume?) to this project / these developers, and (b) it's clearly something they feel is a nice-to-have feature, because they not only chose to spend time making it, but also announced to HN when they had a working PoC. If people never posted something to HN until they were "100% complete", I think this place would be a lot less interesting than it is.
Feel free to open a GitHub issue if you have suggestions for how we can do better.
Its worth mentioning that the option is in a hidden menu
Say you're already running a video sharing site and your servers are serving up all the content to the clients. So, you add your servers as seeders. The client comes in with support for webRTC, requests packets in order, gets your servers as seeders along with a couple other people watching the video and everyone goes along their merry way.
The rare portions don't seem to be an issue because your servers are always seeds, always running, and already have the capacity to support all the demand.
Is this not a win/win to reduce some bandwidth consumption?
Your scenario is more or less the same as what we have today for those swarms that are comprised of many peers on the desktop and a few high-speed always-on seedboxes that already act like some kind of CDN.
The more seeders there are, the better, in any situation. The question is whether the swarm we're talking about is whether you can expect some seeders to be relatively long-lived (in which case streaming is ok) or if we are in a free-for-all (in which case streaming is not). Not all swarms are of the first type, far from it.
If you substitute built-in robustness with servers then yes, of course it will still work. But you're weakening the decentralized nature of the protocol by doing so.
Until browser apps can be given permission to access the file system, this will be yhe case.
Spec: https://www.w3.org/TR/FileAPI/
MDN: https://developer.mozilla.org/en-US/docs/Web/API/File_and_Di...
>The API doesn't give you access to the local file system, nor is the sandbox really a section of the file system. Instead, it is a virtualized file system that looks like a full-fledged file system to the web app. It does not necessarily have a relationship to the local file system outside the browser.
>What this means is that a web app and a desktop app cannot share the same file at the same time. The API does not let your web app reach outside the browser to files that desktop apps can also work on. You can, however, export a file from a web app to a desktop app. For example, you can use the File API, create a blob, redirect an iframe to the blob, and invoke the download manager.
Is it possible to "see" from JavaScript whether or not the download manager has downloaded the file completely, so that the file can be removed from FileAPI storage?
Unfortunately, this is a non-standard API implemented only by Chrome. You have to use other APIs like IndexedDB and WebSQL (also deprecated and non-standard) to get a working solution in all browsers.
This deficiency is really holding back the web.
You can then use idb.filesystem.js to add api support for firefox etc. Search the file above for "is_chrome" for a few idb.filesystem.js-specific quirks.
Looking at that page, it looks like firefox will ship with support in version 50?
tests:
https://instant.io/#74ce2f164e3d9ec5d5ee72c9aafc0cf5860e3d92
https://instant.io/#c241674dc3b257637abfcb08203303fc25de007f
[0] https://www.browserleaks.com/webrtc
Edit: From feross I get that WebRTC no longer leaks ISP-assigned IPs when using VPNs.
So does uBlock allow web torrents, while blocking webRTC leaks? I doubt it, because peers need to know public IP address. Unless you run a VPN client in the router, anyway.
This is incorrect.
WebRTC data channels do not allow a website to discover your public IP address when there is a VPN in use. The WebRTC discovery process will just find your VPN's IP address and the local network IP address.
Local IP addresses (e.g. 10.x.x.x or 192.168.x.x) can potentially be used to "fingerprint" your browser and identify across different sites that you visit, like a third-party tracking cookie. However, this is a separate issue than exposing your real public IP address, and it's worth noting that the browser already provides hundreds of vectors for fingerprinting you (e.g. your installed fonts, screen resolution, browser window size, OS version, language, etc.).
If you have a VPN enabled, then WebRTC data channels will not connect to peers using your true public IP address, nor will it be reveled to the JavaScript running on the webpage.
At one point in time, WebRTC did have an issue where it would allow a website to discover your true public IP address, but this was fixed a long time ago. This unfortunate misinformation keeps bouncing around the internet.
There's now a spec that defines exactly which IP addresses are exposed with WebRTC. If you're interested in further reading, you can read the IP handling spec for yourself.
https://tools.ietf.org/html/draft-ietf-rtcweb-ip-handling-01