Can't seem to find the connector code either. It's a little strange to go back to map reduce when presumably YARN is available and tez or spark could be used.
Can't seem to find the connector code either. It's a little strange to go back to map reduce when presumably YARN is available and tez or spark could be used.
The existing IPFS implementation involves a lot of overhead in memory, CPU use, and latency [and as you can see in the MapReduce bar graph, it's slower], but overall, it improves performance when bandwidth is the bottleneck.
IPFS is pretty similar to BitTorrent, just more practical.
With IPFS, you can effortlessly link from one tree to another, already existing tree.
In BitTorrent, you have to include a .torrent file, or an infohash in some file, but no software that I know of will easily follow that link.
This linking ability is an extremely useful property, that allows you to cheaply create a copy of a merkle tree, with a subset of the data replaced. The tree will operate identically to one created from scratch.
You also don't need to hold all the data to "patch" the tree, which I imagine is useful in this Hadoop filesystem.
Unfortunately there is no real API for operating on torrents. You could say WebTorrent is pushing in the right direction to fix this, but there doesn't seem to be much adoption for it.
YARN just refactored the scheduling / resource management out of MapReduce. Hadoop 2+ still uses the term "MapReduce" for that particular application, and that application runs on a YARN cluster, just like Tez (I think) and Spark can. I don't see anything to indicate they're NOT using YARN.