BitTorrent v2 (2020)
blog.libtorrent.org
blog.libtorrent.org
Hopefully we'll start seeing hybrid torrents/magnets in trackers soon. Although these are mostly automated so I'm not sure how fast the tools will catchup.
One thing I'd like qBittorrent and other torrent clients to do is to allow users to upgrade v1 torrents to hybrid torrents if they're fully seeded, this would allow cross polination for old torrents where peers are scattered between the original torrent and batch torrents, or where trackers add extra files causing the v1 hash to change.
https://github.com/transmission/transmission/issues/458
Elsewhere someone quipped:
> That transmission bug is 4 years old and they've never announced plans to prioritize it (granted the effort is surprisingly high; they've never allowed empty dict keys in bencoding). I think either they either don't know or don't care that their market share makes them an adoption blocker.
Absolutely.
The truth seems to be that between 2014 and mid-2021 the "transmission dev team" was a single (probably part-time) maintainer:
https://github.com/transmission/transmission/graphs/contribu...
1. https://stackoverflow.com/a/65299426/2700296
2. https://github.com/webtorrent/magnet-uri/pull/43#issuecommen...
"Enforcing these encodings also help make it more likely that two people creating a torrent from the same files, end up with the same info-hash and the same torrent."
In the discussion of backwards compatibility, it appears that v1 collections of identical content encoded with different piece sizes will all become the same v2 collection when upgraded.
Granted, that's something like 0.4% of the new torrents made every day, but still, it's not a rounding error anymore.
The page for the crawler I started with is long gone, but this appears to be a reasonable copy on github:
https://github.com/zlzlovezl/simDHT-1
I'm sure you can find other copies, with features added. Looking around, it seems like quite a few people have taken this and made it their own over time.
It's far from perfect in this state, but a fun toy to play with. I will say, from a point of painful experience: The way this crawler works tends to make some of the various people that track such things think you're sharing pretty much all the things. If that's an issue where you live, don't run it from your home connection.
One of these days I should find somewhere to put all these bits of metadata I've been downloading.
Keep it up!
https://archive.org/details/torrent_metadata_archive_sample
I'll start uploading monthly archives.
The first few months will be the biggest, then it should settle down to something like 15-30gb a month.
That's not the protocol's fault, but it's pretty much one of the biggest reasons most common people use the protocol: to pirate.
Get caught pirating? Probably get some fines, right?
Sheesh, please do a more thorough research. It's technically not DMCA (because it's not called DMCA), but Germany and Japan do have similar laws analogous to DMCA, and I'm pretty sure that further digging would reveal more countries that have an analogous process written in law.
Though really you could argue that the Berne Convention already covers the most obvious copyright violations you get from torrents which means that effectively every country on Earth has laws against bare-faced copyright infringement. The WIPO Copyright Treaty is most notable for Article 11 which effectively prohibits the circumvention of technical measures (DRM).
It may be practically difficult to litigate someone in another country, but it has happened in the past[2].
[1]: https://en.wikipedia.org/wiki/WIPO_Copyright_Treaty [2]: https://en.wikipedia.org/wiki/The_Pirate_Bay_trial
I think a major factor there is that IPFS supports pluggable transports, so it can easily be used in both backend and browsers environments directly, while WebTorrent (powered by WebRTC) can't directly communicate with most traditional BitTorrent clients (which use TCP/UDP directly) without going through a bridge node.
It's a novel layer 1 crypto with some interesting tech behind it. A lot of decentralized nodes inspiration from BT went into it as well.
slight pedantism here, but "when" not "if"
Increases in computing power alone can't render SHA256 insecure; that can be proven by thermodynamic/size-of-the-universe arguments. You need an actual cryptographic weakness.
Fascinating. Could you please recommend a source where a curious layman could learn more about this?
In the time since SHA1 came under attack, a lot more research has happened on hash function security. If this has shown anything, it's that SHA256 is more secure than people used to think.
People seem to have this widely believed idea that "all crypto will be broken eventually". But in pretty much all cases you can trace back that broken crypto was known to be weak for a long time. SHA256 is not known to be weak (with the slight caveat that it doesn't protect against length extension attacks, but this is a known property that matters only in very rare circumstances).
You're not gonna break SHA256 with faster computers.
Breaks of symmetric crypto or hashes that are practical require an actual cryptographic weakness, whose existence is not a given.
Then there's quantum computing, but that also doesn't break symmetric crypto outright, it just makes it weaker. I'm not sure if anyone has run the math, but we'd probably still be talking galaxy sized quantum computers to get anywhere.
I have been constantly surprised that there isnt a phone/company that specifically markets to "buy this temp phone to cross a border"
--
Also, WRT Quantum computing, what are your thoughts on the talk by D-Wave CEO where he said "we can reach into different dimensions" -- and soon after, D-Wave went super silent?
What is the state of quantum currently - did they discover something super secret?
Specifically, modern symmetric crypto (and hashing which is a related but different thing) is fine. It's public-key crypto that has always been the worry thanks to Shor's Algorithm, and all modern widely used crypto there is indeed vulnerable if an actual scalable general quantum computer could be constructed. There are a variety of post-quantum cryptographic algorithms under development including some ones that would be ugly slow but likely effective bandaids if required, but that's still a not totally unreasonable concern for certain threat profiles. But yes GP was totally wrong about "all methods are broken when compute becomes faster".
>I'm not sure if anyone has run the math
"The math" here is just Grover's Algorithm, that lowers the cost of generic brute forcing to O(sqrt(n)). So a 128-bit symmetric key could be cracked on average like 2^64 or a 256-bit key like 2^128. Obviously this is utterly trivially countered by doubling the exponent. Even 2^128 is still ridiculous, and going to a 512-bit key brings us right back to 2^256 which is impossible. 128-bit keys should indeed be phased out entirely (and that seems to be well on the way anyway), and perhaps 256-bit too at some point (even if it's only the very paranoid or very long termers, it's also not big stretch on modern hardware at all, so eh), but fundamentally yes symmetric crypto is fine.
So like lets say a year had 256 days and you want to find colliding birthdays. First you would check how many birthdays in 0-127 vs 128-255. This takes just two counters - very little storage. Then lets say the former has more. You'd try 0-63 vs 64-127. Etc.
But yeah this will of course miss collisions. But I think if you have enough children (where enough is still pretty close to sqrt(n)) it should have a good chance of finding a collision.
And subdividing into many buckets at a time will miss less collisions, but require more storage - a tradeoff can be used.
https://pthree.org/2016/06/19/the-physics-of-brute-force/
The total energy output of a supernova is enough to count up to 2^220, making a lot of generous (invalid) assumptions. You'd need 2^36 supernovas to just count up to 2^256, never mind actually running SHA256. At minimum.
There are circa 2^38 stars in the Milky Way, and they're not going to all go supernova. Add in the actual cost of compute and it just isn't happening, not in this galaxy.
For 128-bit crypto we can look at a more earthly calculation. Using the same math as that article, but at room temperature, and taking the total solar irradiance on the Earth as an energy source, it would take this long to count up to 128 bits (calculated using Google):
((2^128) * (1.38064852 * ((10^(−16)) (ergs / K))) * (298 K)) / ((1361 (W / (m^2))) * (pi * (radius of Earth^2))) = 8.04911615 seconds
Definitely more on the plausible side, but we're already moving away from 128-bit crypto and there's a staggering number of generous assumptions being made here; we aren't going to be getting thermodynamically ideal computers using a significant fraction of the total solar irradiance on the Earth any time soon.
Just to give you an idea of how far away we are from that, looking at actual SHA256 calculations:
https://www.iea.org/data-and-statistics/charts/efficiency-of...
22222 MH/J is the highest, or 4.5 × 10^-12 joules per hash. That's a factor of 2^30 worse, so if we used the total solar irradiance of the Earth to power Bitcoin miner style ASICs, it would take about 272 years to go through 128 bits' worth of brute forcing something of similar complexity to SHA-256.
The factor is more like 2^36 relative to the first "temperature of outer space" calculation. That happens to be about the number of galaxies in the observable universe, so using current technology, it would take somewhere on the order of all the stars in the observable universe going supernova to power through a single SHA-256 brute force.
If your algorithm is not linear in the size of the key/hash space, it's not a brute force algorithm. It's a cryptographic break.
There is no proof that such algorithm does not exist, nor is there any proof that it does. Therefore you can't claim it definitely does. That's my point.
See current research on SHA256
https://en.wikipedia.org/wiki/SHA-2#Cryptanalysis_and_valida...
https://dusted.codes/sha-256-is-not-a-secure-password-hashin...
https://security.stackexchange.com/questions/34256/sha256-se...
The password hashing remark is correct, but irrelevant. SHA256 is not a good password hash, because it's not made to be a password hash. Don't use it for passwords.
The weaknesses that broke MD5/SHA1 were known since 1994. No similar weakness is known for SHA256.
Btw cheers! I recall you from correspondence years ago regarding fuzz testing linux packages.
No such thing happened.
You're likely referring to this paper quoted in Wikipedia: https://eprint.iacr.org/2016/374.pdf
This is 1. not about SHA256, but about the truncated sha2 algorithms and, more important, 2. it's an attack on reduced-round versions of these algorithms. That's a common thing in cryptoanalysis to do. You're basically saying "we have no way of attacking the full algorithm, so let's build a version which is much less secure and try to attack that". That's a valid thing to do for research purposes, but it needs to be interpreted correctly. "I can break this massively weaker version of the algorithm" is very different from "I have shown a practical attack on the real algorithm".
> You're proposing to add extra complexity for some hypothetical scenario that is unlikely to happen
I disagree. It's evidence that the given path is not fruitful for breaking SHA256 in the future. When numerous researchers worldwide have attacked a function, and only broken X out of N rounds, for X << N, and future research hasn't been able to improve X for years, that's pretty good evidence that the technique used isn't going to continue to apply for more rounds. The existence of a multi-billion dollar bug bounty for breaking SHA256 (Bitcoin) that's gone unclaimed for years is further evidence that it's quite strong.
Because it wasn't designed for it. For password hashing you want a hash that has a salt (so that the same password on two accounts doesn't have the same hash on the database) and is as slow as possible (that is, fast enough to validate on logins) to increase brute-force time.
Historically Bcrypt was a good option, but I think Argon2[0] is the current best option.
Point taken about hash calculation speed, though.
Also, if you use SHA256 bare without salt you are vulnerable to precomputed dictionaries. There is a fascinating way to make these dictionaries shorter called Rainbow Tables.
Better key derivation functions take a lot of time and memory to make it costly for the attacker to a lot of guesses at the password and they have built in salt. Examples would be SCrypt or Argon2.
A password hash is a function with 3 inputs: the password, the difficulty factor, and a customization structure (most often just a salt, sometimes more). A cryptographic hash is a function with 1 input: the message.
Password hashing functions have variable performance in time and usually memory and cache use, controlled by the difficulty factor. Cryptographic hashes have fixed performance.
Password hashing functions take a salt and possibly other customization data (a secret "pepper", a fixed domain-separation string if it's also a key derivation function, etc).
It's possible to build password hashes from cryptographic hashes. One must be very careful about encoding the multiple inputs into a single "message" for the cryptographic hash, to avoid cannonicalization attacks.
Argon2id is a good password hashing function.
I'm also not sure I'd say the security community as a whole feels that cryptographic agility is an anti-pattern, there are plenty of vocal voices on all sides of that argument. I personally agree that you gain very little from cryptographic agility but have to pay a fair amount, but I don't think it's as settled as you make it out to be.
I totally agree with this assessment, and rereading my comment, I say that as though it’s a lot more settled than it is.
I think it’s probably harmless to tag those sorts of bits of data, but also JWT has shown us that having that metadata in band does lead to people using it for those kinds of decisions whether they’re supposed to or not.
> Like the urn:btih: prefix for v1 SHA-1 info-hashes, there’s a new prefix, urn:btmh: for full v2 SHA0256 info hashes. For example, a magnet link thus looks like this:
> magnet:?xt=urn:btmh:<tagged-info-hash>&dn=<name>&tr=<tracker-url>
> The info-hash with the btmh prefix is the v2 info-hash in multi-hash format encoded in hexadecimal.
https://torrentfreak.com/bittorrent-inc-confirms-acquisition...
and that after the creator of bittorrent left he created his own cryptocurrency:
> The Chia Network was founded in 2017 by American computer programmer Bram Cohen, the author of the BitTorrent protocol.[4] In China stockpiling ahead of the May 2021 launch led to shortages and an increase in the price of hard disk drives (HDD) and solid-state drives (SSD).[5] Shortages were also reported in Vietnam.[6] Phison, a Taiwanese electronics manufacturing company, estimated that SSD prices will rise by 10% in 2021 as a result of Chia.[7] Hard drive manufacturer Seagate said in May 2021 that the company was experiencing strong orders and that staff were working to "adjust to market demand".[6] In May 2021 Gene Hoffman, the president of Chia Network, admitted that "we’ve kind of destroyed the short-term supply chain" for hard disks.[8] Concerns have also been raised about the mining process also potentially being harmful to drives' lifetime.[9]
Thankfully the coin seems to have collapsed completely since May 2021 (https://www.coinbase.com/price/chia-network)
But you can see why people are skeptical
That doesn't appear to be the case:
https://www.coingecko.com/en/coins/chia
Currently rank #257 with a market cap of $247 million.
The whole altcoin market has been in limbo/declining since May last year.
Last update was also yesterday: https://www.chia.net/2022/01/10/understanding-the-changes-in...
People are still struggling to articulate a legitimate use case for cryptocurrencies.
I made some money mining BTC way back in the day with a GPU. I bought a new computer for $1,900 last year and much more than paid for it mining ETH. I’ve spent years now listening to all the claims from BTC, BCH, BSV and a dozen more coins and web3 proponents. But I’ll be damned if I can articulate a legitimate use for it that is better than what we have today or that doesn’t fail to take human nature into consideration.
What's wrong with DeFi?
Edit: Oops. I didn't see there are two "past" buttons. I guess one of them does do that. TIL.
The "if there are any" part is more important than the actual links.
v4.4.0 changelog: FEATURE: Support for v2 torrents along with libtorrent 2.0.x support (glassez, Chocobo1) - https://www.qbittorrent.org/news.php
[1] https://torrentfreak.com/biglybt-is-the-first-torrent-client...
[1] https://github.com/picotorrent/picotorrent/releases/tag/v0.2...
[2] https://biglybt.tumblr.com/post/629947579144290304/2500-rele...
That's exactly what chunking based on a rolling hash solves. You set the average size of chunks and the content controls the exact boundaries.
"the idea is to select blocks not based on a specific offset but rather by some property of the block contents"
https://en.wikipedia.org/wiki/Rabin_fingerprint#Applications
The part that surprises me from reading this is the issues raised from moving from SHA1 to SHA256, specifically because that's 20 vs 32 bytes respectively. This has created compatibility issues.
The same thing happened with Git and it's a nightmare.
I honestly don't understand how this happens. We've gone through this many times for many years. Replacing your hash function should be built in from day one and for any project to come along in the 21st century and not do this is really gross negligence.
I really wonder how this happens. Like is it just hubris about how this hash function will be different? Or are engineers just so in love with the optimizations they get to make by assuming, say, 20 byte hashes? I really wish I knew.
Anyway, I'm a little sad BitTorrent only really found traction for pirating media. Whether you approve of that or not, it's still an astounding technical achievement.
You aren't really gaining anything, and you might as well just have new versions of the protocol, and save all the pointless ceremony, complexity and security holes.
There is no way (beyond maybe embedding some kind of executable code in the protocol itself, which seems like an insane idea) to have support from "day one" for algorithms and functionality that don't even exist yet.
> all you need is the root hash of the tree.
but on the other hand:
> the .torrent file must still contain these piece hashes
So what are we saving here? If a piece hash doesn't match the downloaded data, we re-download the piece.
> The .torrent file size is not smaller for a v2 torrent, since it still contains the piece hashes, but the info-dictionary is, which is the part needed for magnet links to start downloading.
Also it seems that it'll be easier to detect peers which are sending bad data and correct it. I guess this is because previously you could fetch a piece from more than one peer but because the hashing was done at the piece level you didn't know which part of the block was wrong (meaning you'd have to redownload the whole thing). But now you can -- with a bit of tracking -- tell which block of a piece was sent from which peer and then redownload it from another (likely blocking the original peer), because you know which block hash failed to validate. (EDIT: This is correct -- [1] explains the issue and how torrent clients have worked around this problem in v1 torrents.)
Merkle trees have some other benefits (you can cache parts of the tree in such a way that repeatedly checksumming data where lots of leaves are unchanged is cheaper, and you can efficiently prove that a given leaf was part of the hash), but I don't know if that's going to be useful for BitTorrent.
Could this potentially make cracking down on illegal content easier?
"How long will it take copyright lawyers find out about v2-only torrents?"