> The only advantage of uv is to have support for parallel async extraction.
This isn't true, the biggest speed ups are for the warm cases where we've already unpacked the files into the cache, as notatallshaw mentions above. We can also make low-level optimizations during resolution, e.g., in version parsing, that are not possible in pure Python code.
> If pip extracted multiple files/wheels in parallel without being blocked by the GIL, pip could easily match or outcompete uv.
I'm a bit confused by these claims about the GIL? The expensive IO operations release the GIL.
> When the user edits one file, all files are modified simultaneously
This is why we default to reflinks or copy-on-write semantics when creating environments, not all file systems support it but it's becoming more common.
> Hypothetically, a simple deduplication of binary files (.dll .so) should achieve 50% of the savings without significant drawbacks
We also explored this (see https://github.com/astral-sh/uv/pull/19694) and the linked pull request has a table comparing to this strategy.
> We have optimization for empty files (0 bytes) because there is nothing to write and checksum.
Interesting, I would be very surprised if this made a significant difference? but I'll take a look.