Show HN: fcp – A significantly faster alternative to cp(1), written in Rust
github.com
github.com
Feel free to ask me any questions about the project!
On macOS, this is augmented by regular files being copied using the OS's fclonefileat and fcopyfile mechanisms under the hood, which allows even very large files to be copied near-instantly (when compared to copying them block-by-block as classic cp does).
[1] https://docs.brew.sh/Acceptable-Formulae#niche-or-self-submi...
Traditionally the reason to use threads in small-requests i/o is to get multiple outstanding requests in flight in the hardware io stack which some hardware can use to get you more IOPS throughput. (They might end up hitting different disks in RAID, or taking advantage SSD internal parallelism, or whatever tricks spinning disks use to get more throughput with TCQ)
(Re your MacOS remark, Linux has support for copy-on-write file copies as well, see reflink and FICLONE / copy_file_range)
I would assume that is the case here too. The threads are likely hiding the latency of the IOPS. I can't see any CPU load statistics, so it is hard to say, if it is indeed CPU bound or not.
A quick run with time shows user+sys being only slightly higher than real, so with 6 threads for my 6 cpus, it uses a single cpu. (and practically no time in user-space)
CPU cores can't read/write files, those are IO operations by the storage device. The CPU could have 64 cores but on a system with a single hard disk that can do just one concurrent operation.
So, perhaps, take a look at https://github.com/rust-cli/man or something
(I agree that a real manual is necessary too. Maybe you could convert one from the online docs?)
For hard drives this is the opposite of what you want. For hard drives, especially when copying large trees, you only want one thread to access the disk, and preferably sort its stat()s and open/read/close() by inode.
Also this tool seems to lack most options that cp(1) has.
A better approach would be to use information about location of data extents. Luckily there's already a crate that can help with that: https://github.com/the8472/platter-walk
And if you only need the `d_type` instead of the full stat struct, then you get that for free from getdents(2) on some filesystems, this is what platter_walk does.
OK, this is something that's bugged me for a long time. I've been in the field forever, but working on Windows. So I've seen this kind of thing many times before, particularly when referencing man, and it seems completely cryptic.
The way linux people occasionally put a parenthetic number after a generic command seems to me one of the ways that normal civilians, and even pros like me, are intimidated into staying away from serious linux usage.
So what gives? What the heck is the parenthetic number for?
Roughly, 1 is unprivileged shell commands, 2 is syscalls, 3 is library functions, and 8 is administrative commands.
It’s night and day working with people who know how to be approachable in development, but also realize you need to understand some basic flags for everyday tools and not turn to Google or Stack Overflow for what is directly in your manpages.
Admittedly those who cannot stay approachable concoct all sorts of shell hell for everyone, even those who are experienced.
Excellent question, by the way. I wish we’d get some more candid questions like these on HN when relevant to the thread, like this one. It’s a learning opportunity for all readers.
Stack combines the reference (which flags) with the expertise (of educated guesses by sometimes wiser users).
I would never trust a man page over a good Stack answer. (edit: as a web dev who uses *nix all day, but has no desire to become a full fledged sysadmin... Just give me the tool I need to get the job done without sacrificing performance or security. I don't need to know every possible gnu tool and flag and distro dependent options...)
But I have a bigger issue with the fact that developers who use Stack Overflow are more often than not just going to copy and paste and not understand what they’re doing.
And that’s predominantly how the site is used.
Yeah, there are plenty of manpages that are verbose, but frankly a lot of them are easier to parse than reading hacked together Stack Overflow answers that don’t explain how the poster arrived at their answer.
They literally don’t show their work.
As for lacking most of cp(1)'s options, that's a deliberate choice. My goal with fcp is to make the most common use case of cp(1) (i.e. copy some files/directories with no options) fast, not to replace cp(1) completely.
On a SATA hard drive using ext4, I copied a directory with a large hierarchy of file and 216MB overall. After 2 warm-ups, I launch this simple benchmark.
time cp -r moodle X
13.76s s (0.06s u + 2.44s k) 18% CPU
time fcp-0.1.0-x86_64-unknown-linux-gnu moodle X
46.12s s (0.15s u + 2.47s k) 3% CPU
I emptied linux cache before each run with `echo 1 > /proc/sys/vm/drop_caches`, but this does not remove the HDD buffering. I suspect that fcp would be worse in real life conditions (no warm-up).This does not mean the fcp is bad or not performant. But it some cases, it can be 3x slower.
One copy operation is the little files and deep hierarchy. The other copy operation is just a big file. Both operations start at same time and use the same tool.
I'm curious what cp and fcp do to the cache as they operate.
Does it mean you are doing this operation enough times to trick the kernel to cache everything in RAM? Or you run it 3 times and chose the fastest timings?
I think there are sometimes arguments to be made that, for the user, that is run time. When using benchmarks to compare programs like this, though, I think isolating the "actual" (i.e. useful) program execution is productive.
Second run: fcp binary is cached by the OS in RAM so it starts up faster.
Third run: just to make sure :)
This is just one of the things of a type of cache that can be warmed up.
I learned this the hard way when dismantling and abusing a dead HDD. Had to vacuum all the tiny shards from the floor.
Great old video. Thanks for sharing. Terrifying goose though.
Btw, copying files is mostly bound by disk speed (or network if share), not so much by CPU, so it doesn't seem to be a valid reason to rewrite it in Rust for cpu performance...
i don't think it was ever meant to be "exactly GNU/coreutils grep", and I don't think it ever even tried. The API is totally different, i don't think it's ever been the same, it scratches a similar-but-not-identical itch.
https://github.com/BurntSushi/ripgrep/blob/master/FAQ.md#pos...
I definitely can understand the fatigue with "rewrite it in Rust" though, as there is a lot of that.
Something you might be interested in is that people have actually gone through the effort of rewriting the actual coreutils to spec in Rust (https://github.com/uutils/coreutils).
[1] https://lemire.me/blog/my-sayings/
Don't take my word for it, Prof. Lemire is the expert on performance.
echo 1 > /sys/devices/system/cpu/intel_pstate/no_turbo # intel
echo 0 > /sys/devices/system/cpu/cpufreq/boost # amd
together with disabling noisy background services that's the most significant reduction in noise. There are further steps that can be taken but for CPU-bound microbenchmarks they only provide incremental improvements in my experience.IO benchmarks are trickier. SSDs have thermal throttling too, there's the OS write cache, the SLC cache performance cliff and other factors.
On full load, any machine hits thermal constraints, so, considering clock changes a problem, a desktop is problematic as well. A desktop will certainly have a higher ceiling, and/or a laptop may have too much of a low one, but it's incorrect to assume that the problem is inherent to laptops.
The solution is therefore, which is canonical in serious benchmarks and it's not mentioned, is to ensure that the machine runs on a fixed clock rate. This is easily accomplished, at least on desktops, by changing the BIOS. I wasn't personally able to do this from the O/S (Linux).
Using a machine without control on the clock (which a cloud one can be, even if bare metal) is subject to generate inaccurate benchmark results.
https://www.kernel.org/doc/html/v4.12/admin-guide/pm/intel_p...
E.g., if you're using the Intel pstate driver in active mode, you can disable turbo via:
echo "1" | sudo tee /sys/devices/system/cpu/intel_pstate/no_turbo
(and re-enable by writing "0".)GNU Coreutils' cp, the thing you use on most GNU/Linux systems, is not the "Classic Unix cp".
It has options for copy-on-write cloning via kernel-specific methods.
COW copying is opt-in probably because it's not a true copy. If either file sustains damage, both are toast.
summary: this is a nice cow/cloning wrapper for use on a single ssd partition!
-- but cloning not copying between volumes or partitions is .. and when you actually want to copy with fcp, it can be very slow as detailed in other comments.
I want fragmentation to be as little as possible so it wouldn't be too hard to manually recover a file if I delete something accidentally or damage the file system.
Those layers may conceal a real problem (if there is one) so that you can't see it even though it exists, or they can create problems, which you will then mistakenly blame on the "hard drive controller" even if it was not at fault.
So you've created a situation where you can at best confuse yourself into not understanding what's actually going on, and at worst satisfy yourself nothing is wrong when actually there's a catastrophe awaiting you. The worst of all worlds.
ln -s /path/to/source /path/to/destThat said, neither of these are truly the same as an edit to one will impact the other.
i.e. btrfs snapshots use this feature.
When copying between different filesystems it still makes sense to add the --sparse=auto flag to keep a copy of a sparse file sparse, too.
My alias for cp is:
cp -i --reflink=auto --sparse=auto