Comparing dd and cp when writing a file to a thumb drive
distrowatch.com
distrowatch.com
In the case of copying files to a mounted file system, I’ve sometimes found it faster to use a tar pipeline than cp when copying data to an USB stick or SD/microSD card.
Instead of:
cp -r ~/wherever/somedir/ /media/SOMETHING/
I would do cd ~/wherever/
tar cf - somedir/ | ( cd /media/SOMETHING/ && tar xvf - )
And it would be noticably faster.Not the same use case as linked article, but wanted to bring this up since it’s somewhat related.
Just take care that cp, tar and rsync each have slightly different handling of extended attributes and sparse files.
(By the way, I believe "tar -C /path" is the canonical way of doing "cd /path ; tar" without resorting to subshells.)
First thing that makes it so weird is how it assigns different meaning between paths including vs not including trailing slash. Completely different from how most command line tools I am used to behave in Linux and FreeBSD.
That alone is enough to remind me every time I try to use rsync why I don’t like and generally don’t use rsync.
With rsync, destination paths are defined by the command itself and the command is idempotent. I don't usually need to know what the destination is like to form a proper command (though oopsies happen, so verify before using --delete).
With cp/mv - the result depends on the presence and type of the destination. E.g. try running cp or mv, canceling then restarting. Do you need to change the arguments? Why?
mkdir s1 s2 d1
touch s1/s1.txt s2/s2.txt
# this seems inconsistent
cp -r s1 d1 # generates d1/s1/s1.txt
cp -r s2 d2 # generates d2/s2.txt
mkdir s1 s2 d1
touch s1/s1.txt s2/s2.txt
# I don't use this form, but it is consitent
rsync -r s1 d1 # generates d1/s1/s1.txt
rsync -r s2 d2 # generates d2/s2/s2.txt
# Same as above but more explicit
rsync -r s1 d1/ # generates d1/s1/s1.txt
rsync -r s2 d2/ # generates d2/s1/s2.txt
# I prefer this form most of the time:
rsync -r s1/ d1/ # generates d1/s1.txt
rsync -r s2/ d2/ # generates d2/s2.txt
I simply try to use trailing slashes wherever permitted and the result is amply clear.How would you separate "merge this directory into that one and update files with the same name" from "copy this directory into that directory"?
Same principle as rm requiring --no-preserve-root if you actually want to nuke / - right now it's way too easy to accidentally and destructively do the wrong thing.
I would make "copy this directory into that directory" out of scope for the tool.
Let’s imagine a tool similar to rsync, but less confusing, and more in tune with what I want to do, personally. There are for sure a bunch of things that rsync can do, that this imagined tool can’t. That’s fine by me.
Let’s call the tool nsync.
It would work like this:
nsync /some/src/dir /some/dest/dir
And running that would behave exactly the same with or without trailing slashes.
I.e the above and the following would all be equivalent:
nsync /some/src/dir/ /some/dest/dir/
nsync /some/src/dir/ /some/dest/dir
nsync /some/src/dir /some/dest/dir/
And what would this do? It would inspect source and dest dirs. Then it would copy files from source to dest for which last modified was greater in source dir than in dest, or where files in source dir did not exist in dest dir.
In other words, it would overwrite older files that had older last modified time stamp, and it would copy files that did not exist.
Like rsync it would also work with ssh (scp/sftp). Maybe some other protocols too, but only if those other protocols supported the comparisons we need to make. Prefer fewer protocols, and this way of working over trying to be the subset of what works across a gazillion protocols.
If a file exists in dest but not in source dir, it is kept untouched in dest. Not deleted. Not copied back to source dir.
Then there would be one other mode; destructive mode. The flag for it would be -d.
nsync -d /whatever/a/b/c/ /wherever/x/y/z/
This would work similar to the normal mode. But it would remove any files in dest dir that are not present in source dir. Before actually deleting anything it would list all of the files that will be deleted, and ask for keyboard confirmation. [y/N] so that you have to explicitly hit y and then enter. Enter alone will be interpreted as no.
You would be able to override the confirmation with the -y argument.
nsync -dy /whatever/a/b/c/ /wherever/x/y/z/
And that’s it. That’s what I would want rsync to be for me.
There probably are some programs that behave exactly like this. I’ll eventually write one too. It’ll have a user base of 1. Me.
https://github.com/chapmanjacobd/journal/blob/main/programmi...
the only strange thing about rsync is that it follows _BSD_ syntax
https://wiki.archlinux.org/title/rsync#Trailing_slash_caveat
Indeed, it does!
cd "$( mktemp -d )"
mkdir a b c
touch a/f1 a/f2 a/f3
cp -r a/ b/
ls b
Resulting content of directory b on one my FreeBSD machines f1 f2 f3
Guess I usually1. Don’t use cp -r so often, and
2. When I do use cp -r apparently I don’t put a trailing slash
Cause I only ever had problems when trying to use rsync, not when using cp or mv :S
Would be happy to learn the real reason.
With a tarpipe, you can block on two IOPs at a time, and they're decoupled by the pipe buffer.
This primarily makes a difference because the kernel cannot issue the I/O for small files ahead of time, like it does when you sequentially read a large file, so you do actually end up blocking and waiting.
My guess: speed up is due to buffering involved in the latter's case, not sure though.
In general writing to disk is handled asynchronously by the kernel (`write` just copies to a buffer and returns), but metadata operations like creating files are not, so this should help the most for many small files.
In any case, the tests in ioblksize.h indicate that bs=4M is far too large and may perform worse than the default for cp/cat (128KiB). There is a script there that should clear things up for more modern systems.
The point about fdatasync is superfluous as you can run 'sync' yourself, or unmount the filesystem.
dd uses a default block size of 512 bytes, according to the manual page: calling write(2) directly means you need to choose a buffer of some size.
"bs=" sets both input and output block sizes, which probably isn't the best idea in this case.
Block sizes are a tricky subject:
https://utcc.utoronto.ca/~cks/space/blog/unix/StatfsPeculiar...
https://utcc.utoronto.ca/~cks/space/blog/tech/SSDsAnd4KSecto...
Thank you and everyone else in this thread for the know-how and references! I would like to eventually write an article (or at least heavily commented source) to hopefully explain all this nonsense for other people like me.
("Eject" is macOS's term for "unmount filesystem")
"eject" is a Linux command which attempts to safely disconnect removable media.
https://manpages.ubuntu.com/manpages/noble/en/man1/eject.1.h...
TFA doesn't mention filesystems at all, sort of jumps in where we find the block device. Things could become messy if the device were mounted while the copy is attempted.
<infile pv >outfile
And it seems to work very well plus gives a good progress indicator.I'm all for anything not 'dd' or cargo-culted like balenaEtcher
This all depends on a certain format of ISO that I can't recall
There's no magic. Checksum the drive and the file after
Using cp seems overkill? It's really really designed to copy files between file systems and has a lot of logic to handle different cases and different optimisations which don't matter when writing to block devices.
As you say, there's no magic. I want a program which just calls 'open' on a path I give it and then uses 'write' to write my data to it. cp does so, so much more related to file systems, shell redirects aren't a separate program I can run with sudo, but I can trust dd to do the job. And it has a (kinds bad) progress monitor to boot.
The default 512 byte block size is unfortunate though.
It can be "cat", which is often cargo culted in it own right. Or even "dd", which can work with stdin/out.
I usually prefer not to use shell redirects if there is a command that take filenames. That's because it gives more control to the app. The app knows how the file will be used, the shell doesn't, so it can open it the most appropriate way, avoid overwriting the output file if something goes wrong, output better error messages, etc... Now, if you don't trust the app (for example if you fear it will modify your input file), then shell redirects may be the better option.
With bash you could maybe use the non-standard `read -N` option but I'm not sure about null bytes.
It can be really quit surprisingly fast too, 4GB/sec reading /dev/zero writing /dev/null. ( yes 40gbit/sec, binary copies, in a shell script!!! ). 2-3GB copying real data.
Oh boy but the edge cases and quirks.. there are so many reasons why this is a bad way to do things..
Curiously, I double the speed bu unsetting and recreating the buffer variable for each op. There also seems to be cases where ksh will buffer reads from a pipe to allow limited seeks, but also if you do a read -N from a pipe with too large a size (>8k I think, I'd have to check) and that read can't be completed because the source finished writing less than that, then that data is gone. Less than 8k, you can still read it with a subsequent read -n.
Probably the most actually useful thing I learned was that by ksh93 creates its pipes with unix sockets, not pipes, which means the buffer size is set by /proc/sys/net/core/[w|r]mem_default, rather than a fcntl call. That makes it easier to tweak from a script, and also makes ksh pipes faster by default for many streams, compared to most other shells ( depending on block sizes, and the difference goes away if you tweak up the pipe buffer size )
Don't get me wrong, not something I'd use in real life, but it was fun anyway.
while LANG=C IFS= read -d '' -r -n 1 x ;do printf '%c' "$x" ;done <dat1 >dat2
At least ksh and zsh have equivalent functionality though the details may differ.
Not dash or ancient sh.
I don't pretend it's practical, just possible.
I actually do have a neck beard this month.
while LANG=C IFS= read -d'' -r -N 1 x ;do printf '%c' "$x" ;done <dat1 >dat2
With -n it handles 0x00 fine but screws up on 0x0A !
Use read in a loop, with special care with LANG and IFS to make all bytes meaningless. Except there is no way to avoid null being special, but you can handle null by making null the delimiter for read, and only reading one byte at a time. So even though you can't store an actual null in a variable, you can still detect that there was a null and print a new one back out, and since you only read one byte at a time, you do that for each individual input byte and strings of nulls are not collapsed.
It looks like a lot, but, read is a builtin, and at least in bash and ksh and zsh so is printf, and although this is a loop, it's actually not even a sub-shell. If you edit variables inside the loop, they are still there after the loop, ie, you never forked a child.
while LANG=C IFS= read -d '' -r -n 1 x ;do printf '%c' "$x" ;done <junk1.rnd >junk2.rnd