The Cult of DD
eklitzke.org
eklitzke.org
-> links to wikipedia page with direct discription of lineage back to 5th ed research unix
"That weird bs=4M argument in the dd version isn’t actually doing anything special—all it’s doing is instructing the dd command to use a 4 MB buffer size while copying. But who cares? Why not just let the command figure out the right buffer size automatically?"
Um -
a) it is 'doing the special thing' of changing the block size (not buffer size)
b) Because the command probably doesn't figure out the right size automatically, much like your 'cat' example above which also doesn't
c) And this can mean massive performance differences between invocations
> Another reason to prefer the cat variant is that it lets you actually string together a normal shell pipeline. For instance, if you want progress information with cat you can combine it with the pv command
Umm:
dd if=file bs=some-optimal-block-size | rest-of-pipeline
that was hard.>If you want to create a file of a certain size, you can do so using other standard programs like head. For instance, here are two ways to create a 100 MB file containing all zeroes:
$ uname -sr
OpenBSD 6.0
$ head -c 10MB /dev/zero
head: unknown option -- c
usage: head [-count | -n count] [file ...]
well.. guess that wasn't so 'standard' after all..
I must be using some nonstandard version... $ man head |sed -ne 47,51p
HISTORY
The head utility first appeared in 1BSD.
AUTHORS
Bill Joy, August 24, 1977.
$ sed -ne 4p /usr/src/usr.bin/head/head.c
* Copyright (c) 1980, 1987 Regents of the University of California.
Hmm..> So if you find yourself doing that a lot, I won’t blame you for reaching for dd. But otherwise, try to stick to more standard Unix tools.
Like 'pv'?
edit: added formatting, sector size note, head manpage/head.c stuffs.. apologies.
EDIT: fixed man section references
It's not back in the day, it's still true. In Linux, block devices are kernel-cached by default, unless opened with O_DIRECT flag.
In general UNIX case (for example, FreeBSD), they aren't: https://www.freebsd.org/doc/en/books/arch-handbook/driverbas...
So, in FreeBSD "dd bs=1" will fail if it involves any disk device: disk driver will return EINVAL from read(2) or write(2) because I/O size is not divisible by physical sector size. "cat" with buffer size X (which depends on implementation) will work or not depending on divisibility of X by physical sector size, and other random factors, like short file I/O caused by delivery of the signals.
Summary: dd(1) still has its place and author of original article is getting it wrong.
For those wondering Linux uses a unified buffer / page cache so there isn't a coherency issue. The buffer cache entries typically point to the corresponding entry in the page cache if it exists. The biggest reason the two are separate but correlated is that the block size isn't always the same as the page size.
The raw/cooked device thing crossed my mind but thought it would distract from the point-by-point here..
Since coreutils-8.24[1], dd "accepts a new status=progress level to print data transfer statistics on stderr approximately every second."
> Because the command probably doesn't figure out the right size automatically . . . this can mean massive performance differences between invocations
For anyone who's wondering, here are two good threads on determining optimal block size: https://superuser.com/questions/234199/good-block-size-for-d... http://stackoverflow.com/questions/6161823/dd-how-to-calcula...
As an aside: On BSDs (incl. macOS), SIGINFO is also able to be sent interactively by the line driver when you type ^T (like ^C sends SIGINT.) Kind of lame that Linux doesn't follow suit [or even have SIGINFO], or we'd see a lot more programs that build in useful "prod me for an update" hooks, the way they already have "prod me to reload my config" SIGHUP hooks.
(You probably know this as you use OpenBSD, but) something I really like about BSDs is that nost of the core commands respond to ^T with progress info ofsome kin, dd included.
sudo dd .... of=/dev/...
But there's no trivial cat equivalent: sudo cat ... > target
Will open target as your current user anyway. You can play around with tee and redirection of course. But that's getting more complicated than the original. sudo sh -c 'cat some.img > /dev/sdb'
or even more baroque: cat some.img | sudo tee /dev/sdb > /dev/null
is a pain by comparison, and the `sudo sh -c` variant has env implications when spawning a sub-shell.I have an ARM/linux installer script that writes the u-boot image to a specific offset before the first partition:
dd if=${UBOOT_DIR}/MLO of=$LO_DEVICE count=1 seek=1 bs=128k
dd if=${UBOOT_DIR}/u-boot.img of=$LO_DEVICE count=2 seek=1 bs=384k
This is admittedly somewhat esoteric, but it seems like a stretch to say `dd` does not have some place, especially when transferring binary data in very specific ways. :w !sudo tee %
does the trick. (What "w!" does is send the buffer into the given shell command as stdin.) sudo cp image.iso /dev/sdbBasically, you can use cp wherever you use dd, as long as you're not changing any low-level parameters (e.g. starting 500 bytes into the file or something).
cp backup.tar.bz /dev/sda
Nowadays I would know enough to at least get the contents of the backup.tar.bz back. Back then, this was the end of both my / partition (or any other partition) and the backup of my music collection.Still, that didn't end my love affair with Unix. It did make me a whole lot more careful though.
Ouch. That just hurts seeing that line.
sudo (cat ... > target) sudo bashAs far as I'm concerned, dd is lower-level than most of the other utilities and provides more control over what's happening.
The author does have a point that the syntax is strange though.
Use noerror, but forget sync? Corrupt output file if there is an error. Use a bigger bs so it's not slow as treacle? A single faulty sector blows away a whole bs of data, and your output image may get unwanted padding appended to the end. Recoverable error? dd's not going to retry.
Use ddrescue or FreeBSD's recoverdisk(1). They're faster, they're safer, they're more effective, and they're easier to use.
[1]: http://www.gnu.org/software/ddrescue/ddrescue.html [2]: http://www.garloff.de/kurt/linux/ddrescue/
Funnily enough, I ended up using it to accidentally name the wrong drive in the argument, and lost years of photos, music, video etc. though I suppose I can't blame dd for that :)
I now use full paths for destination as well as source.
cat image.iso | pv >/dev/sdb
could be rewritten as pv < image.iso > /dev/sdb
A related mistake is the Useless Use of Echo, since any command of the form echo "foo" | bar
can be written using here strings as bar <<< "foo"
or even bar <<WORD
foo
WORD
[1] http://porkmail.org/era/unix/award.htmlIn the rare case where the volume of data is large enough to make the efficiency hit noticeable on large machines, rewriting a pipeline to eliminate a leading cat makes sense. In all other cases, it is a premature and unnecessary optimization.
<image.iso pv > /dev/sdb pv image.iso >/dev/sdbHuh? pv can cat stuff on it's own, and it will be able to make a progress bar based on the filesize
pv image.img > /dev/sdb pv < image.iso > /dev/sdb
would actually need to buffer the complete file, before commencing to write to the device (and showing the progress), which would defeat the whole idea of showing progress. $ ls -l /proc/self/fd/0 < /tmp/x
lr-x------ 1 user users 64 <date> /proc/self/fd/0 -> /tmp/x <image.iso pv >/dev/sdb //SYSPRINT DD SYSOUT=*
//SYSLIN DD DSN=&&OBJAPBND,
// DISP=(NEW,PASS),SPACE=(TRK,(3,3)),
// DCB=(RECFM=FB,LRECL=80,BLKSIZE=3200),
// UNIT=&SAMPUNIT
//SYSLIB DD DSN=SYS1.MACLIB,DISP=SHR
//SYSIN DD DSN=&SAMPLIB(IEWAPBND),DISP=SHRAh, I miss elements of the mainframe days.
[0] https://www.ibm.com/support/knowledgecenter/zosbasics/com.ib...
Because it can be a lot slower. dd is low level, hence powerful and dangerous.
And, if we are going down that rabbit hole, you don't need cat[1]
“The purpose of cat is to concatenate (or "catenate") files. If it's only one file, concatenating it with nothing at all is a waste of time, and costs you a process.”
(This sounds like a zen koan somehow.)
In a UUOC avoidance case, it's the current process which reads, generally via stdin. Say, the shell, or dd itself with an 'if=' parameter.
Which I strongly suspect you know.
:-D
dd is a tool. dd can do a lot more then cat. dd can count, seek, skip (seek/drop input), and do basic-ish data conversion. dd is standard, even more standard then cat (the GNU breed). I even used it to flip a byte in a binary, a couple of times.
New-ish gnu dd even adds a nice progress display option (standard is sending it sigusr1, since dd is made to be scripted where only the exit code matters).
> Actually, using dd is almost never necessary, and due to its highly nonstandard syntax is usually just an easy way to mess things up.
Personally I never messed it up, nor was confused about it. This sentence also sets the tone of the whole article, a rather subjective tone that is.
edit: Some dd usage examples: http://www.linuxquestions.org/questions/linux-newbie-8/learn...
http://www.linuxquestions.org/linux/answers/Applications_GUI...
But I guess they do...
http://stackoverflow.com/questions/1734243/in-c-how-do-i-pri...
I've seen that kind of brokenness from programs trying to find their binary image on disk. Don't do it, it's bad.
ioctl(STDIN_FILENO, BLKGETSIZE64, &size)It's not the same thing as trying to walk the FS to look for the filename is silly.
pv file.bin | dd of=/dev/sdb
is nice too. Honestly while the OP sparks a nice conversation, it is actually complaining about the use of dd out of lack of knowledge.You wanna format a usb key? Google this, copy/paste these dd instructions, it works, move on with your life.
You wanna format a usb key using something related to cat you once saw and didn't fully understand? Have fun.
Both approaches have their weak points, but in any OS the answer to "How do I format a usb key" should not start with "Oh boy, let's have a Socratic dialog over 10 years on how to do that."
"Why do we do it like that? I dunno, that's how I learned how, how do you do it?"
Embarrassingly, it took me a long time before I started reaching for man pages instead of Google. That has probably has had the biggest effect on tightening up my command line fu.
find is another tool that seems to get only one specific use case that ignores its rather large and useful toolset.
Now, sometimes when people watch me work in a shared session they comment on my "peculiar" (to them) usage of flipping between -h --help and man $command, because there's a whole lot of switches I have memorized over time, but even more that I just have good reference points for.
But, bar none, what I've noticed among my peers is that the people that have always bowed to quick google solutions never really have taken the time to learn what they're doing. They almost always seems to be the 'quick fix', 'get it working now, sort it out later' types.
Also note that there are still unix systems out there which do not support byte-level granularity of access to block devices. On those devices you must actually use a buffer of exactly the size of the blocks on the device. Heck, linux was like this until at least v2.
An essential tool for low level repair, like when you can guess the partition table values but there is no partition table anymore.
dd's "highly nonstandard syntax" comes from the JCL programming language, but it's really just another tool to read and write files. At the end of the day it's not more complex or incompatible than other unix tools. For example, you can also use tools like `pv` with dd no problem to get progress statements.
That's the beauty of Unix.
Everything is a file. Thus every program that can work with files, can in fact work with everything.
It's actually very liberating.
That's the Unix way: The customer, eh, user is always right.
I once wanted to clean up backup files created by emacs (they end in the tilde character) by typing "rm <asterisk>~" - except what I did type was "rm <asterisk> ~".
(On the upside, I learned a valuable lesson that day.)
What happens when you did a sparse file? And cp?
https://wiki.archlinux.org/index.php/sparse_file
C.f. fallocate(1,2)
It better :-)
But it all comes from the unix idea of everything is a file.
2. "UNIX was not designed to stop you from doing stupid things, because that would also stop you from doing clever things."
- Doug Gwyn
$ dd if=/dev/urandom count=1000 bs=1000000 | pv -s 1000000000 > foo
214MiB 0:00:16 [13.1MiB/s] [========================>
] 22% ETA 0:00:55
Compare to ^T: $ dd if=/dev/urandom of=foo count=1000 bs=1000000
load: 1.76 cmd: dd 80097 running 0.00u 0.89s
11+0 records in
11+0 records out
11000000 bytes transferred in 0.947316 secs (11611752 bytes/sec)
load: 1.76 cmd: dd 80097 running 0.00u 1.68s
22+0 records in
22+0 records out
22000000 bytes transferred in 1.746013 secs (12600134 bytes/sec)
load: 1.76 cmd: dd 80097 running 0.00u 2.28s
31+0 records in
31+0 records out
31000000 bytes transferred in 2.392392 secs (12957742 bytes/sec)
load: 1.76 cmd: dd 80097 running 0.00u 2.83s
38+0 records in
38+0 records out
....Linux, not. I wonder why.
But it looks like the answer is just "it was complicated to implement so Linux didn't add it."
SIGINFO works on gnu dd last I tried it.
I certainly agree the syntax of the arguments is strange, due to its age, but I don't agree that learning it is difficult or a waste of time.
All I've learned is that the author doesn't like dd well enough to learn it.
* Using "cat source > target" instead of "cp source target"
* Using "cat source | pv > target" instead of "pv source > target"
* Using "head -c 100MB /dev/zero > target" instead of "truncate -s 100MB target"
If you mess up the syntax on a dd invocation, a nice thing happens: nothing.
Use a shell command and pipes, and your command better be perfect before you hit return.
Usually it's not past the easily-reproducable system partition yet or on a data disk that is backed up regularly so I can recover in an hour or so..
"One of the biggest regrets of my life."
The command line should be reserved for times where you need the fine grain control to do something that DD is meant to do. A GUI should implement everything else in a reliable way that doesn't break half the time or crash on unexprected input.
The better half of my computer use happens in the command line interface, way more efficient use of my time.
Also, I only need to remember one progress command for my entire operating system: control+t. I also get a kernel wait channel from that which is phenomenally pertinent to rapidly understanding and diagnosing what the heck a command is doing or why it is stuck.
I hate what Linux has done to systems software culture.
And if you use dd then you probably should specify a bigger block size than the default of 512 bytes.
But yeah, most usage is obsolete.
so you get the best block size for reads and writes. I can't speak to what the shell does, though.
(dd's corresponding options do not suffer from this problem.)
Grab the first N bytes vs. grab everything starting from the Nth byte.
Also it's a victim of dd's bizarre non-Unix syntax - an option like "--status" or "--progress" would be more in keeping with expectations.
Some kinds of devices are structured such that each write produces a discrete block, with a maximum size (such that any bytes in excess are discarded) and each read reads only from one block, advancing to the next one (such that any unread bytes in the current block due to the buffer being too small are discarded). This is very reminiscent of datagram sockets in the IPC/networking arena. dd was developed as an invaluable tool for "reblocking" data for these kinds of devices.
One point that the blog author doesn't realize (or neglects to comment upon) is that "head -c 100MB" relies on an extension, whereas "dd if=/dev/zero of=image.iso bs=4MB count=25" is ... almost POSIX: there is no MB suffix documented by POSIX, only "b" and "k" (lower case). The operator "x" is in POSIX: bs=4x1024x1024.
Here is a non-useless use of dd to request exactly one byte of input from a TTY in raw mode:
file:///usr/share/doc/bash-doc/examples/scripts/line-input.bash
Wrote that myself, back in 1996; was surprised years later to find it in the Bash distribution.
Though fio is better because it can work in parallel.
cat image.iso | pv >/dev/sdb
just do pv image.iso >/dev/sdb