The cult of dd (2017)
eklitzke.org
eklitzke.org
Well, because it doesn't. At least on the Linux version I used in the past, it defaulted to 512 byte blocks or something similarly small and it commits every block individually, leading to really slow performance with USB sticks with controllers not smart enough to bundle writes together. I wouldn't be surprised if that incurs some heavy write amplification at flash level too. Perhaps it's smart enough to figure a better block size now but this is where that habit comes from.
Another thing, creating sparse files with the seek option (simply put files containing zeros that are not actually written to the disk nor taking space, but do turn up all these zeroes when you read the file). Also something not duplicated with cat or head.
What I like about dd is that it can do pretty much all disk operations in one simple tool. Definitely worth learning all of it IMO.
i think he's talking about cat being the command that figures it out automatically
I wonder what cat does in terms of buffers, I kinda doubt it has any special optimisations though I would guess the shell redirect might have. As that's really the thing doing the work there, not cat. Edit: Nope, I'm wrong there!!
Also that command does more than just specify the memory buffer like he says. That's my point, it's useful for tuning which can be super helpful with huge images.
It can also lead to some dangerous gotchas as well when working with files. But with full disk images these don't apply generally.
No? The shell redirection is just
int tmp_fd = open("/dev/sdb", O_WRONLY|O_TRUNC, 0666);
dup2(tmp_fd, STDOUT_FILENO);
close(tmp_fd);
(plus error handling and whatnot)`cat` ends up with a file descriptor directly to the block device, same as `dd` does; the only difference is whether the `open()` call comes before or after the `execve()` call.
I really doubt cat is smart enough to figure out a suitable block size though. At least with dd you can specify one.
It seems to try, at least
https://github.com/coreutils/coreutils/blob/master/src/cat.c...
- >=9.1 (2022-04-15) : Use the `copy_file_range()` syscall and let the kernel figure it out
- >=8.23 (2014-07-18) : max(128KiB, st_blksize(infile), st_blksize(outfile))
- >=8.17 (2012-05-10) : max(64KiB, st_blksize(infile), st_blksize(outfile))
- >=7.2 (2009-03-31) : max(32KiB, st_blksize(infile), st_blksize(outfile))
- at least as far back as 1996 : max(st_blksize(infile), st_blksize(outfile))
(In my psuedo-code, `st_blksize(fd)` is the `ST_BLKSIZE(buf)` of the result of `fstat(fd, buf)`.)
I recall that GNU cat used a constant value that was tuned for read performance on common hard drives, giving no consideration to write performance. Looking at the GNU cat sources today, it doesn't look at all how I remember; I'd have to study it a bit to tell you what it does.
Edit: Hrmm, it seems I was mis-remembering. Perhaps I was thinking of the minimum values (32KiB, 64KiB, 128KiB below)? Or perhaps I was thinking of a BSD cat? Anyway:
How GNU cat sizes its buffers, by version:
- >=9.1 (2022-04-15) : Use the `copy_file_range()` syscall and let the kernel figure it out
- >=8.23 (2014-07-18) : max(128KiB, st_blksize(infile), st_blksize(outfile))
- >=8.17 (2012-05-10) : max(64KiB, st_blksize(infile), st_blksize(outfile))
- >=7.2 (2009-03-31) : max(32KiB, st_blksize(infile), st_blksize(outfile))
- at least as far back as 1996 : max(st_blksize(infile), st_blksize(outfile))
(In my psuedo-code, `st_blksize(fd)` is the `ST_BLKSIZE(buf)` of the result of `fstat(fd, buf)`.)
Not sure what cat does but like I said in my other comment its not really cat itself that does the disk writing in that command but rather the shell redirect. Edit: Nope, I'm wrong there!!
I interpreted "the command" in the original article as "the command you end up running, whether it be `dd` or `cat` or something else."
But you're mistaken about the shell redirect. In either scenario it's the command (`dd` or `cat`) making the write() syscall to the device. The shell passes the file descriptor to the command, then gets out of the way.
In that case I guess dd does call a sync() on every output block? Because it's definitely slower and the LED pattern on a USB stick is also much more 'flashy' when using 512 bytes.
Only if you tell `dd` `oflag=sync` or `oflag=dsync`.
https://git.savannah.gnu.org/cgit/coreutils.git/tree/src/iob...
Unfortunately dd has not just footguns but foot cannons that are amplified by the mistakes people often make with string escaping, odd file names, loops, null checking, and conditionals in bash.
It uses a syscall per block write (which would slow it down if you use 512 byte blocks instead of 8M for example), but the OS does the file buffering and final writing to the device, unless you pass the fsync or fdatasync options to dd.
Edit:
here's writing to a old and slow stick and you can see that dd is fast and then the OS has to transfer the buffered data onto the stick:
dd if=some-2GB-data status=progress of=/dev/sdX bs=8M
1801256960 bytes (1,8 GB, 1,7 GiB) copied, 7 s, 257 MB/s
0+61631 records in
0+61631 records out
2019556352 bytes (2,0 GB, 1,9 GiB) copied, 378,59 s, 5,3 MB/s
And the stick is placed in a USB 3.2. port on a fast machine ;-0Correct, and that's one more point for dd compared to head/tail (which are fine commands by themselves).
But wouldn't help much in my example, where I used an (very) old "high speed usb 2.0 stick" with 4 MByte/s write speed to demonstrate the difference between buffering and actual writing.
"Who cares?" --> people that care are people that want to have control over that option, as not every tool is written intelligently.
Suppose you have 2 hours of UHD RGBA32 video, and need 5 minutes of footage from the halfway mark:
dd if=in.raw bs=$(( 3840 * 2160 * 4 * 24 )) skip=3600 count=300 | ffmpeg ...
This will be a lot faster than pointlessly catting the first 3 terabytes!Here's a variation. Suppose the video file has a 1M header you need to skip:
{ dd bs=1M skip=1 count=0; dd bs=796262400 skip=3600 count=300; } < in.raw | ffmpeg ...
The first dd invocation does nothing more than seek stdin ahead 1M, so the second can operate on full 1 second chunks of video. Useful!Now, are there Uncalled For Uses of dd, just like there are Useless Uses of cat? You bet.
tail -c +STARTOFFSET $FILE | head -c MAXLENGTH
There are more than one way to do stuff with UNIX tools.
But tail might read up to pipe buffer size more data than is actually required. Also tail + head approach have an overhead of copying data between processes.
You could seek to second, keyframe, etc, and it would continue to work for formats that don't have fixed frame sizes.
It's true that it is quite a lifehack if you often seek to frame in raw though.
ffmpeg -f rawvideo -r ntsc -s 1920x080 -ss 3:00 -i some_file -t 5:00 output_fileBack in 1979 I was an intern at NBS (now called NIST), in their computing standards division. Amongst other machines, we had a couple PDP-11 systems running v6 Unix (it was glorious).
The smaller system had removable 5MB disk packs, and we did periodic backups with dd to a spare disk. It was horribly, horribly slow, and we lived with it for months until someone realized that
dd bs=10 ...
wasn't copying ten blocks at a time, it was copying ten bytes at a time. Whereupon backups got LOTS snappier.(That was a fun position. One of the resident Unix gurus handed me a copy of Lion's notes on my first day and basically said, "Read this and see me later", and I got to maintain and build out a whole lab full of microcomputers).
# My partition table is in sector 0 and my Linux partition starts at sector 62500
dd if=random-linux-image.img of=/dev/mmcblk0 bs=512 count=62499 seek=1 skip=1
How do you make "cat" seek 512 bytes into its stdout before writing? You can't. And even if you used a temporary file, I think the above is better than a pipeline of "head" and "tail" with uses of $(()) to convert from sector to byte units.
Once you have learned to use the tool in such situations a few times, you mentally start to associate it with "disk operations". So you start to use it also in simpler situations where simpler tools would be sufficient. You start teaching its use to others. I don't think it's that strange, or bad in any way.
sudo dd if=/dev/sda of=/tmp/boot.1 bs=512 count=1
sudo head -c 512 /dev/sda > /tmp/boot.2
md5sum /tmp/boot.1
06d6f2aa3e7c33ad06282f70aa4e133b /tmp/boot.1
md5sum /tmp/boot.2
06d6f2aa3e7c33ad06282f70aa4e133b /tmp/boot.2
> dd if=random-linux-image.img of=/dev/mmcblk0 bs=512 count=62499 seek=1 skip=1
is not the same as:
> sudo dd if=/dev/sda of=/tmp/boot.1 bs=512 count=1
You are missing the "seek" and "skip" arguments which the GP challenges you to make `head` do.
sudo head -c 1024 /dev/sda | tail -zc 512 > /tmp/boot.2
Read 1024, keep last 512
sudo tail -zc +513 /dev/sda | head -c 512 > /tmp/boot.2
I use dd when redirect doesn't have permission.
This is very common with u-boot and eMMC/microSD devices on SBC (pi clones, etc).
It’s also generally valuable to be precise and explicit when messing with boot loader data in general.
I think the author is under appreciating these use cases.
works, fyi
> Dd is super useful for incrementally writing boot loader to /dev device endpoints when offsets are required (multiple images).
How exactly does one use cat to overwrite only a specific offset? Say, I have a 512 bootloader that needs to be placed starting at 2048 bytes into the image; how would I invoke cat to do that?
The ability to specify that an exact amount of bytes should be written or read in a location is an important task. More so on special files, like device files. Even provisioning an SD card with the image block aligned is crucial, and that is not speaking of very early boot loaders that need to be available in a specific byte offset.
This discussion hopefully educates the public that this kind of tools are foundational and their use cases should be learned:
You never know when you need to read a special byte offset in a hacky one liner in a debug environment. Analysis productivity matters, and knowing analysis tools is key for that.
The proposed alternative is to use the program "pv". In my opinion, the program "pv" deserves the epithet "obscure", not dd.
I have been using Linux for decades, but the program "pv" has not been installed on any of the systems that I have used.
On the other hand, dd is a part of coreutils, so it is always available.
I frequently use dd when writing on raw devices precisely for the progress option, because the images written may have sizes from tens of GB to several TB, so the duration of the operation is not negligible, sometimes it may take hours, because on most SSDs the writing speeds drop to quite low values for multi-GB data sizes.
Moreover, since SSDs have replaced HDDs, the size of the write buffer has become much more important than before. There are many SSDs where certain buffer sizes can increase the writing speed a lot, typically when the buffer size matches the size of the erase block, which might be of 128 kB or 64 kB, or of another similarly large size. The right buffer size must usually be determined by experiments, so there is no chance that a standard copying program will choose it by default.
Incidentally, we get another Useless Use Of `Cat something | pv > outfile`, instead of `pv < infile > outfile` (or even `cat < infile > outfile`, let the shell do its job ffs)
Why bother memorizing when you can man dd?
Proud member of the cult! I've used dd for all kinds of things:
- Extract the magic bits from a file
- Copy the disk MBR
- Make a disk image with a specific block size and confirm it wrote correctly
- Create loopback disk images
- Over-format floppy disks for Linux distros that need every extra byte
- "Securely" delete files (overwrite exact file size with random junk)
- Copy data more efficiently by changing block size
Some of dd's useful functions: - format fixed-length records from newline-separated input
- transform lowercase to uppercase and vice-versa
- create sparse files
- don't truncate data
- don't stop processing on errors
- skip and seek in files, input and output, separatelyI've seen a lot of hate for dd lately and I don't really understand why, except that maybe people are getting hung up on the nonstandard arg format and unaware of the tool's versatility. I'm not convinced that cat or tail are better (or even as good) for the examples listed in TFA, and if I couldn't remember something as easy as "status=progress" there's no way I'm going to memorize those cat and head pipe contraptions.
dd if=file1 of=file2 bs=1 skip=p seek=q count=n conv=notrunc
Very useful.If you don't need seek then you can at least use ibs=1 instead, as skip/count operate on input blocks, but this will still read one byte at a time even though it will aggregate the reads into larger output blocks.
It would be really nice if we had a dd2 tool that offered similar options but defaulted the block size to a new "auto" value, treated seek/skip as byte values, and treated count as a byte value if the input block size is "auto".
dd allows you to specify your seek, skip and count values as bytes instead of blocks using the iflag/oflag options seek_bytes, skip_bytes and count_bytes.
So to read the first MB of data 100GB into a file you can:
> dd if=/tmp/file1 bs=2M skip=100G count=1M iflag=count_bytes,skip_bytes
It is also in Coreutils 8.22 in RHEL7.
Edit: https://man7.org/linux/man-pages/man1/dd.1.html
Edit 2 : I just found this in the coreutils 9.0 changelog:
dd now counts bytes instead of blocks if a block count ends in "B".
For example, 'dd count=100KiB' now copies 100 KiB of data, not
102,400 blocks of data. The flags count_bytes, skip_bytes and
seek_bytes are therefore obsolescent and are no longer documented,
though they still work.`dd` has purposes, so the author must not do anything real:
0. Clearing volume tables
1. Creating files of fixed sizes
2. Varying block sizes
3. Partial reads
4. Unaligned reads
5. Differing block sizes
6. Character set translation
If coreutils' `cp`, `mv`, `dd`, and their friends were modified to support `sendfile(2)` on Linux, these would then use `splice(2)` zero-copy kernel transfers under-the-hood. Already, they're likely hitting caching at some layer but there are advantages when copying between 10+ GbE network devices and/or tmpfs. Furthermore, the major shells `bash`, `zsh`, `fish`, `busybox`, `dash`, and `tcsh` would be advised to use `sendfile(2)` where possible. There's really no need to play bucket-brigade "fire drill" with data when there's a firehose that's free.
For bulk copying of /dev/zero to /dev/null, `dd` is fastest (block sizes of 256kB up to 4MB are about the same speed), then `pv` (block sizes of 256kB up to 4MB are about the same speed), then `head`.
But it's true, unless you're dealing with an actual tape drive, there aren't a lot of things in the world anymore that depend on a specific block size; the idea that you'll get a faster transfer if you (say) exactly match the erase block size of your SD card doesn't seem to hold much water anymore.
And debian's bash-completion of `dd` has been broken for several releases, I despair of it ever working right again.
It feel like this is gonna lead to another thing like thefuck...
Doesn't `status=progress` work? (The "obscure option to GNU dd" that the article mentions)
No but if you use a much lower block size with cheap flash it does get horribly slow.
Makes sense because cheap eMMC does not have cache RAM to combine write commands and no NCQ (native command queueing) either. So it has to execute each write as it gets it. I bet you can kill flash pretty quickly this way with a block size of 1. The write amplification will be huge.
The problem with the cat method is that you don't really know what it's doing under the hood in terms of write sizes. Probably it will pick something smart but it depends on the OS there and possibly the shell.
Sure sign that the post wasn't researched at all, since dd is one of the very oldest UNIX utilities and is even in the POSIX standard.
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/d...
The "pv" command they recommend instead is, by contrast, not standard and is nowhere near universally available. Also, as just about any not-a-noob knows, the author's "cat" suggestion will not work on devices that expect whole blocks, except perhaps by accident ... and it's not a good habit to rely on accidents. Options like skip and seek also allow you to do things like write a single block at an arbitrary location, which can be very useful sometimes and is something cat/head/etc. can't do.
Reacting to anything one doesn't understand, without even trying to understand it, as strange or obscure or weird (all their words) is a very deeply incurious attitude. That kind of crap doesn't even belong here, per guidelines.
P.S. Hey kid, learn how to use a proper link instead of a Wikipedia title.
$ which pv
pv not found
I've been using Unixes of various flavours since the 1980s and had never heard of it until I read this article. It's certainly more obscure to me – as in not discoverable at all – than dd status=progress.Apparently I've been slumming it by watching long jobs using Activity Monitor on macOS.
dd is pretty standard if you know it. Just because it isn't a standard to you, doesn't mean it isn't for everyone.
There are a lot of commands I don't use regularly that I saw once, which appeared strange to me. There were different way to do what tey did. Yet I didn't feel compelled to write a salty blog post about it.
2) (GNU) dd has support for progress bar and even if not enabled, one can get statistics by sending USR1 signal to it.
And the BSDs have siginfo, so at least most of the Free OSs are covered.
The Linux kernel doesn't have SIGINFO. GNU dd uses SIGINFO on platforms that have it, or SIGUSR1 otherwise. The default action for SIGUSR1 is to kill the process. So it makes sense that on platforms that do have SIGINFO no one would bother to override that default SIGUSR1 behavior.
load: 0.15 cmd: sleep 52109 [nanslp] 0.27r 0.00u 0.00s 0% 2132k
mi_switch+0xc2 sleepq_catch_signals+0x2e6 sleepq_timedwait_sig+0x12 _sleep+0x1d1 kern_clock_nanosleep+0x1c1 sys_nanosleep+0x3b amd64_syscall+0x10c fast_syscall_common+0xf8
SIGUSR1's default is to terminate the process, which makes it awkward to use as a SIGINFO alternative on platforms without it.The OP knows that. The article says that with pv, you have something that works for every command instead of needing to remember dd-specific syntax.
One of the intended use cases of cat is reassembling files after they have been split with the "split" command!
Perhaps it "survives" because people have been using it for decades for specific tasks for which it still still works just fine. Just because there are other ways to achieve the same results for some of the tool's use cases doesn't mean that said tool should be done away with. If you don't like it no one is forcing you to use it, and calling those who still use dd "The Cult of Dd" is just ridiculous, in my opinion.
ln only gets confusing because the file that exists is the target of the link, so it may be natural to think of the link as going from the last name on the command line to the earlier ones, but really the mental model should be in names that either exist or are to be created.
The `help` text for GNU coreutils `cp` starts with `Usage: cp [OPTION]... [-T] SOURCE DEST`.
The `help` text for GNU coreutils `ln` starts with `Usage: ln [OPTION]... [-T] TARGET LINK_NAME`.
The man pages & help text aren't conducive to building the `cmd old new` mental model. `dd` makes it more explicit.
cat id_rsa.pub | ssh $host 'dd of=.ssh/authorized_keys oflag=append conv=notrunc' # Cat version with progress meter
cat image.iso | pv >/dev/sdb
The progress meter here is mostly meaningless. You'll see the initial progress go very quickly (because you're writing to in-memory cache), and once it's "done", you'll have to wait some amount of additional time for a final `sync` to complete (and if you forget to do that, you might remove the drive while writes are still in progress).The best way to write an image to a drive is like so:
dd if=foo.iso of=/dev/sdx bs=4M oflag=dsync status=progress
`oflag=dsync` bypasses the write cache, so your progress bar is actually meaningful. It also guarantees that writes are actually completed[1]. Yes, that 4M block size may be improved by manual tweaking, and it would be nice if that happened automatically. I'm sure tools to do this exist, but they're not installed ubiquitously by default. Older versions of dd do not support `status=progress`, and as a workaround you can do: pv foo.iso | dd of=/dev/sdx bs=4M oflag=dsync
(alternatively, you can set up a bash for-loop that periodically sends SIGUSR1 to dd)[1] Unless the drive has onboard dram cache etc., but this is rare for removable media
P.S. If you use "/dev/sdx" in example commands, it will fail when someone blindly copy-pastes without reading anything, instead of erasing their whole OS
pv <image.iso >/dev/sdbdd is one of the most standard Unix tools around. It's been a core part of Unix for 40 years. It has its issues, but "non-standard" isn't one of them.
pv some.iso >/dev/sdb
Or if you don't care about progress bar, you can use cp: cp some.iso /dev/sdb < some.iso > /dev/sdb % cat > foo
fred
barney
wilma
betty
% < foo > bar
% cat bar
fred
barney
wilma
betty
% diff foo bar
%I just happened to explore this on an M1, so I had a shell that would do things this way.
Update: Back at the computer now, and I cannot replicate it with bash:
rascul@smarts:~/mm> cat > foo
fred
barney
wilma
betty
rascul@smarts:~/mm> < foo > bar
rascul@smarts:~/mm> cat bar
rascul@smarts:~/mm>
I'm curious what shell you did that with.The zsh way seems more properly fitting with the philosophy of the Unix shell to me. It would be an uphill battle getting everyone else to change that behavior, though.
blah | sudo dd of=/some/file
where: blah | sudo cat > /some/file
wouldn't work. blah | sudo tee /some/file > /dev/null
would though. Both are probably fine.blah | sudo sh -c 'cat > /some/file'
Eg:
Dd if=/dev/zero of=image.img bs=1k count=1 seek=999999
Will give you a 1gb file of zeros that takes up 1kb of disk space until you start to write to it, at which point it will transparently “expand on write”.
truncate -s 1g image.img
Or fallocate if you want it to complete instantly without disk writes but still allocate the space.cat can't do that.
> dd if=/tmp/file1 bs=1 skip=50000000 count=10000000
or you can use any block size you want and treat skip and count as byte counts, not block counts.
> dd if=/tmp/file1 bs=4M skip=50M count=10M iflag=skip_bytes,count_bytes
Option 1 will run at Kilobytes per second as it is transferring 1 byte at a time (my test gives me 933kB/s)
Option 2 will run a 100's MB per second as it is transferring 4MB at a time (my test gives me 202MB/s).
Edit: What I should have said at the top of the post was "If you want to extract data from a file whose size is not an integer block multiple". In the above example you could have used a Blocksize of 10MB for the same result. You cannot do this for an oddly size extraction (say 1234567bytes).
Why doesn't the man page say anything about these flags? They are completely obscure becasue of that.
One other nuggets of wisdom which is very poorly documented in the man page: if you have a long running dd process (I often have 12+ hr dd processes reading 10+TB tapes) you can send a USR1 signal to check the progress. Very useful to check if your tape drive has had a hardware failure.
Information on both this and the byte counts are in my man page (GNU coreutils 8.30 on Linux Mint 20) around lines 130.
Debian testing, coreutils 9.1.
It is also in Coreutils 8.22 in RHEL7.
https://man7.org/linux/man-pages/man1/dd.1.html
Edit: I just found this in the coreutils 9.0 changelog:
dd now counts bytes instead of blocks if a block count ends in "B".
For example, 'dd count=100KiB' now copies 100 KiB of data, not
102,400 blocks of data. The flags count_bytes, skip_bytes and
seek_bytes are therefore obsolescent and are no longer documented,
though they still work.So you can simply do something like:
dd if=foo skip=nB count=mB bs=4M of=bar
To get m bytes from offset nApparently Linux has a facility for 'optimal IO size' of block devices (see 'lsblk -o NAME,OPT-IO') but on my system I only have a value for Linux md RAID devices (which happen to be RAID0, and the OPT-IO value is the stride). All of the other devices have OPT-IO 0.
Perhaps more work needs to be done to bubble some value up from the hardware.
You read about these fake SD cards or USB sticks on sale, where the controller asserts a much larger memory size than the card actually has? You can check such cards with dd, e.g.:
for i in $(seq 1 9999) ; do echo -n "$i "; echo "Record Number $i" | \
dd ibs=1K count=1 obs=256K seek=$[i-1] cbs=4K copy=block,sync of=/dev/sdX
if [[ $? > 0 ]] ; then break ; fi
done
This writes numbered and blank padded blocks (due to the cbs=...) at the start of every 256K block of a device. Adjust your parameters to taste and match the purported size of your card or stick and later retrieve the first block to see if it contains "Record Number 1" or some other number due to wrapping around during writes. (and sure, ibs and count can be removed in the above example, I added them to demonstrate useful options in case the input isn't a simple echo command)Anyway, I'd like to see how the author would handle all those blocking & unblocking, conversion, padding or sync requests available with dd with his head/tail approach. And what about the syntax? OK it was a a joke due to IBM's JCL, but I had to learn infix, postfix and even postfix notation for math (i.e. 5+3, 9! or f(x,y) and integrals and ...) which sometimes too looks like JCL? And have to remember if I need to use -c or -m even for simple tools like wc.
While importing a DB dump is easy to see the progress with dd.
dd if=my.sql status=progress | mysql mydb
I don't remember the exact conditions that trigger it, but `dd` without `iflag=fullblock` can result in
dd: warning: partial read (16384 bytes); suggest iflag=fullblock
I'm able to reliably trigger this with $ cat /dev/zero | openssl enc -aes-128-cbc -pbkdf2 -k foo | dd status=progress bs=1M count=100000 of=/dev/null
but not when I omit the `count=...` for some reason (maybe it isn't showing the warning in that case because it doesn't matter - apparently the effect this has is one of the "blocks" being smaller, and thus fewer bytes being copied, but it doesn't add padding or anything stupid like that, see https://unix.stackexchange.com/questions/121865/create-rando...).I wish we had a cat-like tool for writing into files, for the "cat foo | do-something | sudo dd of=/dev/something" use case.
… | sudo cp /dev/stdin /dev/somethingNot the best of examples, but it comes up all the time. Sure, I could use "tee /some/path/blah.tar > /dev/null", but that has a greater risk of me forgetting the redirect, and getting garbage to the terminal.
dd if=/dev/urandom count=1024 bs=1 | base64
Need it to be in a particular subset of characters?
dd if=/dev/urandom count=1024 bs=1 |tr -dc 'A-Za-z0-9'
Note: if you're on a mac, its tr is kind of broken. Do export LC_ALL=C first.
head -c 1024 < /dev/urandom
head [OPTION]... [FILE]...
head -c 1024 /dev/urandom | base64
LC_ALL=C tr -dc 'a-zA-Z0-9-_\$' </dev/urandom | fold -w 20 | sed 1qDoing this kind of thing is a) rare and b) usually you have more things on your mind at the time, so you tend to be quite conservative and not think to experiment.
It doesn't really do anything particularly arcane (unless you count EBCDIC conversions) but it does things that are useful and often necessary that other equally-standard utilities can't. Isn't that good enough?
> Doing this kind of thing is a) rare
Probably not as rare as you think. Maybe it's not all that useful for writing applications, but for the quite large number of people who have to do provisioning (either bare-metal or virtual) and such it's a different matter. There's a lot to be said for tools that are ubiquitous, well standardized, and flexible.
$ apropos search
...
grep (1p) - search a file for a pattern
... $ apropos search
apropos: Command not found.
Aside from the obvious, man pages just aren’t the answer most of the time.Part of the issue is that a man page is a reference. This is a necessary kind of documentation, but sometimes I want a tutorial and sometimes I want a cookbook. References are great for some things, but discovering new tools isn’t generally one of them. Obviously, you don’t have to write man pages that way, but almost all of the ones I see in the wild are.
Another part is that man pages as a resource have really atrophied in my experience. Lots of new CLI tools don’t have them at all, and lots of systems don’t have them installed even for older/core tools. The why is harder. I suspect that it’s a mix of tooling (e.g. needing to learn groff/troff to write them), search (your man pages aren’t indexed by Google by default like your GitHub readme), and culture. But I’m not sure.
I think dashed arguments, that is, getopt style, just add unnecessary noise. This is usually not so bad. But gets pretty annoying when you have an actual language exposed in the args, I'm looking at you iptables. in iptables the readability would be far improved if they had just left off the dashes.
Then there is the unholy abomination that is --key=value what is that -- prefix doing for anyone?, just drop it and use dd style key=value args and everyone(ok perhaps just me) will thank you.
So my conclusion is, ditch cat and just use dd. The bonus is that dd works in both modes: file to pipe, and pipe to file, downside is that dd can't concatinate files. but who actually does that(joke).
Eventually you'll probably need both but I think the point of the article is that being closer to the Unix/POSIX way will give you more reusable knowledge.
That's certainly true when you're starting out at least. The more small tools you can combine without thinking too much about it the more usable Unix is.
But you are right about the power of unix being that it is a very expressive system exposed in a fairly simple syntax.
We have fallen far from the tree. But I think the real genius of the original unix was in what it did not do. there was a real desire to keep it a fun simple usable system. and it was and still is, as is proven by it's huge popularity and influence even today 50 years later.
On macOS and BSD’s it’s as easy as pressing ctrl-t to send a SIGINFO signal to have dd show the status. On Linux it’s a tiny bit more complicated indeed. But knowing how to send signals to running processes is also a trick valuable of learning.
So, this is an approximation of a command pipeline I run several times per year, when I happen to need a secret of an approximate length:
dd if=/dev/urandom bs=6 count=1 | base64
Tune the "bs=6", depending on how long you want your (guaranteed typeable) secret to be. Every 3 bytes in the input will give you 4 characters in the output and keeping the input block size a multiple of 3 avoids having the output ending in "=".It MAY be possible to replace this use of dd with other shell commands. But, since I needed to learn enough of dd to cope with tapes, I use that.
That's just as easy with dd, decompress|dd is especially useful:
xzcat linux.img.xz | dd of=/dev/sdb bs=1M
Throw pv in there if you want to.
> here are two ways to create a 100 MB file containing all zeroes
And here's a better way: truncate -s 100m zero.img
Other than the missing dash/double-dash, what exactly is "highly nonstandard" about `dd`?
"f" has been a shortcut for "stream" for decades
fprintf
the `arg=value` syntax is common to many command line tools grep --color=auto
; and i/o are literally the abbreviations for, well, Input/Output. So we have "input-stream=" and "output-stream=". How is that "highly nonstandard"? Because the dashes are missing?> But otherwise, try to stick to more standard Unix tools.
`dd` is literally part of POSIX, it doesn't get more standard than that;
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/d...
People barely use dd for anything but copying disk images -- but it's kind of a clunky tool for that. It's got a weird unusual syntax, it requires knowing weird arcane details like what size sectors a device likes working with, and it's easy to shoot yourself in the foot with it.
You'd think that after decades somebody would have made a handy tool for actually writing disk images as an end user. It could use more standard arguments, do sanity checks for destination device contents and mounted filesystems, and automatically determine the optimal block size for the destination device.
It's almost like if we decided that a chisel is the normal tool to use as a flat screw driver.
There are lots of tools out there for creating disk images/boot disks etc.
https://ubuntu.com/tutorials/create-a-usb-stick-on-ubuntu#4-...
Fedora recommends using Fedora Media Writer, and other tools like Unetbootin are very popular.
Though if you’re simply initializing from /dev/zero you’re probably better off with truncate to create a sparse file: https://man7.org/linux/man-pages/man1/truncate.1.html
Can’t beat a program that doesn’t even write the bytes.
The Cult of DD - https://news.ycombinator.com/item?id=13896675 - March 2017 (171 comments)
I wouldn’t call that a cult. More of a “this one off command I run once every five years isn’t something I need to learn or understand because any ‘better’ alternative won’t stick with me anyway.”
Otherwise I use Etcher, or the Pi flasher utility that has really cool extra features.
No I do not care that Etcher is 80MB. Not my app, I didn't pay for it, I don't have to deal with the code, I'm not going to complain.
Don't believe me? Use DD to create a sizable sample file (let's say 2 gigs or so) from rand and try copying that file to your disks with different block sizes (powers of two). Write a small script, and run all your tests with the time command to compare results. Go make dinner, and by the time you're back you should know the optimal block size you should run your fucking huge disk transfer at.
dd is perfect as it is. It does exactly what you tell it to do.
Can you even use ‘>’ io redirects with sudo?
There exists in such a case a certain institution or law; let us say, for the sake of simplicity, a fence or gate erected across a road. The more modern type of reformer goes gaily up to it and says, “I don’t see the use of this; let us clear it away.” To which the more intelligent type of reformer will do well to answer: “If you don’t see the use of it, I certainly won’t let you clear it away. Go away and think. Then, when you can come back and tell me that you do see the use of it, I may allow you to destroy it.”
Ideally you'd want to discover that the building was there, then confidently tear down that fence, but having a fence that nobody dares to touch because the building has been forgotten and we can't find why the fence was even there isn't great either.