Useless Use of "dd" (2015)
vidarholen.net
vidarholen.net
One of the great things about `dd` is that you have a lot of control how the input and output files are opened. You can bypass the page cache when reading data by using iflag=direct, which would stop this from happening.
That is, without iflag=direct, dd will repeatedly ask the kernel to copy bs bytes from the kernel's buffer cache into dd's address space, and then ask the kernel to copy bs bytes from its address space into the kernel's buffer cache.
Edit: correcting for nitpicker
This gives you the best of both worlds, in terms of the shapes/finishes that forming can achieve, but with the workflow benefits of 3D printing.
Similarly, the new hotness is compression/cold forging using 3D printed moulds and carbon fibre (and other materials). Very very cool.
there was some hueristic in there that tried to prevent it, but it wasn't very good.
It would take forever and always end up with an i/o error. I figured the new PC had somehow wonky usb ports or something. But when that happened on another "known good" box, I figured it was the flash drive. But Disk Utility in MacOS worked well.
Then I tried increasing the block size to 1M and everything went smoothly. It even took less time to write the whole image correctly than it took it to error out before.
That being said, probably don't use too small of a block-size - this will eat up CPU in system call overhead and slow down copy regardless of target media type.
I always used dd since it provides more direct control over the transfer stream and how it's transported. I also call dd as "direct-drive" sometimes due to these capabilities of the tool.
"By specification, its default 512 block size has had to remain unchanged for decades. Today, this tiny size makes it CPU bound by default. A script that doesn’t specify a block size is very inefficient, and any script that picks the current optimal value may slowly become obsolete — or start obsolete if it’s copied from "
There are ways to get block size of a device. Multiply it by 2 to 4 (or more), open it directly, and keep your device busy.
The blog post is oblivious to nuances about the issue and usefulness of "dd" in general.
With modern SSDs, "sector/block size" is rapidly approaching vagueness of cylinder/head/sector addressing scheme, as used a couple of decades ago on venerable spinning/magnetic disks.
That is - it is definitely a thing, somewhere deep down, but software running on host CPU trying to address those, wouldn't necessarily end up addressing the same thing as user had in mind.
If you want a concrete example - look no further than "SLC mode" cache - where drive would have a number of identical flash chips, but some of them (or even a dynamically-allocated fractional number of chips) would be run at lower bits per cell count, for higher speed/endurance. However erase- and write- blocksize for the chip is expressed in cells, not bits. What that means is - cache and main storage of the very same SSD would have different blocksize (in bits/bytes).
I don't think it's nitpicking. We're discussing there. We're technical people, and we tend to point out different aspects/perspectives of a problem, and offer our opinions. That's something I love when it's done in a civilized manner.
Regarding to remaining part of your comment (I didn't want to quote to not make it look crowded), I kindly disagree.
The beauty of SSDs are they have a controller which fits into the definition of black magic, and all flash is abstracted behind that, but not completely. Hard drives also are almost in the same realm.
Running a simple "smartctl -a /dev/sdX" returns a line like the following:
Sector Sizes: 512 bytes logical, 4096 bytes physical
This means I can bang it with 512 byte packs and it'll handle it fine, but the physical cell (or sector) size is different, which is 4kB. I have another SSD, again from same manufacturer, which reports: Sector Size: 512 bytes logical/physical
So, I can dd it and it'll just handle it just fine, but the first one needs a bs=4kB to minimize write-amplification and maximize speed.This is completely same with USB flash drives. Higher end drives will provide full SMART (since they're bona-fide SSDs), but lower end ones are not that talkative. Nevertheless, a common denominator block size (like 1024kB, because drives also can be composed of huge, 512kB cells, too) allows any drive to divide the chunk to physical flash sector sizes optimally, and push data to it
In the SLC/xLC hybrid drives' case, controller does the math to minimize the write amplification, but again having a perfect multiple of reported physical sector size makes controller's work much easier, and makes things way smoother. Either because the reported physical size is for the SLC part which is you're hitting for the most cases, or the controller is already handling multi-level logistics inside the flash array (but thinking in terms of block sizes since the this is how it works on the bus side regardless of the case inside).
The subtlety of opening the destination file first, and then writing into it, was what made dd 'special' (and it would open things in RAW mode so there wasn't any translation going on for say, terminals) but that is lost on people. Bypassing the page cache and thus not killing directory and file operations for other users of the system is a level even below that. Only the few remaining who have done things "poorly" an incurred the wrath of the other users of the system sitting in the same room really get a good feel for that :-). Fortunately for nearly everybody these days they will never have to experience that social embarrassment. :-)
[1] Well unless you had noclobber set in which case it would error out.
Please keep in mind that `ditto` is a file copy and archive utility, not a block copy utility like `dd` (which is also available on macOS).
An online man page for ditto: https://ss64.com/osx/ditto.html
Early Windows NT was awful with this, pegging the system with a cascade of disk IO at unpredictable times, often for ten seconds or more.
Can anyone suggest ways to avoid blowing the file cache on Windows with large copies? Is this even a problem anymore?
There remain some useful applications of dd. These may of course be achieved by other mechanisms, but typically less conveniently.
1. Read a specific number of blocks or bytes from a source:
dd if=/dev/hda of=/root/mbr bs=512 count=1
This will make a copy of, say, your master boot record (first 512 bytes of your first disk drive) and stash it in your /root directory.2. Read from specific bytes of a file
dd if=mydata skip=1k bs=32 count=1
Reads 32 bytes after the first 1024 (1k) bytes of "mydata".3. Write to specific bytes of a file
dd if=source of=target seek=10k bs=512 count=1 conv=notrunc
That should write 512 bytes from "source" beginning 10k into "target". (I've not tested this, you should verify.)4. Create a sparse file. Sparse files appear to have a nonzero size, but take up no space on disk, until data is actually written to them. These are often used as "inflating" dynamic filesystem images for virtual machines.
dd if=/dev/zero of=sparsefile bs=1 count=0 seek=20000M # Create 20 GB sparse file
5. Case conversions. Sure, you could use tr(1), but where's the sport? dd if=MixEdCaSE of=lcase conv=lcase # Convert to lower case
dd if=MixEdCaSE of=ucase conv=ucase # Convert to upper case
6. ASCII / EBCDIC conversions dd if=ebcdic of=ascii conv=ascii # ebcdic -> ascii
dd if=ascii of=ebcdic conv=ebcdic # ascii -> ebcdic
When reading to or from IBM data tapes, you might find blocking / unblocking conversions useful. I've done this, but it's so long ago that I don't trust my memory on that any more. Odds are good you'll not have to worry about this.There are other useful applications as well, though these are not typically encountered very often. Do feel free to explore and attempt these on safe media.
Note that this can also be done using truncate(1).
Commands like `cp --sparse=always` aren't worth such a blanket disrecommendation though; if you are working directly at a console as opposed to scripting you typically don't need portability.
What does the (1) mean beside dd? I see this with man pages. Is it a version identifier?
Edit: thank you both for taking the time to share. I appreciate the quick response.
man 1 <name> # executable: <name>(1)
man 3 <name> # function: <name>(3)
Without the distinction argument, man will throw its arms up in the air and give up.However, if there’s no ambiguity (such as with ‘dd’ where only the executable exists), you can drop the section number parameter when running man. man will then search all the sections, see that ‘dd’ only exists in section 1, and go from there.
I also use the notation as a convention to indicate that I'm referring to a Unix command (or library function, etc.). E.g., "cat" might refer to a feline, but "cat(1)" should more clearly refer to the Unix / Linux command.
(Where I've got proper markup, I'll typically set commands or references in monospace using backtick notation: `dd`, `cat`, etc.
You can do that here by putting them on a separate line (two newlines, separate paragraph) and indenting by a couple characters.
This is all ancient knowledge and well in place back in the mid 1980's long before Linux or any of that stuff. The `man` page sections used to be different 3-ring binders all printed out sitting on a table in the computer lab. The 'sections' were just different binders of documentation. Get off my lawn!!!
At least in man-db and FreeBSD implementations, it actually shows you the first page it found by default.
I just tested, and `man man` opened man(1) despite man(7) existing.
man exe <name>
man lib <name>There's nothing necessarily preventing man from allowing string codes as a replacement, but man interprets (fully) non-numeric arguments as pages to go. So `man cat dd` (on Ubuntu) will first open cat(1), and when you exit (with 'q'), will prompt if you want to continue. If you say yes, it'll open dd(1).
That has the side effect than `man exe dd` would be interpreted as opening exe(#) (printing "No manual entry for exe") followed by dd(1).
$ man -k ''
And if you just want to see the ones for high-level documentation (i.e. section 7), then $ man -s 7 -k ''
does the trick. Sections 7 and 5, in particular, are full of hidden gems.edit: `man -k ''` might be a GNUism too. Just tried on a BSD derived OS and got back nothing.
the unix manual was commonly printed out, and these section numbers were a necessity for looking things up. Collation was section number, and then alphabetical
when you were first introduced to unix, you sat down with the manual and read it.
Especially handy when you've fed in a huge amount of JSON (sometimes all on one line because, y'know, why not) into jq and you get the inscrutable output:
parse error: Invalid numeric literal at line 1, column 236162512I'll keep that in mind, json-parsing being among my current hobbies...
tail --bytes=+236162512(edit: I do have to concede that the `tail|head` version is much faster than the `dd` - ~11s vs ~65s in my quick test with that ^skip)
Large blocks -> efficient I/O. Within reason.
The specific recipies should be vetted. The stated goals can be achieved with proper invocation.
dd bs=440 count=1 if=/usr/lib/syslinux/bios/mbr.bin of=/dev/sda
Even in the age of EFI/GPT, that still gets used often (usually VM providers that only offer MBR boot). head -b 440 /usr/lib/syslinux/bios/mbr.bin > /dev/sdaFun fact, the -c usage comes from ksh, where head is a shell builtin.
head -b 512
Will also copy the first 512 bytes in case you want to avoid dd for clarity. I have actually used that for moving mbrs around.What's keeping other implementations from adding these features?
Though in this case I think previous posters are mistaken. My GNU implementation of head supports -c and not -b.
Perhaps it's the fact that the longopt is --bytes that caused the confusion.
I simply miss-rememered.
The thing is that dd issues a read() for each block, but is doesn't actually care how many bytes it gets back in response (unless you turn on fullblock mode).
This isn't really a problem when you're reading from a block device, because it's pretty uncommon to get back less data than you requested. But when you're reading from a pipe, it can/does happen sometimes. So you might ask for five 32-byte chunks, and get [32, 32, 30, 32, 32]-sized chunks instead. This has the effect of messing up the contents of file you're writing, with possibly destructive effects.
To avoid it, use `tee` or something else. Or use iflag=fullblock to ensure that you get every byte you request (up to EOF or count==N).
Today I found the magical dd command that causes an NVMe drive to run at almost full speed:
# dd if=/dev/nvme3n1 bs=4096k iflag=direct of=/dev/null status=progress
959090524160 bytes (959 GB, 893 GiB) copied, 178 s, 5.4 GB/s
The trick to getting this throughput is telling Linux to do an insanely large IO (4 MiB). The drive can't do 4 MiB reads - the largest IO it can handle is 2 MiB. # nvme id-ctrl /dev/nvme3 | grep mdts
mdts : 9
# echo '4 \* 2^9' | bc -l
2048
More in the tread starting here:https://twitter.com/OMGerdts/status/1514376206082269191?s=20...
$ sudo dd if=/dev/nvme0n1 bs=4096k iflag=direct of=/dev/null status=progress
...
9667870720 bytes (9,7 GB, 9,0 GiB) copied, 20,2714 s, 477 MB/s
But when reading from empty space (as in, non-written yet blocks), it goes at full PCIE x4 Gen4 speeds. $ sudo dd if=/dev/nvme0n1 bs=4096k iflag=direct of=/dev/null status=progress skip=450000
...
16710107136 bytes (17 GB, 16 GiB) copied, 2,42802 s, 6,9 GB/s
I have another nvme drive - Force MP510 - and it doesn't care if data was previously written or not. When reading from it, I get ~full x4/Gen3 speeds of 3.5GB/sPS: nvme smart-log shows 100% available spare, and 0% percentage_used, so it doesn't seem to be wear-related.
is it possible that the controller is over-heated? Do you have the NVME drive under a GPU?
Yes, reading "evicted" or "trimmed" space [0] specifically, is much less work - only flags are read, not actual media.
They write in small blocks (eg, 4K), but can only erase in very large blocks (eg, 2M). Writing to an empty device is easy, but eventually you have to overwrite something, and you can't just replace a 4K block with another. You have to take a contiguous block of 2M, wipe the entire thing, rewrite whatever part of it was useful, then do the write you actually wanted to.
This effect is known as "write amplification", and it means that in bad cases you need to do many times more work than the host system requested.
Modern high end SSDs have various ways of dealing with this like a RAM cache, a SLC cache and extra reserved space to always have some spare room, but there are still limits.
This is what TRIM (confusingly also known as 'discard' in some contexts) is for -- the SSD operates in blocks and has no way to know that some chunk you've written to before is now useless to you because you deleted the file it belonged to, until you give it the command to overwrite that block. TRIM tells the drive "these parts can be recycled", and allows it to create empty blocks in advance. So make sure to TRIM once in a while.
Sadly TRIM wasn't well specified initially in regards to what performance characteristics it should have -- some old drives can get stuck on it for a while. So while many filesystems support emitting TRIM automatically, it can cause severe performance issues on some drives and the recommendation is to do it as a maintenance task on a timer instead.
TL;DR: Run `fstrim /mountpoint`. Wait a while before testing if it changed anything. The drive isn't obligated to do the work immediately.
It gets a bit more complex than that due to layers. LVM and dmcrypt can filter out TRIM requests. You can use `lsblk -D` to check the support status.
This is called garbage collection. It may be happening at any time in the background but becomes more frequent and perhaps in the write path as the drive fills. 2M is an example size - the actual size will vary by drive model and it is rarely disclosed.
> Modern high end SSDs have various ways of dealing with this like a RAM cache, a SLC cache and extra reserved space to always have some spare room, but there are still limits.
There are multiple types of SLC cache as well. Client drives may have a small number of gigabytes of SLC that can absorb a small burst of writes. Client drives may also have pseudo SLC (pSLC) that is called pSLC, TurboWrite (Samsung), etc. With pSLC, when there's about 30% of the drive's NAND that is erased, the drive will use that as SLC. So, a 1 TB drive will use about 300 GB of space as a 100 GB SLC cache.
How performance degrades as the drive fills varies widely between drive models. Some drives start to have significant read and write performance degradation long before the drive is half written. Others will maintain fairly consistent read performance (maybe within 90% of that which is advertised) regardless of how full the drive is. For instance, the original version of the Samsung 980 Pro maintains close-to-spec read speeds regardless of how full it is but write performance drops from about 5200 MB/s to about 1300(?) MB/s the moment it hits 70% allocated.
Datacenter and enterprise class drives tend to have lower peak performance than client drives, but their performance is much more consistent regardless of how full they are.
If you are buying a client NVMe drive for speed, buy one that is larger than you need and set aside at least 30% of it in unpartitioned (or unused partition) space. This will prevent the OS from writing to 30% of the drive thus keeping plenty of space for pSLC and similar optimizations. This will also increase the life of the drive as garbage collection is likely to have to rewrite the same data less frequently, resulting in a lower write amplification factor.[1]
1. See Over Provisioning at https://semiconductor.samsung.com/consumer-storage/magician/
> This is what TRIM (confusingly also known as 'discard' in some contexts)
But wait, there's more terms for the same concept: trim (ATA), unmap (SCSI), deallocate (NVMe) are interace-specific ways that Linux performs discard.
> Run `fstrim /mountpoint`
Or if there is no filesystem, `blkdiscard /dev/nvmeXn1[pY]`.
The degradation you are seeing is rather extreme - it seems to be performing worse than a cheap SATA SSD. I'd say this is worth having a conversation with Kingston.
$ echo 'like this' | dd bs=1 speed=10
“Well why does the entire internet say to use dd then?”
Because they copy from each other just like you copied from them. Just use cat.
Great statement. This brings me some much needed internal clarity on my own thoughts and actions.
https://www.executiveforum.com/cutting-off-the-ends-of-the-h...
tl;dr nobody in 2 generations knows why they cut the ends off the ham before cooking it, until they talked to grandma, who said her pan was too small
A Unix thing that's been posted to HN for a decade, that's almost literally the same story:
Understanding the bin, sbin, usr/bin , usr/sbin split
https://news.ycombinator.com/item?id=3519952
tl;dr /usr/bin is separate from /bin because someone had a small hard disk once
There was a problem about integrating support for different compression applications, with each one needing to get a new letter in the tar command!
x - extract
z - gzip format
f - file (must be last so it parses the filename arg correctly
I also add "vv" in the middle so it lists every file as it goes, so I can see it work instead of just waiting with no output.
That said if you don't know whether or not you need it, you don't. And those who do would know not to listen to your advice (for what they need to do) anyway, so it's not bad general advice to give.
UNIX’s simplicity of “everything is a file” continues to surprise me, even though it shouldn’t any more. I think a lack of confidence in my understanding of the basics, leaves me tempted to copy from others as you mention.
We’ve gone full circle.
https://en.m.wikipedia.org/wiki/Cat_(Unix)#Useless_use_of_ca...
Anything after the "cat file |" can't overwrite the file.
cat file | cat > file
(This won't, as you might hope, leave the file unchanged.)Re-read what you typed before hitting enter, and keep backups of important things.
More likely is you are messing around with some new/infrequently used utility and you pass in the arguments incorrectly and specify "file" as the output instead of the input as intended.
At the end of the day do whatever pattern works for you - for me it's doing "cat file |" at the start of a pipeline and "> outfile" at the end.
I also avoid globbing inside a pipeline as it can be dangerous too.
Because you usually need root to write to drives, and you can't `sudo` redirect. `sudo dd if=whatever.img of=/dev/whatever` is nicer and easier than `sudo sh -c "cat whatever.img >/dev/whatever"`.
Your example would then become something like...
cat whatever.img | sudo tee /dev/whatever > /dev/null # https://stackoverflow.com/questions/82256
Just like in household plumbing, the 'tee' command basically takes the input and sends it to more than one place. Naturally, running 'sudo tee' will let you send things all over but as another user.
All that said, I won't speak to the comment "why does the entire internet say to use dd". I've never actually found the "whole internet" to agree on much of anything (-:
I just tried to copy an ISO to a USB drive with `rsync`, and it didn't work.
Looking it up, I read these comments[0]:
"rsync operates on files which are on a filesystem. It doesn’t do comparisons between blocks. If you want to back up a block device to a file, use dd"
"I completely understood the use case. rsync (still) does not operate on block devices, so it’s not the solution here."
"Agreed, rsync has never operated on block devices."
So I thought, "Ah. There's File Data, which is what `rsync` deals with, and RAW BLOCK data, which is what scary hardcore tools like `dd` deal with."
Then the notion that `cp` and `cat` can deal with both species of data is confusing. :p
0: https://old.reddit.com/r/linuxadmin/comments/eappzm/rsync_bl...
That's the thing; it isn't. Author forgot to explain (if he knows that at all) that /dev/sda2 on Linux is not a raw device. It's a block device.
So if dd is hailed as something to use on /dev/sda, that's not an example of being hailed for a raw device.
dd's capability to control the read/write size is needed for classic raw devices on Unix, which require transfers to follow certain sizes.
E.g if a classic Unix tape needs 512 byte blocks, but you do 1024 byte writes, you lose half the data; each write creates a block.
The raw/block terminology comes from Unix. You have a raw device and a block device representing the same device. The block device allows arbitrarily sized reads and writes, doing the re-blocking underneath. That overhead costs something, which you can avoid by using the raw device (and doing so correctly).
Tapes are still annoying on both in their block-ness, iirc?
Tapes aren't block devices at all - that's why you can't mount them :-)
Tapes by nature are block devices due to only being accessible in block increments, and writing them in transparent way is a bit more problematic than with disks, as being able to just re-read a block to do read-update-write cycle isn't guaranteed - or necessarily easy.
And yes, there used to be a time where you could mount tape as filesystem ;)
Heh, I remember finding out this on the good old days when I was finding out how to rip cds and shitting bricks. I mean I read it on a BBS and thought "that can't be right"; I was expecting to find something like the Nero suite back in windows.
Much more recently, I enjoyed the same kind of amazement on the bash tcp pseudo devices.
There's very little reason not to use it, even if it's just to get a nice progress view instead of just the current amount of data copied.
It's been invaluable as an intuitive tool to recover data from failing disks/drives.
Never thought to use it as a day to day dd, but will give that a try. Thanks for the idea!
> If an alias specifies -a, cp might try to create a new block device rather than a copy of the file data. If using gzip without redirection, it may try to be helpful and skip the file for not being regular. Neither of them will write out a reassuring status during or after a copy.
> dd, meanwhile, has one job*: copy data from one place to another. It doesn’t care about files, safeguards or user convenience. It will not try to second guess your intent, based on trailing slashes or types of files.
> However, when this is no longer a convenience, like when combining it with other tools that already read and write files, one should not feel guilty for leaving dd out entirely.
When I first learned that dd was not magical, I started using cp but I made some mistakes with partitions number and whatnot (nothing serious).
Maybe it's just the wierd syntax or the fact that I treat dd differently, I'm just more cautious and don't press enter automatically. Of course it's silly but to me at least, that's a good reason to keep using dd.
> dd if=/dev/sda | gzip > image.gz
This actually serves at least two approximately legitimate purposes: firstly, it ensures that reads from sda are aligned to a (at least nominal, ie 512-byte) disk block, which doesnt matter for normal, kernel-supported drives like IDE/SATA/most USB (which is almost certainly what sda is), but avoids bespoke devices (or their drivers) trying to do something clever when gzip asks for 1 or 7 or 17 bytes. (And writes to poorly-designed devices/drivers can be even worse.)More importantly, like useless use of cat, it prevents gzip from trying to delete sda when it's done, which is something it will in fact do:
$ echo test > /tmp/sda
$ gzip /tmp/sda
$ cat /tmp/sda
cat: /tmp/sda: No such file or directory
(For gzip specificially, you can also prevent this by writing `gzip </tmp/sda`, but I've occasionally run into tools that try to 'intellegently' handle stdin file descriptors that point at 'real' files, so I feel better having a separate process blocking the way.)And there's also the fact that deleting the source file by default is always the wrong thing for it to be doing. If I want to deal with corner cases like being almost out of disk space, I can pass --delete-source explicitly.
But whatever. I’m sure 90% of code stems from something someone read and copied or came most easily to them. It works.
‘dd’ was really useful for finicky media like tapes and doing EBCDIC translation. It’s still great when you combine bs and count. Blow away an MBR, make an xGB file, etc.
It’s a Swiss Army knife. It can do a lot, but isn’t the best tool for most things. I still love it. Probably just muscle memory.
And it’s nicer that the recent ‘dd status=progress’ (if I remember that option right).
You’re 100% correct. I’ve grown more tolerant of these ceremonial uses. There really aren’t many folks that understand shell. I see it in CICD, etc all the time, too.
The one left that kills me that I see in vendor scripts all the time is ‘command; ret=$?; if [ “$ret” -ne 0 ]; then foo; fi’
If you’re handling return code cases, then cool. But folks don’t realize ‘test’ aka ‘[‘ is just another command.
But then again, so much Java, etc. I read is the same copy/paste. Those can be worse because example code isn’t prod ready, where bad shell usually does the right thing, just awkwardly or inefficiently.
But as you point out, so much isn’t thinking about what you’re doing, just mimicking.
In FreeBSD (and presumably other UNIX implementations) they aren't: https://docs.freebsd.org/en/books/arch-handbook/driverbasics...
So, in FreeBSD "dd if=/dev/ada0p1 of=/dev/null bs=1 count=1" will fail: disk driver will return EINVAL from read(2) because I/O size (1) is not divisible by physical sector size (usually 512). "cat" with buffer size X (which depends on the implementation) will either work or not depending on divisibility of X by physical sector size, and other random factors, like short file I/O caused by delivery of a signal.
Summary: dd(1) still has its place and author of original article is getting it wrong.
I was at a satellite office with all windows PCs for the day, so used a live disk to get a decent environment to get things done. Only problem was that the DVD drive kept spinning down and every time I did something that was not cached it made me wait forever.
nohup + while loop + sleep 4s + raw dd read from CD for the win :)
EDIT: Reading this article it sounds like dd has no "special" ability to access the disk in a raw way. But surely that's what the nocache option is for...
Dropbox had some promotion for their new photo storage service, where they were giving away free extra space up to 10gb to encourage you to store your photos.
Some clever cookie told me you could make 10gb of ‘empty’ jpeg files and put them in your Dropbox folder.
The Dropbox app would compress this 10gb of files for upload (usual behaviour for the app in those days) and increase your storage permanently for free.
I can’t remember the command but it was something like
dd if=/dev/zero of=/path/to/dropbox/1.jpg bs=10000000
Bam! 1 10gb jpeg file full of zeros that compress down to a few bytes. Instant 10gb free. Now uploading required!
I think I still have that dropbox account, I learnt about dd and dev/{null,zero,one} that day and never used them again.
Bravo, you've found something (marginally) more upsetting than using ed:)
yes 'y' | dd of=response bs=1 count=1
yes 'e' | dd of=response bs=1 count=1 seek=1
yes 's' | dd of=response bs=1 count=1 seek=2A few years ago, "128K" was usually the fastest choice. Today, on faster systems, "512K" has a slight edge.
I could not tell you why, though. Try it for yourself.
For USB also make sure to power it off before removing otherwise you might lose data. Eg. by
udisksctl power-off -b /dev/sdX"udisksctl power-off" checks nothing is using the drive, commits buffers to storage, deconfigures the drive, and powers it off. Unfortunately it may kill more devices than you want in some scenarios due to the way killing the port works. Also for normal USB drives I don't think® poweroff is any different than an immediate physical pull post unmounting. The upside is it'll be completely disabled the instance the command completes, i.e. it can't even be written to raw as it is no longer powered even though plugged in. Not sure how helpful that is in reality though.
I usually just umount the drive which syncs and prevents further writes (well, to the mount at least) but doesn't try to kill it for me. Most likely sync is "good enough" for the vast majority of use though.
sync; sync; sync
I feel more comfortable with just unmounting and powering off nowadays. But I have no hard evidence.Hardly ever user more than one USB drive same time, so not sure whether there could be side effects with powering off one of them. Well, USB being USB, nothing would surprise me too much.
And even if you use it without it being needed, it's not a big deal. It doesn't add much overhead, if any.
What I normally do is just do a 'watch -n 30 killall -USR1 dd' in another window which triggers regular progress updates :) That why I don't use ^T
I use FreeBSD also but indeed it supports it now.
I personally suggest recommending people the Disks application included with Ubuntu Desktop and Fedora/Centos Workstation. It shows icons representing internal disks, SD cards or flash drives so they know what device they want to work with. If they want to take their time they can see all the information about the drives and partitions, they can start discovering and asking questions and reading up on how computers use disks right from there if they want to. And if they don't want to, it's just extra confirmations that it's the correct disk or DVD that they want to put their image onto. Then when they're sure about the device they can create a disk image and restore a disk image in that same application!
Well, it pays off a lot to know a complex tool and using it for easy stuff gets you in the habit of using it, reminds you of the arguments etc.
dd is really really useful.
Edit: Looks like O_SYNC (oflag=sync) is needed, too. Should update my 'sddd' alias.
"The O_DIRECT flag on its own makes an effort to transfer data synchronously, but does not give the guarantees of the O_SYNC flag that data and necessary metadata are transferred."
While I don't think this is relevant to block devices, I see no harm in including the flag in my alias, either.
I can't find anything about it in the man page - does anyone know the options? I'm assuming it doesn't mean simply running dd with seek to output another file and then running that.
It's not something you'd realistically use, but it's better to be limited by imagination than tooling.
The by-label trick helps a bit, but I still don't like it. Etcher exists. Etcher will give an alert when it's done flashing. Etcher will verify what it wrote. Etcher will predict how much time is remaining.
The raspi imaging utility goes even farther and gives you configuration options. There are many other special purpose flasher utilities like that.
Where dd really shines is in a script, for making empty images of a certain size. But even then... there are tools like truncate to make a sparse file instead.
The only other time I ever need the CLI is for directly creating a compressed image, but that's a somewhat uncommon task for me. And I would not be surprised if one of the GUIs had it by now.
There actually was a Window system named “W”; it was the window system for the “V” operating system (IIRC). “X”, therefore, was the successor to “W”.
status=progress is also quite handy. You can splice a "| dd status=progress |" in the middle of a pipeline to get a sense of what's happening.
cat < file | somecmd | cat > file2
#vs
somecmd < file > file2 cat file | grep ...
or
cat file | awk '{blah}'
for ad-hoc stuff even though I am fully aware that I could easily avoid the cat simply because having the regex or awk code at the very end makes it easier to read, edit or extend with additional pipeline elements. ~$ <frongle.txt grep fringle >frungle.txtI would hope both commands boil down to exactly the same action. Any overhead from invoking cat (isn't that a shell built-in?) should be negligible.
<file1 somecmd >file2
I wish that shells would just optimize the useless cat behind the scenes, so I could use it without soliciting complaints :)Archive link of this; at least for me, the original site isn't responding.
cat /dev/sda | pv | cat > /dev/sdb
This could be replaced by just: pv < /dev/sda > /dev/sdb
Or in this case even: pv /dev/sda > /dev/sdb </dev/sda pv >/dev/sdbI came as far as comparing the image files, which were identical, but I did not have the time to compare the disks. So, this still is a mystery for me.
On a side note: If you ever need to do a full disk clone nowadays, I can only recommend clonezilla. Using FAI (Fully Automatic Installation) would be even better instead of doing image clones, but sometimes, that's not an option.
It can be set to skip over bad blocks to copy as many healthy blocks as possible, and then go back and retry the bad blocks over and over until they either read successfully or you stop it.
I used it once to recover a 1TB NTFS drive which can't be mounted by Windows. It only failed to recover 12kB in 1TB, which corrupted an MP3 file, which can be easily replaced.
It always had issues of course: e.g. if the source was a 100MB drive and you dd'ed to a blank 200MB, you'd get a new 100MB drive and all the rest was lost because it was a low level copy. Similarly any bad sectors on the old drive became bad sectors on the new and any bad sectors on the new were simply used as if they were OK.
cp myfile.iso /dev/sdb
compared with this one: dd if=myfile.iso of=/dev/sdb bs=32M
because implementation of cp have a fixed buffer, so if the amount of data is big and the disks fast, using cp you are calling more read() and write() syscalls than necessary, slowing down the copy process.Last time I researched this, a long time ago: `lz4` is (was?) your best friend.
Oh, but - to follow the submitted article's "rant" - if one wants to use `cat` instead of `dd`, shall there be freedom... Sometimes people like some kind of form-aided "readibility and clarity and method", which leads you to forms like `cat filename | grep` .
dd if=/dev/zero of=/mnt/test.data bs=1M count=1024
Overall, a very informative article though!
When writing to a USB sd card reader, sometimes the data wouldn't be written entirely without calling 'sync' but dd can perform the sync itself with 'fsync'
dd if=/dev/sda | gzip > img.gz
dd if=img.gz | gunzip | dd of=/dev/sda gzip < /dev/sda > img.gz
gunzip < img.gz > /dev/sda
Those should give you better block sizes than default dd without bs=128k or whatever.cat /dev/mmcblk0 | pv -abtrN Processed | gzip > img.gz
cat img.gz | gunzip | pv -abtrN Processed > /dev/mmcblk0
gzip -c /dev/sda > img.gz
For your second command: gunzip -c img.gz > /dev/sda
are what I suspect the blog post would recommend.why using these "oldschool" if / of parameter
how its meant to be used
dd < /dev/zero > /dev/sdx
like others already mentioned: parameter "bs" ~ block-size
dd < /dev/zero bs=1M > /dev/sdx
or parameter "count" ~ number of blocks of specified size
dd < /dev/zero bs=1M count=100 > /dev/sdx
show progress ~ utility "pv"
dd < /dev/zero bs=1M | pv > /dev/sdx
etc.etc.
why do people write such articles w/o at least consulting the manpages!?
and why is such a mediocre article on the frontpage of HN!?
just my 0.02€
mysqldump -u mysql db | ssh user@rsync.net "dd of=db_dump"
pg_dump -U postgres db | ssh user@rsync.net "dd of=db_dump"
(not useless, though ...)I regularly do `cat ~/.ssh/id_rsa.pub | ssh foo@bar tee -a ~/.ssh/authorized_keys` without repercussion. For larger files I usually use rsync as you can resume interrupted transfers
mysqldump -u mysql db | ssh user@rsync.net "cat > db_dump"Where does dd's name come from then? Its man page does not tell.
> dd actually has two jobs: Convert and Copy. A post on comp.unix.misc (incorrectly) claimed that the intended name “cc” was taken by the C compiler, so the letters were shifted in the same way we ended up with a Window system called X. A more likely explanation is given in that thread as pointed out by Paweł and Bruce in the comments: the name, syntax and purpose is almost identical to the JCL “Dataset Definition” command found in 1960s IBM mainframes.
I can't remember specifics, but I dd ed the first 512B to a file and pointed the windows OS loader to that file. The specifics I don't remember are why I wanted to use the windows loader instead of LILO.
$ gnome-disks --restore-disk-image=/etc/motd
Behind the scenes it calls udisks, which uses polkit to ensure that the user is authorized to write to the target disk.Overall, the chances of a user typoing and wiping out the wrong disk are greatly reduced. They also get to see progress etc.
dd if=/dev/urandom of=/dev/null
?So?
The problem with this article and the useless-cat article is that 1) they create a false right vs wrong dichotomy and 2) it gets amplified and used as a tool by the small minded as a citation for "the correct way".
A post on learning your tools is fine. But calling something useless and needlessly adding branches to one's decision tree doesn't empower anyone except jerks.
Stay useless.
Were any of the cited examples harming anything? No they weren't.