I don't know why dd remains the standard for so many tutorials when there are competent GUI tools that can take all the guesswork out of it. It's like a perpetual hazing ritual for new Linux users.
I don't know why dd remains the standard for so many tutorials when there are competent GUI tools that can take all the guesswork out of it. It's like a perpetual hazing ritual for new Linux users.
You can also do something like
dd if=file.iso of=/dev/disk/by-id/ata-Samsung_SSD_840_EVO_120GB
Where the last part is basically the brand and name of your USB stick, not much guess work required. Though I've never seen that in any tutorials.In short, I think ignorance is the simpler answer ;)
Not even then. The prospect of making an occasional mistake, while low, is clearly much higher than when you are selecting the disk using its name. Being experienced doesn't excuse you from the responsibility of using a less risky tool.
Usually on Linux I only use /dev/disk/by-id/foo in scripts and config files. But using it when doing routine stuff with dd is a pretty good idea too.
It’s too bad that macOS and FreeBSD hasn’t anything similar to my knowledge. And since I use both of these operating systems so much, and so often for things involving dd, I think in my case I don’t gain much from doing by-id for dd on Linux unfortunately.
Even this is faulty solution. By default, it only lists the devices with their "/dev/sd*" names.
I've also been doing paths through `/dev/disk/by-id/` in the ops documentation and scripts for my startup.
It's a tiny convention for us promote, and one of countless good practices we exercise, but this particular one can save much misery, at negligible cost.
And early, precarious startups are one of the contexts in which some of these tiny negligible-cost good practices seem to pay off especially well: in an early startup, it's easy for me to imagine a data loss, bad downtime, or missed opportunity ending the company (when a more-established company might be able to weather a mess-up better).
$ lsscsi
[0:0:0:0] disk ATA ST2000DX001-1NS1 CC41 /dev/sda
[3:0:0:0] cd/dvd ASUS SH-224FB 1.00 /dev/sr0
[4:0:0:0] disk Generic- Multiple Reader 1.11 /dev/sdb /dev/disk/by-id/ata-Samsung_SSD_860_PRO_1TB_S42NNF0K123456N -> ../../sdchttps://perfectmediaserver.com/
It's full of other great nuts-and-bolts advice.
I feel it's important to mention you want by-id as there are 5 permutations of /dev/disk/by-*:
/dev/disk/by-id
/dev/disk/by-partlabel
/dev/disk/by-partuuid
/dev/disk/by-path
/dev/disk/by-uuidFor folks just getting started or more comfortable with a GUI, I’d recommend giving USBimager[0] a look. It does exactly what’d you’d expect based on the name, it performant, and they have native apps. No affiliation, just a fan of a KISS app done right.
Though if it's just for ISOs, ventoy is fantastic, just drag and drop the file and no burning at all :)
I ran into this very issue when trying to make a bootable Windows 10 USB on macOS. No amount of fiddling with dd, unetbootin or Etcher resulted in a bootable USB. Despite being principled about it, I had to admit defeat and just pulled an old <4GB Windows 10 iso and flashed that to the stick.
I know I could have installed Linux through a virtual machine and got it done that way, but that seemed horrible overkill. Oh well.
This has always worked for me to boot UEFI installers
Recent windows ISOs have a file that is >4GB, so you can't have the partition formatted as f32.
exfat isn't compatible with uefi.
These are the two largest files in that image.
5.0G ./sources/install.wim
534M ./sources/boot.wim
I wonder how Microsoft's media creation tool deals with this... Dism /Split-Image /ImageFile:C:\folder_name\sources\install.wim /SWMFile:C:\folder_name\sources\install.swm /FileSize:3800
I copied that from a ZDnet article with more details. Wasn't sure about posting the URL but searching will find it. 4.0G ./sources/install.esd
381M ./sources/boot.wim[0] https://flathub.org/apps/details/org.fedoraproject.MediaWrit...
Wouldn't one (at least partial) solution here be, that the kernel should refuse to let you write directly to a block device in use by a mounted filesystem? (Maybe combined with some special ioctl/whatever to bypass that restriction if you ever really need to.) Then, if you are running dd from an OS running on your main drive, the kernel will refuse to let dd overwrite the main drive, but will let it overwrite the flash drive (which presumably is not mounted, and anyway shouldn't be if you are about to overwrite it)
I have written a few wrappers like that on my own system preventing me from making a few common mistakes of mine (like scp'ing a file locally to a filename ressembling an IP address instead of on a remote server ;)
I think this could be addressed by the idea of a "parent block device". So /dev/sda1 is mounted, then its parent /dev/sda would be classified as mounted, but /dev/sda2 would not be (assuming there is no partition mounted there.)
I'm sure one could work something out that would work for device-mapper, LVM, etc as well
> Secondly, the Linux kernel has a very strong backwards compatibility guarantee
What about a sysctl knob? Turn it on, you get this new behaviour, turn it off, you get the backwards compatible behaviour. Each distribution can decide what to default it to. If it defaults to off in Linus' tree, that should satisfy his backwards compatibility concerns.
> With your idea, many disk management programs will be broken.
There would need to be some escape hatch, e.g. an ioctl, to allow unsafe writes. And disk management programs would have to be patched to invoke that escape hatch. A distribution wouldn't ship the sysctl as defaulting to on until it had patched all the disk management programs in that distribution. And, if you download a third-party tool, either its developers have patched it to use that ioctl, or else you can just temporarily turn off the sysctl knob while you use it.
While dd isn't the best tool for writing images to devices (if only because of its arcane and bizarre command-line syntax), it is a valuable tool when you want to recover data from media that has errors. The 'conv=noerror' option is a lifesaver for when you want to recover something from your media.
But overall I agree: plenty of people recommend using dd just because "it's always been done that way" and often also because it makes them look smart :-)
ddrescue -b 2048 -d -r 3 -R -v /dev/sr0 image.iso image.log
If you try to read with a larger block size than the media's native block size, you'll get errors for chunks of that size even when part of the data may have been recoverable.
The above is true whether or not the OS has access to the media's raw ECC data.
From experience of recovering bad DVDs and BRs, there were discs I had to pass through 100+ times to get a valid read on all blocks.
This is also because, the underlying hardware will always do a full ecc block read (only way for it to determine that it read the block correctly, to read whole block and verify it), so any smaller reads are pointless.
the way optical media works is that the optical media reads bits from the drive in ECC block size and then verifies / fixes the block and passes that back to the OS if its has a valid block, otherwise returns an error to the OS. hence, my logic that I describe below.
optical media is an unreliable medium in general and hence depends on the ECC codes to ensure blocks are read correctly and they are used a lot. back in the day of CD and DVD burning there were fancier burners that provided apis for reading the error correcting stats into user space (i.e. how many of different types of errors were corrected), dont know if they still exist. It was never 0 across the board, but that's how the medium was designed, not to require it be 0 across the board.
As you said, the nature of the medium requires ECC (side note: modern hard drives do too). So if I ask for a 2048 byte sector, the drive has to read the ECC. So why ask for more than that? It already knows the sector boundaries. In other words, if I tell `dd` to use a block size of 2048+$ECC, won’t that actually work a sector and a half (well, 1 + $ECC/2048) at a time?
[0]: In fact, unlike CDs, I don’t even know or have any way of finding out how many ECC “bytes” there are in my hard drives' sectors
simplistic case, imagine we have 1 ECC block of 16k, but we read at 2k, so we'll number the 2k blocks 0-7
T0 - read block 0, fails T1 - read block 1, succeeds! T2 - read block 2, fails T3-T7 repeat for blocks 3-7, all fail
in practice if we read a 16k bock at T1, we would be golden (and finished). Instead we did 8 steps, and only got 1/8 of the data.
This is becaue the OS doesn't have a concept of the hardware's ECC block size, so the optical hardware in a sense virtualizes it, and the OS will just keep on rereading the same ECC block on the media and possibly continue to get errors.
you just care about the size of data the ECC is protecting. if the ECC protects 16K or 32K of data, you want to read on those physical boundaries. as then you'll read a whole ECC block and it will either pass or fail. If it passes, you never have to try to read that ECC protected block again (and maybe fail).
Of course, there is one hitch to my scheme. Ensuring that you always read on ECC protected block boundaries. I'm pretty sure if you use ddrescue to always read the right block size it will, but not 100% (why not? perhaps the ECC protects data not visible to the end user in some way (say the first block is only 8kb, not 16kb in practice).
on the issue of hard drives, there is a lot more going on that puts you at the mercy of the firmware (relocatable sectors and the like). I did lose a RAID5 once (1 drive totaly died, and then in rebuild, another drive threw and error) , and I was able to use ddrescue to recover all but 4K block on it. as I was using a 128KB stripe size, that meant I probably lost somewhere between half MB and a MB of data - if the 4K was contained within a single stripe or not (probable it was). I was content with that. never did discover what data if at all was corrupted, but I was able to recover the raid5.
It is not arcane, you just have to read docs, like for every command line tool. Different cli applications have different syntaxes usually influenced by their domains, eg compare find, tcpdump, iptables, cut, docker.
I used dd a lot to burn images on USB. And that command is simple as expected.
I didn't say it wasn't simple. But it is very weird.
Do you realize that anyone can make a cli tool with whatever syntax they like? :) Also, see "man xm", which uses quite similar approach. And there are probably many other examples.
The tool is originally meant to do various conversions of formats of data on 8-track tapes and both its name and syntax is reference to (arguably less bizarre) syntax used by IBM's JCL to produce contents of tapes that need that kind of conversion to be usable on unix.
I hate tutorials that start out with “I’m gonna show you how to do X. I like to use Y, but since the Y tool isn’t installed, we have to get it. Start by editing sources.list. You might need to be root to do this. Here is how you do that.” Halfway through the tutorial we’re still Yak shaving.
I actually like dd's command-line syntax the most out of the Linux coreutils, and I wish more programs used a similar "key=value" argument system. It's pretty easy to remember too, since there's really only a few keys you need to remember to do 99% of what dd typically gets used for.
One trick to be sure that you've typed correctly, start with 'echo':
# echo dd if=xxx of=yyy [enter]
This way you'll be able to check that command will be executed with the parameters you want. Especially helpful if you're running loops, e.g:
# for i in xxx; do echo yyy; done
I would never use GUI for that, as you can't be sure what it will execute, no matter what it shows.
You miss the point entirely then. The point is to use a command line because what you think is competent for a GUI tool isn't nearly as intuitive as you think especially when a GUI isn't available.
The root user is special in that it can overwrite devices. Look at what you’re typing before using its powers.
ls /dev/disk/by-id
instead of /dev/sd*Other than that, someone can replace the whole of Etcher with a simple bash script that will list available block devices with nice descriptive names taken from sysfs, allow the user to select the one he wants to flash to, and handle flashing of potentially compressed image and verifying the result automatically.
It would probably not be much longer than 100 LoC, at least on Linux.
If a user blindly copy-pastes and destroys their main drive, that's a very good lesson to thoroughly double check before running destructive operations.
Also the advantage of dd is that you can run it with sudo easily, while with cat it will not work, since the redirection is done by the shell and not the cat binary itself, and you either have to open a root shell, or pipe the output of cat in `sudo tee filename >/dev/null` that is less than ideal.
Also I think the problem is of Linux that lets you write on disks that are mounted. In MacOS is forbidden and you must umount the drive (so overwriting your root partition by error is impossible).
I just ran gparted:
$ gparted
and instead of complaining that I didn't specify an argument, it chose the first one: /When I realized what I had done, I couldn't recover the partition table, but I managed to rsync everything elsewhere. I'm sure there was a better/safer way but everything was recovered in the end.
I've often used whatever disk utility came with my distro. Or, rufus (rufus.ie) if I'm on windows.
No later than yesterday I copied my wife's entire (Windows 8.1) HDD to an SSD using the free (GUI) version of "Macrium Reflect". I'm pretty sure you can make the exact same mistake: copying destination unto source instead of the contratry. I used a Windows software and not DD for it was a Windows computer and I didn't feel like booting a live Linux CD to do the dump but under Linux, I always use dd.
Is it really that hard to learn that in dd the 'i' in "if" means "input" and that the 'o' in "of" means "output"?
The problem with "making things simple using a GUI" is that you typically totally lose the ability to do not just advanced things but even "average" things. For example for read-only media I always write the checksum on the media itself, using a sharpie. That way I can easily verify that my disk (or the copy) is ok doing a dd and piping into sha256sum (there are some gotchas to keep in mind but it works fine when you know how to do it).
Like, say, a Debian install ISO. I like to have these on read-only medium and make sure the checksum matches the official one (so I prefer a write-once / read-only DVD than a read/write memory stick).
How do you that with the GUI? Piping into a cryptographic hash?
I can understand that people prefer GUI over command line for many things but I think that people imaging entire disks are at least power users and can learn the difference between "input" and "output". Heck, maybe it's even doing them a service to teach them to use the CLI. Maybe it's the opportunity to teach them about piping, about cryptographic hashes, etc.
And once again: you can totally screw up with a GUI too.
I don't mean it in a bad way at all but... Linux on the desktop (which I use since 20 years) ain't exactly enjoying a huge market share compared to Windows or OS X and I think that's fine. If a Linux user cannot or isn't willing to learn dd, maybe that user is better served by Windows or OS X. And really: I don't mean it in a bad way. I don't think Linux has to "win" the Desktop war. I don't think it's wrong not to use Linux.
But I do think it's wrong to believe users willing to learn Linux cannot learn the difference between input and ouptut.
Also a strong case can be made that someone for whom it's a "disaster" to overwrite it's main drive is one hard disk failure away from disaster anyway.
I'm for educating users, not baby-feeding them with tools that are going to keep them in the dark and reinforce their bad practices (like not doing proper backups and hence being "one hard drive failure" away from disaster).