Cloning a Laptop over NVMe TCP
copyninja.in
copyninja.in
On the destination laptop:
$ nc -l -p 1234 | dd of=/dev/nvme0nX bs=1M
On the source laptop: $ nc x.x.x.x 1234 </dev/nvme0nX
The dd on the destination is just to buffer writes so they are faster/more efficient. Add a gzip/gunzip on the source/destination and the whole operation is a lot faster if your disk isn't full, ie. if you have many zero blocks. This is by far my favorite way to image a PC over the network. I have done this many times. Be sure to pass "--fast" to gzip as the compression is typically a bottleneck on GigE. Or better: replace gzip/gunzip with lz4/unlz4 as it's even faster. Last time I did this was to image a brand new Windows laptop with a 1TB NVMe. Took 20 min (IIRC?) over GigE and the resulting image was 20GB as the empty disk space compresses to practically nothing. I typically back up that lz4 image and years later when I donate the laptop I restore the image with unlz4 | dd. Super convenient.That said I didn't know about that Linux kernel module nvme-tcp. We learn new things every day :) I see that its utility is more for mounting a filesystem over a remote NVMe, rather than accessing it raw with dd.
Edit: on Linux the maximum pipe buffer size is 64kB so the dd bs=X argument doesn't technically need to be larger than that. But bs=1M doesn't hurt (it buffers the 64kB reads until 1MB has been received) and it's future-proof if the pipe sizes is ever increased :) Some versions of netcat have options to control the input and output block size which would alleviate the need to use dd bs=X but on rescue discs the netcat binary is usually a version without these options.
Otherwise, i.e. when dd has to read the entire input file because there is no "count=" option, "iflag=fullblock" does not have any documented effect.
From "info dd":
"If short reads occur, as could be the case when reading from a pipe for example, ‘iflag=fullblock’ ensures that ‘count=’ counts complete input blocks rather than input read operations."
it pretty much seems iflag=fullblock is a requirement if you want the counts to work, even though the failure times might be rare
Larger writes will be more efficient, however, if only due to reduced system call overhead.
While not necessary when writing an image with the correct block size for the target device, even partial block overwrites work fine:
# yes | head -c 512 > foo
# losetup /dev/loop0 foo
# echo 'Ham and jam and Spam a lot.' | dd bs=5 of=/dev/loop0
5+1 records in
5+1 records out
28 bytes copied, 0.000481667 s, 58.1 kB/s
# hexdump -C /dev/loop0
00000000 48 61 6d 20 61 6e 64 20 6a 61 6d 20 61 6e 64 20 |Ham and jam and |
00000010 53 70 61 6d 20 61 20 6c 6f 74 2e 0a 79 0a 79 0a |Spam a lot..y.y.|
00000020 79 0a 79 0a 79 0a 79 0a 79 0a 79 0a 79 0a 79 0a |y.y.y.y.y.y.y.y.|
*
00000200
Partial block overwrites may (= will, unless the block to be overwritten is in the kernel's buffer cache) require a read/modify/write operation, but this is transparent to the application.Finally, note that this applies to most block devices, but tape devices work differently: partial overwrites are not supported, and, in variable block mode, the size of individual write calls determines the resulting tape block sizes.
How about `truncate -s 512 foo`?
Plus, a it'd be well deserved coffee break.
Considering I'd be going at GigE speeds at best, I'd add "oflag=direct" to bypass caching on the target. A bog standard NVMe can write >300MBps unhindered, so trying to cache is moot.
Lastly, parted can do partition resizing, but given the user is not a power user to begin with, it's just me nitpicking. Nice post otherwise.
dd is installed on every system, and if you don't have nc you can still use ssh and sacrifice a bit of performance.
dd if=/dev/foo | ssh dest@bar "cat > /dev/moo"I just copy that block device with "dd", that's all. It's just a dumb pipe encapsulated with TCP, which is already battle tested enough.
Moreover, if I have fatter pipe, I can tune dd for better performance with a single command.
I strongly recommend against oflag=direct as in this specific use case it will always degrade performance. Read the O_DIRECT section in open(2). Or try it. Basically using oflag=direct locks the buffer so dd will have to wait for the block to be written by the kernel to disk until it can start reading the data again to fill the buffer with the next block, thereby reducing performance.
I won't be bothered in a home network.
> Clonezilla are vastly more moving parts
...and one of these moving parts is image integrity and write integrity verification, allowing byte-by-byte integrity during imaging and after write.
> I strongly recommend against oflag=direct as in this... [snipped for brevity]
Unless you're getting a bottom of the barrel NVMe, all of them have DRAM caches and do their own write caching independent of O_DIRECT, which only bypasses OS caches. Unless the pipe you have has higher throughput than your drive, caching in the storage device's controller ensures optimal write speeds.
I can hit theoretical maximum write speeds of all my SSDs (internal or external) with O_DIRECT. When the pipe is fatter or the device can't sustain that speeds, things go south, but this is why we have knobs.
When you don't use O_DIRECT in these cases, you see initial speed surge maybe, but total time doesn't reduce.
TL;DR: When you're getting your data at 100MBps at most, using O_DIRECT on an SSD with 1GBps write speeds doesn't affect anything. You're not saturating anything on the pipe.
Just did a small test:
dd if=/dev/zero of=test.file bs=1024kB count=3072 oflag=direct status=progress
2821120000 bytes (2.8 GB, 2.6 GiB) copied, 7 s, 403 MB/s
3072+0 records in
3072+0 records out
3145728000 bytes (3.1 GB, 2.9 GiB) copied, 7.79274 s, 404 MB/s
Target is a Samsung T7 Shield 2TB, with 1050MB/sec sustained write speed. Bus is USB 3.0 with 500MBps top speed (so I can go %50 of drive speeds). Result is 404MBps, which is fair for the bus.If the drive didn't have its own cache, caching on the OS side would have more profound effect since I can queue more writes to device and pool them at RAM.
This matters in the specific use case of "netcat | gunzip | dd" as the compressed data rate on GigE will indeed be around 120 MB/s but when gunzip is decompressing unused parts of the filesystem (which compress very well), it will attempt to write 1+ GB/s or more to the pipe to dd and it would not be able to keep up with O_DIRECT.
Another thing you are doing wrong: benchmarking with /dev/zero. Many NVMe do transparent compression so writing zeroes is faster than writing random data and thus not a realistic benchmark.
PS: to clarify I am very well aware that not using O_DIRECT gives the impression initial writes are faster as they just fill the buffer cache. I am taking about sustained I/O performance over minutes as measured with, for example, iostat. You are talking to someone who has been doing Linux sysadmin and perf optimizations for 25 years :)
PPS: verifying data integrity is easy with the dd solution. I usually run "sha1sum /dev/nvme0nX" on both source and destination.
PPPS: I don't think Clonezilla is even capable of doing something similar (copying a remote disk to local disk without storing an intermediate disk image).
I noted that the bus I connected the device has 500MBps bandwidth theoretical, no?
To cite myself:
> Target is a Samsung T7 Shield 2TB, with 1050MB/sec sustained write speed. Bus is USB 3.0 with 500MBps top speed (so I can go %50 of drive speeds). Result is 404MBps, which is fair for the bus.
The 5Gbps ports are just marketed as "USB 3.1" instead of "USB 3.0" these days, because USB naming is confusing and the important part is the "gen x".
USB 3.0, USB 3.1 gen 1, and USB 3.2 gen 1x1 are all names for the same thing, the 5Gbps speed.
USB 3.1 gen 2 and USB 3.2 gen 2x1 are both names for the same thing, the 10Gbps speed.
USB 3.2 gen 2x2 is the 20Gbps speed.
The 3.0 / 3.1 / 3.2 are the version number of the USB specification. The 3.0 version only defined the 5Gbps speed. The 3.1 version added a 10Gbps speed, called it gen 2, and renamed the previous 5Gbps speed to gen 1. The 3.2 version added a new 20Gbps speed, called it gen 2x2, and renamed the previous 5Gbps speed to gen 1x1 and the previous 10Gbps speed to gen 2x1.
There's also a 3.2 gen 1x2 10Gbps speed but I've never seen it used. The 3.2 gen 1x1 is so ubiqitous that it's also referred to as just "3.2 gen 1".
And none of this is to be confused with type A vs type C ports. 3.2 gen 1x1 and 3.2 gen 2x1 can be carried by type A ports, but not 3.2 gen 2x2. 3.2 gen 1x1 and 3.2 gen 2x1 and 3.2 gen 2x2 can all be carried by type C ports.
Lastly, because 3.0 and 3.1 spec versions only introduced one new speed each and because 3.2 gen 2x2 is type C-only, it's possible that a port labeled "3.1" is 3.2 gen 1x1, a type A port labeled "3.2" is 3.2 gen2x1, and a type C port labeled "3.2" is 3.2 gen 2x2. But you will have to check the manual / actual negotiation at runtime to be sure.
It's not intended to be used by-design. Basically, it's a fallback for when a gen2x2 link fails to operate at 20Gbps speeds.
Fun fact: "openssl sha1" on a typical x86-64 machine is actually about twice faster than "sha1sum" because their code is more optimized.
Another reason I don't bother to use tee to compute the hash in parallel is that it writes with a pretty small block size by default (8 kB) so for best performance you don't want to pass /dev/nvme0nX as the argument to tee, instead you would want to use fancy >(...) shell syntax to pass a file descriptor as an argument to tee which is sha1sum's stdin, then pipe the data to dd to give it the opportunity to buffer writes in 1MB block to the nvme disk:
$ nc -l -p 1234 | tee >(sha1sum >s.txt) | dd bs=1M of=/dev/XXX
But rescue disks sometimes have a basic shell that doesn't support fancy >(...) syntax. So in the spirit of keeping things simple I don't use tee.Instead of the "fancy" syntax I used
mkfifo /tmp/cksum
sha1sum /tmp/cksum &
some_reader | tee /tmp/cksum | some_writer
Of course under the conditions mentioned throughputs were moderate compared to what was discussed above. So I don't know how it would perform with a more performant source and target. But the important thing is that you need to pass the data through the slow endpoint only once.Disclaimer: From memory and untested now. Not.at the keyboard.
dd followed by sha1sum on each end is still very few moving parts and should still be quite fast.
In a data center it's not (this is when I use clonezilla 99.9% of the time, tbf).
I guess most people don't have faster local network than an SSD can transfer.
I wonder though, for those people who do, does a concurrent I/O block device replicator tool exist?
Btw, you might want also use pv in the pipeline to see an ETA, although it might have a small impact on performance.
Besides, I don't think anyone really has a local network which is faster than their SSD. Even a 4-year-old consumer Samsung 970 Pro can sustain full-disk writes at 2.000M Byte/s, easily saturating a 10Gbit connection.
If we're looking at state-of-the-art consumer tech, the fastest you're getting is a USB4 40Gbit machine-to-machine transfer - but at that point you probably have something like the Crucial T700, which has a sequential write speed of 11.800 MByte/s.
The enterprise world probably doesn't look too different. You'd need a 100Gbit NIC to saturate even a single modern SSD, but any machine with such a NIC is more likely to have closer to half a dozen SSDs. At that point you're starting to be more worried about things like memory bandwidth instead. [0]
[0]: http://nabstreamingsummit.com/wp-content/uploads/2022/05/202...
Doesn't make sense at first glance. There's no head to move, as in an old-style hard drive. What else could make random write take longer on an SSD?
You can check literally any SSD benchmark that tests both random and sequential IO. They're both vastly better than a mechanical hard drive, but sequential IO is still faster than random IO.
So if large sequential writes mean that you only write full whole blocks, that can be done much faster than writing the same data in random order.
RMW cycles (though not in-place) are common for writes smaller than a NAND page (eg. 16kB) and basically unavoidable for writes smaller than the FTL's 4kB granularity
Any modern (last few decades) SSD has a lot of sauce on top of it to reduce this penalty, but random writes are still a fair bit slower - especially once the buffer of pre-prepared replacement pages runs out. Sequential access is also just a lot easier to predict, so the SSD can do a lot of speculative work to speed it up even more.
You might be surprised if you take a look at how cheap high speed NICs are on the used market. 25G and 40G can be had for around $50, and 100G around $100. If you need switches things start to get expensive but for the "home lab" crowd since most of these cards are dual port a three-node mesh can be had for just a few hundred bucks. I've had a 40G link to my home server for a few years now mostly just because I could do it for less than the cost of a single hard drive.
They're PCIe 3.0 x8 cards so they can't max out both ports, but realistically no one who's considering cheap high speed NICs cares about maxing out more than one port.
For example: https://unix.stackexchange.com/questions/632267
Note that you can increase pipe buffer, I think default maximum size is usually around 1MB. A bit tricky to do from command line, one possible implementation being https://unix.stackexchange.com/a/328364
Aside, I guess nvme-tcp would result in less writes as you only copy files in stead of writing the whole disk over?
Where does one learn this black art?
For compression the rule is that you don't do it if the CPU can't compress faster than the network.
dd if=/dev/device | mbuffer to Ethernet to mbuffer dd of=/dev/device
(with some switches to select better block size and tell mbuffer to send/receive from a TCP socket)If it's on a system with a fast enough processor I can save considerable time by compressing the stream over the network connection. This is particularly true when sending a relatively fresh installation where lots of the space on the source is zeroes.
https://www.techtarget.com/searchstorage/news/252459311/Ligh...
> The NVM Express consortium ratified NVMe/TCP as a binding transport layer in November 2018. The standard evolved from a code base originally submitted to NVM Express by Lightbits' engineering team.
https://www.lightbitslabs.com/blog/linux-distributions-nvme-...
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
nbdkit file /dev/nvme0n1
nbdcopy nbd://otherlaptop localfilePreviously I cloned but I this time I wanted to refresh some of the configs.
Using a usb-c cable to transfer at 10gb/s is so useful (as my only other option was WiFi).
When you plug the computers together they form an ad-hoc network and you can just rsync across. As far as I could tell the link was saturated so using anything else (other protocols) would be pointless. Well not pointless, it's really good to learn new stuff, maybe just not when you are cloning your laptop (joke)!
Serious question since last time I tried a direct non-Ethernet connection was sometime in the 90s ;)
With Thunderbolt and operating systems that support Ethernet over Thunderbolt, a virtual network adapter is automatically configured for any Thunderbolt connector, so connecting a USB C cable between 2 such computers should just work, as if they had Ethernet 10 Gb/s connectors.
With USB 3 USB C connectors, you must use USB network adapters (up to 2.5 Gb/s Ethernet).
It merely requires one side to be capable of behaving as a device, with the other side behaving as a host.
I.e., unlike PCIe hubs, you won't get P2P bandwidth savings on a USB hub.
It just so happens that most desktop xhci controllers don't support talking "device".
But where you can, you can set up a dumb bidirectional stream fairly easily, over which you can run SLIP or PPP. It's essentially just a COM port/nullmodem cable. Just as a USB endpoint instead of as a dedicated hardware wire.
It's very useful for sending large data using minimal equipment. No need for two cat6 cables and a router for example.
Secondly, I'm not sure if a crossover cable setup will autoconfigure the network, as the poster above says, it has been since the 90s when I bothered trying something like that!
You don't need crossover cables anymore. You can just connect a regular patch cable directly between 2 devices. Modern devices can swap RX/TX as needed.
As for auto-configuration, that's up to the OS, but yeah you probably have to set up static IPs.
I was under the impression you have to set up the devices as each other's default gateway, but maybe I'm the one not up to modern standards this time.
Therefore they can talk directly, without using a gateway.
Their corresponding Ethernet MAC addresses will be resolved by ARP.
The problem is that in many cases you would have to look at the autoconfigured addresses and introduce manually the peer address in each computer.
For file sharing, either in Windows or using Samba on Linux, you could autodiscover the other computer and just use the name of the shared resource.
The transfer itself screamed and I had a terabyte over in a few mins. Also I didn't bother with encryption on this one, so that simplified things a lot.
There's one caveat though: With btrfs it is not possible to send snapshots recursively, so if he had lots of recursive snapshots (which can happen in Docker/LXD/Incus), it is relatively hard to mirror the same structure in a new disk. I like btrfs, but recursive send/receive is one aspect ZFS is just better.
Btw, what kinds of guarantees do you get with the dd method? Do you have to compare md5s of the resulting block level devices after?
Rsync cannot transfer more than one file at a time so if you were transferring a lot of small files that was probably the bottleneck. You could either use xargs/parallel to split the file list and run multiple instances of rsync or use something like rclone, which supports parallel transfers on its own.
6 hours is roughly 10MB/s, so you likely could have gone much much quicker. Did you compress with `-z`? If you could use ethernetyou probably could have done it at closer to 100MB/s on most deviceds, which would have been 35 minutes.
One thing I could have done is found a way to track total progress, so that I could have noticed that this is going way too slow.
Elsewhere I recommended the Thunderbolt4 cable—it screamed at ~10GBs and didn't need any optimization like tar/compression.
> No, I didn't use compression. Would it be useful over a high-bandwidth connection?
Compression will be useful if you're IO bound as opposed to CPU bound, which at 10MB/s you definitely are.
> --adapt[=min=#,max=#]: zstd will dynamically adapt compression level to perceived I/O conditions.
> I wonder what could I have done better.
Used an Ethernet cable? That’s not an impressive throughput amount over local. WiFi has like a million more sources of perf bottlenecks. Btw, just using a cable on ONE of the device => router ~> device can help a lot.
dd doesn't skip empty blocks, like clonezilla would do.
"Carrier-sense multiple access with collision avoidance (CSMA/CA) in computer networking, is a network multiple access method in which carrier sensing is used, but nodes attempt to avoid collisions by beginning transmission only after the channel is sensed to be "idle".[1][2] When they do transmit, nodes transmit their packet data in its entirety.
It is particularly important for wireless networks, where the alternative with collision detection CSMA/CD, is not possible due to wireless transmitters desensing (turning off) their receivers during packet transmission.
CSMA/CA is unreliable due to the hidden node problem.[3][4]
CSMA/CA is a protocol that operates in the data link layer."
https://en.wikipedia.org/wiki/Carrier-sense_multiple_access_...
With the author not having access to an ethernet port on the new laptop, I think my hacky approach might've even been faster because of the slight boost compression would've provided, given that the network speed is nowhere near the speed limit compression would add to a fast network link.
I think I did something like `nc -l -p 1234 | gunzip | dd status=progress of=/dev/nvme0n1` on the receiving end and `dd if=/dev/nvme0n1 bs=40M status=progress | gzip | nc 10.1.2.3:1234` on the sending end, after plugging an ethernet cable into both devices. In theory I could've probably also used the WiFi cards to set up a point to point network to speed up the transmission, but I couldn't be bothered with looking up how to make nc use mptcp like that.
True, I do always just take the NVME disk out of the laptop and put it in a highspeed dock.
Still, if I was the planning ahead kind then using something more declarative like NixOS where you only need to copy your config and then automatically reinstall everything would probably be the better approach.
LUKS2 does not care about the size of the disk. If it's JSON header is present, it will by default treat the entire underlying block device as the LUKS encrypted volume/partition (sans the header), unless specified otherwise on the commandline.
Amazing software (transfer performance wise) written (unfortunately in) Java :) Unintuitive CLI options, but fastest transfer I have seen.
So fast, if I don't artificially limit the speed, my command would sometimes hog the entire local network.
-limit <rate> Restrict the transfer speed at the specified rate. K (KiloBytes/s), M (MegaBytes/s) or G (GigaBytes/s) may be used as suffixes.
[1] http://monalisa.cern.ch/FDT/download.html
PS: It does cause file-fragmentation on the destination side, but it shouldn't matter to anyone, practically.
This reminds me one of my favorite articles
https://blog.codinghorror.com/the-infinite-space-between-wor...
Windows would panic, certainly (because so much drivers & other state is persisted & expected), but the Linux kernel when it boots kind of figures out afresh what the world is every time. That's fine.
The main thing you ought to do is generate a new systemd/dbus machine-id. But past this, I fairly frequently instantiate new systems by taking a btrfs snapshot of my current machine & send that snapshot over to a new drive. Chroot onto that drive, and use bootctl to install systemd-boot, and then I have a second Linux is ready to go.
The system booted, no license issues.
Edit: or backup process
When copying files (e.g. with rsync) you need to watch out for those, yes. The best way to deal with other mounts is not to copy directly from / but instead non-recursively bind-mount / to another directory and then copy from that.
Nope.
7B is because back in WinNT days it made sense to disable the drivers for which there are no devices in the system. Because memory and because people can't write drivers.
Nowadays it's AHCI or NVMe, which are pretty vendor agnostic, so you have 95% chance of successful boot. And if you boot is successful then Windows is fine and yes, it can grab the remaining drivers from the WU.
There may be an occasional exception for when you've doing something weird (iirc the 11th-13th gen Intel RST need a slipstreamed or manually added drivers unless you change controller settings in the BIOS which may bite on laptops at the moment if you're unaware of having to do it).
But even for big jumps you can usually get it working with a bit of hackery pokery. Most recently I had to jump from a legacy C2Q system running Windows 10 to run bare metal on a 9th gen Core i3.
I ended up putting it onto a VM to run the upgrade from legacy to UEFI so I'd have something that'd actually boot on this annoyingly picky 9th gen i3 Dell system, but it worked.
I generally (ab)use Macrium Reflect, and have a copy of Parted Magic to hand, but it's extremely rare for me to find a machine I can't clone to dissimilar hardware.
The only one that stumped me recently was an XP machine running specific legacy software from a guy who died two decades ago. That I had to P2V. Worked though!
These days though I've had a number of cases where some very humble small scoped bios change will cause windows to not boot. I've run into this quite a few times, and it's been quite an aggravation to me. I've been trying to get my desktops power consumption down & get sleep working, mostly in Linux, but its shocking to me what a roll-of-the-die it's been that I may have to do the windows-self-reinstall, from changing an AMD cool-n-quiet settings or adjusting a sleep mode option.
If you do things the Microsoft way, you just sign into your MS account and your files show up via OneDrive, your apps and games come from the Microsoft store anyway.
There's plenty to fault Microsoft for. Like making it progressively more difficult to just use local accounts.
I don't think "cloning a machine from hardware a decade old onto a brand new laptop may require expertise" is one of them.
and happen to live near good bandwidth.
> Like making it progressively more difficult to just use local accounts.
right. because they don't care about your hardware just your payment relationship with them.
> cloning a machine from hardware a decade old onto a brand new laptop may require expertise
you don't see the connection between all these facts? "coping all the work and effort you've collected on your personal machine for a decade _still_ unaccountably requires expertise."
If this were something so unusual that people would rarely want to do it this would be a reasonable point. The fact that this is such an obvious thing people want to do without having to have a monthly subscription first indicates that the product does not cater well to it's selected market segments.
That said, I didn't know you could export NVMe over TCP like that, so still a nice read!
The only problem with that approach is that it is also copies over the .config and .cache folders, most of which are possibly not needed anymore. Or worse, they might contain configs that overrides better/newer parameters set by the new system..
Yeah there are much simpler approaches. I don't bother with container images but can install a new machine from an Ansible playbook in a few minutes. I do then have to copy user files (Borg restore) but that results in a clean new install.
Buy a USB to Ethernet adapter. It will come in handy in the future.
Not so many years ago doing something similar (exporting a block device over network, and mounting it from another host) would have meant messing with iscsi which is a very cumbersome thing to do (and quite hard to master).
That's because bs=40M and no status=progress.
Is there any way to mount a remote volume on Mac OS using NVMe-TCP?
I restore my working setup to a fresh install within minutes.
Is that a mix of conclusion and concussion that requires some ibuprofen ?
I'll be using GRML again tonight to rescue some laptop drives which were retrieved from a fire, to see what can be salvaged, for example. First in forensic mode to see what's possible, then to use dd_rescue on a N95 box.
I just take my nvme drive out of the last one and I put it into the new one.
I don't even lose my browser session.
Edit: If you do this (and you should) you should configure predictable network adapter names (wlan0,eth0).
I've bought a few hard disks in my lifetime. Many years ago one such disk came bundled with Acronis TrueImage. Back in the day you could just buy the darn thing, after certain year they switched it to subscription model. Anyhoo.. I got the "one-off" TrueImage and I've been using it ever since. I've 'migrated' the CD (physical mini-CD) it to a USB flash disk, and have been using it ever since.
I was curious for the solution as I recently bought a WinOS tablet, and would like to clone my PC to it, but this looks more like one too many hoop jumps to something that TrueImage can do in 2h hours (while watching a movie).