Defrag Like It's 1993
defrag.shiplift.dev
defrag.shiplift.dev
Great times!
I think I might have been able to get something in the local public university computer centre. It was the beginning of the internet (14.4 modems min to connect to the university, which was the only internet provider)
Great fun times!
Some PCs just last. A 2.4 GHz Core2 Quad I had, I think from 2008, is still going strong and can run most modern software just fine.
It's amazing how little progress there has been in last 13 years. In the nineties you had to upgrade every two years or so just to be able to run current software.
I think 2021 computer has a good chance of being still completely usable in 2040.
I'm currently building my next computer, and the difference again is just staggering. And that is not even factoring in that the next one has got 12-cores.
It is not that the progress hasn't been slow (I guess it has if you compare it to the pace of the past), but I think it has more to do with computers back then still being very capable.
You can still fire up some CAD software and do proper work on the Q6600. I mean, right until you want to listen to music or do some very light web browsing. Then it will suddenly become quite painful. The amount of waste in today is incomprehensible. No seriously, properly incomprehensible.
I do think I could still manage on Q6600 with 16 GB RAM and an SSD. Might need to downscale VMs and browser tabs, but manageable in a pinch. Heck, in the early nineties I was happy to have an Amiga with 1 MB of RAM and that thing flew in comparison to C64. :-)
I recently set-up my old Q6600 as a station for CNC machinery, that thing flies with LinuxCNC on a small SSD. There's no issue running firefox and a music player at the same time (though I tend to not have other things running when milling something but more by cargo-cult than actual tests of the impact on the real-time behaviour needed for the CNC machine control). Some operations on Inkscape are of course slower than on more recent hardware but otherwise it's still very useable.
And of course you can surf the web. But even with 5GHz turbo it is still slower than a fast connection was in the late 90s (on 90s hardware and browser). And it has nothing to do with that progress hasn't advanced in that time (we are talking orders of magnitude).
Absolute garbage, and entirely representative of all other consumer class software.
In reality though, the increase in performance was barely noticeable.
But I’d be interested in any real data anyway.
Most of the time however, my disk was probably only about 30% fragmented (if I recall correctly from Speed Disk). I was running DOS on a FAT file system where there was no swap file or anything that would cause massive fragmentation except deleting and installing programs, and maybe some temp files created by certain programs. The performance delta for me was barely noticeable on a 40MB Seagate IDE drive.
Why not?
$ virt-make-fs --type=msdos /usr/share/doc /var/tmp/disk.img
$ export file=/var/tmp/disk.img
$ nbdkit eval \
after_fork=' echo 0 > $tmpdir/read ' \
get_size=' stat -Lc %s $file ' \
pread='
diff=$(($4 - `cat $tmpdir/read`))
echo $4 > $tmpdir/read
if [ $diff -lt -10000 ] || [ $diff -gt 10000 ]; then
# simulate a seek
sleep 0.1
fi
dd if=$file skip=$4 count=$3 iflag=count_bytes,skip_bytes
'
$ sudo nbd-client localhost /dev/nbd0
$ sudo mount /dev/nbd0 /tmp/mnt
The tricky thing is probably making a deliberately fragmented disk image for testing using modern tools.I would use process monitor from sysinternals to measure which files are loaded as the system boots. Then use jkdefrag/mydefrag to take said list and rebuild drive into different zones. Try to put small files that are used together in the same area of the disk. Put huge things like 'isos/vhds/acrivezips' etc near the end of the drive. Put small frequently used files near the front in used order if possible. Smash out any gaps if possible. This helps NTFS from creating more fragments. Defrag any files that 7zip unarchives immediately as it has a bad habit of creating highly fragmented files (contig from sysinternals). Try to put file index in its own 'area' as one reason it runs so badly is the drive fragmentation it causes to itself. It got so bad I would put the index on its own partition sometimes. Try to put files that fragment frequently in their own 'zone' as NTFS will tend to pick the next available gap to where the file is (not always).
That made a real noticeable difference in perceived speed.
These days with nvme and ssd I might just defrag the files themselves once a year if I feel like it. There is a difference but the perceived is pretty much not there anymore. The difference is there is a chain of sectors in the MFT that windows will have to traverse if you ask for something in the middle of a file. That is about the only thing you save with newer drives. ext4 has a similar issue. But it seems to be better about picking spots where it will not fragment. It will however fragment badly in low space conditions. Most filesystems will.
First thing the middle of the volume/partition is not the middle of the disk in any multi-partitioned disk, then NTFS $MFT is not (and never has been) in the middle of the partition, its location varies depending on the size of the volume/partiton but above a certain size, around 5/6 GB it is at a fixed offset, usually on LCN #786432 which - on a normal 8 KB/cluster NTFS volume - amounts to offset 6,291,456 sectors or 6,442,450,944 bytes, rarely the middle of the volume, unless its size is around 13 GB.
Seems to me the MFT was in the middle of volumes commonly in older versions, and lately much closer to the beginning of the volume.
Could have something to do with the feature of shrinking a volume which became more common, and you can now more often shrink an NTFS volume by more than half, which was often impossible in earlier years.
If you "think in hex" the address is in the PBR expressed in clusters (VCN) as "0xC0000" (at offset 0x30 or 48 dec you should find "00000C000000000"), which is a nice, round number.
Maybe you remember the "old" NTFS format (NT 3/4) where the mirror of the bootsector (aka $BootMirr, not the $MFT)was exactly in the middle of the volume, but since Windows 2000 this copy of the first sector of the PBR was moved in the "gray zone" at the end (inside the partition but outside the volume[1]).
As a side note (and JFYI) thanks to (from 7 onwards) NTFS resizing capabilities it is possible to "force" the $MFT to very early sectors, see this only seemingly unrelated thread here:
http://reboot.pro/index.php?showtopic=18022
[1] some info on the matter in this other, as well seemingly unrelated, thread:
This now makes me wonder whether the middle disk block (numerically) is equal "seek times" away from disk edges.
SSD drives don't need to be defragged.
But yes these designs generally seemed to be much better than FAT under DOS or Windows at avoiding excessive fragmentation.
And then hiring a personal chef for that person. Sure, the food will taste better, but at what cost? At what cost?
(Actually watching defrag stopped being satisfying around the time that I got my first 50GB HDD. It took too damn long to watch, like a 2 year old playing Tetris in slow motion)
Well, with this new website, you can now enjoy defraging without having to suffer the consequences of a bad file system.
Maybe I should just defrag my SSD's these days. Sure they don't need it's but it's fast and bad for the drive, which gives it that same thrill of gambling with the integrity of my data. Of course I'd need to disable cloud backups to backblaze and local backups to a raid array, but if I try then I think I can still chase the thrill of the 'frag.
If you use a filesystem with checksums, you can also scrub it. The stakes are there: is my data bit rotted? Will the errors resist recovery? It also takes about the same time as the defrag, but again I don't know of any graphical tools.
xfs_fsr -v*edit: not entirely true: the Seagate MACH.2 drive series that has 2 separate head assemblies per drive. They're pretty expensive and hard to get ahold of.
Even "worse", now I an translating some English text adventures into Spanish.
And by playing them OFC.
Today I have more fun doing actual projects (programming, writting, playing with teleco stuff) than badly simulating then in video games.
With DOS/Windows you simulated a life on games. With Linux, you made it real. Seriously. Slackware it's still a "game changer. You don't play Cyberpunk 2077 if you want, you can make it real life. I do that with the PocketCHIP, Gopher and text gaming/roguelikes/MUDs up into a mountain at night. Incredible experience.
Or better, listening to the ISS with a WebSDR and decoding SSTV images with QSSTV. That's the ultimate cyberpunk experience.
In the rare circumstances when I was in need of actual defrag of a heavily polluted drive > 40Gb I opted to take a filesystem aware image with Norton/Symantec Ghost and just re-image it back. Worked fine and by the time of 100-200Gb drives there were a USB HDDs so I wouldn't even need to unscrew the PC case.
A prof had to ask me to set it back because they couldn’t figure it out.
Norton Utilities were slick. Slicker than any MSDOS tools, or even Apple Dos 3.0 or ProDOS (or ProntoDOS for that matter) utilities for that matter.
Seems like Joe has the helpful menu system, but it feels overall more "serious" in comparison, especially with the minimal white on black color scheme.
A perfect clone of the MS-DOS editor for UNIX/Linux does not exist. Maybe this is the new toy project for the "written in Rust" crowd.
Then it could not only regroup files into contiguous blocks, but it could put the most stable files first so that subsequent defrags are faster!
You probably want a few alternating stripes of files and free space. Or do what most people did, which is keep a mostly empty hard drive.
Online gaming with my Dreamcast
The other win is putting files accessed together (like when the computer is booting) close to each other to minimize seeks. I suppose the actual goal of defragmenting is minimizing seek time.
"By clicking Accept, you agree to share the information about your files with us and any relevant third parties [read: anyone who wants to pay us money to get info on you]".
Thank you for spending such quality time with your son
sudo e4defrag /
sudo dd if=/dev/zero of=/zeros; sudo rm -f /zeros
While write intensive, I run this on my ZFS VMs on occasion to keep the zstd compression ratio high. No pretty interface, but it's still satisfying when I see the volume's compression ratio is close to entropy. Even in 2021. Even with all-flash.sudo bash -c 'dd if=/dev/zero of=/zeros; rm -f /zeros'
If you really need to defrag SSD (which somewhat helps IME if it was heavily used up to 100% capacity), copy everything to other storage, wipe it (preferably with NVMe/SATA commands, or at least do a full TRIM), and then copy everything back.
i know that per SMART, disks (hdd and ssd) have to have extra space that will be used if an active sector goes bad, but idk about the total internal remapping on ssd's.
The issue TRIM helps with is the following: while reads and writes are performed at the page level (2-16k), a page can not be overwritten (it can only be written to when empty) and an SSD can only erase entire blocks (128~256 pages).
This means when you perform an overwrite of a page, the SSD controller really has an internal mapping between logical pages (what it tells the FS) and physical pages (the actual NAND cells) and it updates the mapping over the overwritten logical page to a new physical page (in which it writes the “updated” data), marking the old physical page as “dirty”.
So as you use the drive the blocks get fragmented, more and more full of a mix of dirty and used pages meaning the controller is unable to erase the block in order to reuse its (dirty) pages. So it garbage-collects pages, which is a form of completely internal defragmentation: it goes through “full” blocks, copies the used pages to brand new blocks, then queues the blocks for erasure.
The issue with this process is… the SSD only knows that a block is unused when it’s overwritten, historical protocols had no more information since the hard drives had the same unit for everything and could overwrite pages, and thus the FS managed that directly. This means if you delete a file the controller has no idea the corresponding pages are unused (dirty) until the FS decides to write something unrelated there, meaning if you do lots of creation and removing (rather than create-and-never-remove or create-and-overwrite) the controller lags behind and has a harder and harder time doing physical defrag and having empty blocks to write to, this leads to additional garbage collection and thus writes, and all the “removed but not overwritten” blocks are still used as far as the controller knows so they’re copied over during GC even though no one can or will ever read them.
TRIM lets the FS tell the controller about file deletions (or truncation or whatever), and thus allows the FS to have a much more correct view of the actually used blocks, thereby allowing faster reclamation of empty blocks (e.g. if you download a 1GB file, use it, then remove it, the controller now knows all the corresponding blocks are dirty and can be immediately queued for erasure) and reducing unnecessary writes (as known-dirty pages don’t have to be copied during GC, only used pages).
The problem that was described was reclaiming space from a VM. I took the comment about compression ratios to be a measurement of the entire VM disk image, because otherwise if you just want a big meaningless ratio then dump a petabyte of zeroes into a new file.
A VM that isn't set up to TRIM will leave junk data all over its disk image, bloating it by a lot. If the guest OS understands how to TRIM, and the VM software properly interprets those TRIM commands, then it can automatically truncate or zero sections of the disk image. If either of those doesn't understand TRIM, then you need to fill the virtual disk with zeroes to get the same benefit. (And possibly run an extra compaction command after.)
This is at a completely different layer from what you elaborated on, because it's a virtual drive. It's good to have virtual drive TRIM even if you're storing the disk image on an HDD! It's a different but analogous use case to real physical drive TRIM.
For such a long time (as in, well into the tail end of Windows XP) I was stuck with an old Pentium computer that could handle Windows 98SE at most. Had maybe 2GB of HDD to work with which sounded luxurious to me until my father was issued a 4GB flash drive at work. I defragged this every first Sunday of the month to keep things running smooth. Ofc the effect might as well be purely psychological but hey, as I said it did last me a long time.
The mechanical sounds I remember are finer, not like pop snoring no? :)
Every once in a while I hear my AIO spin up and some water start moving around and that's nice (not sure if it's supposed to do that but it's done it since day one). But that's all I have to look forward to other than the fan noise.
Someone also wrote a script to start killing processes before the system can hang when it runs out of RAM.
Isn't this exactly what the OOMKiller does?
Makes you wonder...
https://askubuntu.com/questions/1258371/oom-killer-never-run...
That's why some orgs implemented their own solutions to avoid OOMKiller having to enter the picture, like Facebook's user-space oomd [1] or Android's LMKD [2]
https://devblogs.microsoft.com/oldnewthing/20211111-00/?p=10...
I remember the whole experience was cathartic, and my DOS PC definitely felt faster after doing (though it most likely wasn't).
"Unused" blocks is a probably misleading terms - "partially used" is more precise.
The simulation performs significantly more reads from "Unused" blocks, so I'm not sure, if it's exaggerated/incorrectly modeled, of if the virtual disk is an edge case (I suspect the former case).
Here's another video, with the same behavior: https://www.youtube.com/watch?v=syir9mdRk9s.
Have you thought about using <canvas> and just writing OEM DOS style characters where they need to be like some sort of obscene franken-video-buffer?
You've got contiguous space, speed, total time, wear.
I can't remember which it is, but either the center or the outside edge of a disk read faster, so another common trick was to place all the files necessary to boot the system there, to reduce boot times.
So far as the algorithm, what I observed is the program would look at how much room was needed to place the next file, then clear that many blocks out of the way by moving them to the end of the drive. Then copy the blocks for the file to form a contiguous stretch of blocks. And repeat. You wanted to have as much free space as the largest file (plus a little), but I think some of them were able to move large files in pieces.
Back in the days of Windows NT we had a network share with 8+ million files on it (a lot for the time) and we had a serious fragmentation problem, where it could take a second or more to read a file. Most of the defraggers just gave up, or weren't making any progress, but we eventually found one that worked (PerfectDisk, maybe?)
With PerfectDisk[1] you can decide how you wanted the disk laid out. And it did make a measurable difference in the cases I tested.
For SSDs fragmentation is less of an issue, but it's not completely gone. Especially on Windows with NTFS, where one can run into issues with heavily fragmented files[2][3].
We actually have this very issue with a client in production these days, due to the DB log file causing heavy fragmentation.
[1]: https://www.raxco.com/home/products/perfectdisk-pro
[2]: https://support.microsoft.com/en-us/topic/a-heavily-fragment...
[3]: https://support.assurestor.com/support/solutions/articles/16...
Oh wait, we don't have screensavers anymore either...
disclamiar: did write benchmarks with sd-cards in the past and was astonished by the failure rates. There also was (is?) little correlation to price, seemed like there was a huge variance in the quality of different batches of memory chips.
I don't really know what to do with all my old hard drives. I love SSDs because they don't vibrate and make noise.
i REALLY wonder if a SSD would still work after lying unpowered for ten years, they are made of capacitors after all ...
That’s less work and somewhat easier to program. There still will be edge cases where doing that increases the size of directory structures, so it’s not fully trivial.
If I remember from the last time I used it, it is not nice about it. It may create fragments. It consolidates the free space to the end of the drive.
Does not always work. Especially if the drive in question is active.
TBH I'm a bit uncertain on how today's GRUB goes from the boot sector, which is too small to understand filesystems, to a beast that can load the kernel from practically any filesystem.
I think that GRUB 1 used the 32kB DOS compatibility region to store whatever didn't fit in the MBR. The 32kB was enough for it to boot itself to a state where it understands filesystems.
For UEFI boot, it's the firmware's job to understand the partition table and FAT32 so that it can open up the boot partition and make a list of bootable files. Then it can run one or give the user a menu. So GRUB just has to put a single file in the right directory.
I was actually just thinking about this this morning. After finally getting past single and dual floppy systems I used to defrag my 40MB disk every night and felt rightness with the world. I remember wishing I could choose which files go where because of outer parts of the platter being faster than the inner parts. Whether that was true in practice didn’t matter, just that feeling of fully tuning my machine.
http://csis.pace.edu/~jyuan2/paper/betrfs5.pdf
"In this article, we demonstrate that modern file systems can still suffer from fragmentation under representative workloads, and we describe a simple method for quickly inducing aging.
Our results suggest that fragmentation can be a first-order performance concern—some file systems slow down by over 20x over the course of our experiments. We show that fragmentation causes performance declines on both hard drives and SSDs, when there is plentiful cache available, and even on large disks with ample free space."
Apparently they've never let a filesystem get above 70% full or so, especially one on a fileserver that has seen years of daily use by a dozen+ people.
I've heard it for HFS, HFS+, XFS, ext3, ext4, etc and every time it was an lie. Every filesystem fragments after regular use unless you have a fuckton of free space.
It's pretty hard to justify to one's bosses over-provisioning disk space to enough of a degree to render fragmentation unlikely, and users just expand their data to fit it anyway.
hmm that does not seem right and off by an order 10x...
I mean magnetic hard drives are very similar to vinyl recordings, so it's clear why contiguous file placement has lower latency, and thus why defragmentation is a big performance boost. I believe since SSDs are organized differently (many NAND chips), their access latency doesn't benefit as much from a contiguous file placement.
Furthermore, defrag necessarily has to write. Since SSD cells have a finite number of writes before they do not retain data, defragmentation shortens your SSD's lifecycle.
You phone works simple by asking for data from addresses. Not saying it’s perfect. Files can be fragmented so you may send a few more requests to get them. But it’s not something to lose sleep or run a utility on.
EDIT: I realize that it needs the sounds of the hard drive itself being defraged.
Literally nearly an entire work day lost spent waiting for defrag to finish on my FAT32 hard drive. We used some program with a very minimal UI that was supposed to do it faster/better than the utility built into Windows.
But I'll still take ZFS over any other file system :)
Defrag is relaxing and hands off, but reverse engineering is a hard drug.
I remember now, you could even hear the reverberation in the metal case!