A FAT32 fragmenter
gist.github.com
gist.github.com
I tested it on floppy disks - if you let it run for a while, the remaining free space was so fragmented it was almost unusable. You could hear the poor floppy drive seeking like crazy just to open a tiny text file.
The fun part was defragmenting it afterwards - I miss the days of the graphical defrag that Norton and MS had.
Then it would highlight a span of blocks to indicate reading.
Finally it would write those clusters elsewhere, indicating so with a different color.
Edit: Uhh.. and HN deleted my diagrams. See http://pastebin.com/raw.php?i=nKHddiKe
http://ultradefrag.sourceforge.net/en/index.html
Seems it even offers a terminal interface for the real nostalgic kick.
People often claim that fragmentation doesn't affect SSD drives, but that's not true: http://www.sami-lehtinen.net/blog/ssd-file-fragmentation-myt...
This is slightly related. How contiguously growing files are allocated on different file systems: http://www.sami-lehtinen.net/blog/test-btrfs-ext4-ntfs-simpl...
Relatively speaking the SSD slows down to ~50% of its max speed with fragmentation, but the HDD will be down to around 1%. So fragmentation affects SSDs somewhat, but HDDs are affected much more severely.
(Which SSD was it, and how was the test setup? That's important for comparison purposes, as the layout of the filesystem blocks and how they correspond to the NAND eraseblocks/pages has a huge effect on what fragmentation will do.)
Currently it has good read support and limited write support (can re-write existing files but not append or create new files).
Before writing that code I made some prototype read-only code in Python, to make sure I understand the FS structure properly: https://github.com/ambrop72/aprinter/blob/master/prototyping...
What did you use as your main source of documentation while developing?
I've written a bit more about that before:
https://news.ycombinator.com/item?id=7492318
Append = find an empty cluster and link it into the chain for the file, then write into it. Create is similar except you start with a empty chain and also add an entry to the directory (which is like writing/appending to a file, because directories are files.)
When allocating next clusters you can reduce fragmentation significantly over the dumb "first fit" strategy if you scan the FAT to find contiguous free clusters and apply a next/best/worst fit. Even better if you add an API to allow tuning the amount of "gap" after a file when creating it, based on knowledge of how much it may expand in the future.
I believe that with good tuning and allocation heuristics, FAT32 can outperform the far more complex filesystems (ext*, NTFS, etc.) widely believed to be superior. One idea I've never gotten around to testing out is to modify the Linux FAT driver and do some benchmarking.
I believe that with good tuning and allocation
heuristics, FAT32 can outperform the far more complex
filesystems (ext*, NTFS, etc.) widely believed to be
superior. One idea I've never gotten around to testing
out is to modify the Linux FAT driver and do some
benchmarking.
I don't understand. Has anybody ever claimed that the more complex alternatives were actually faster?NTFS and other filesystems that followed FAT32 are "superior" because they support things like journaling and more robust permissions... things that unavoidably incur (a least) a small performance hit.
One reason for the RAM needs is that I use a block cache and I take care do to "atomically" certain related operations (on the block cache level). These are things like marking a cluster as allocated, decrementing the number of free clusters in FsInfo, and linking the allocated cluster to the end of a chain. These things require multiple blocks to all be present in the block cache at the same time.
I suppose the asynchronous design also contributes to the code and RAM size (though I did not actually measure the code size or compare it to other drivers). It may look like more that it really compiles to. Anyway would be very surprised if you show me another asynchronous driver - I couldn't find any!
Also consider that my code supports VFATs (decoding only, file creation is not supported anyway).
The design is intentional, I have sacrificed some RAM/program size for asynchronous operation and better reliability in case of random write failures (because of the block cache, all writes can reliably be retried later, avoiding corruption as long as they eventually complete).
Update: I estimate the code compiles to about 10kB with gcc for Coretex-M3 using -Os. This is certainly within limits of most ARM uCs and many AVRs.
His has better status information.
It's pretty horrendously bad on spinning rust - you get very, very close to 100% fragmentation, so the disk has to seek for each cluster. A lower bound average on seek time is probably something like 4ms, so your max transfer rate is going to be limited to 125 * cluster size. Depending on the disk size[0] the cluster size may be up to 32KB, so we're looking at maybe 4MB/sec best case. I'd guesstimate it'd increase boot time by somewhere between 10x and 40x.
EDIT: Just realized that I don't think this can be done through normal means. I think it'd need a pathological FAT file system driver to do it, and you'd have to build the image ahead of time to do it.
Now that i would loved to be witness to.
But it depends on how large a block is on the SSD.
Granted that would not work out for SSD, but I know next to nothing on their addressing
In general no, at least not since the early 80s. Although you still hear about cylinders, heads, and sectors, spinning hard drives have effectively been a black box to OSes for 30 years and will internally remap sectors wherever they feel like it (and lie about it if you ask them).
The result is that while you can be reasonably confident that sector 33566 is followed by sector 33567 and reading both will not involve seeking, things like reading the whole cylinders at a time are not worth the effort since you don't know where the sectors are.
Screw it, I can't hear your rules over how awesome this is.
Why do you say that? Of course it isn't.
If making a negative comment you should explain, i.e. "I think this is bad because ..." rather than "This is bad.". If you have a valid point you are helping by raising it.
Similarly when retorting to a negative point with a positive view, "I disagree because ... " is good but just "You are wrong." simply unhelpful noise.
So feel free to provide positive encouragement when you think something is awesome!