Swap on HDD: Does placement matter?
vidarholen.net
vidarholen.net
[1] Example from 2007 https://www.linuxquestions.org/questions/debian-26/debian-in... [2] Example from around 1997 https://tldp.org/HOWTO/html_single/Partition/#SwapSize
I also (vaguely) remember some people putting build partitions closer to the front.
still avoid hitting that thing at any cost. Memory is pretty quick stuff.
If you assume that swap is a crutch the ideally won't be used or if it is used it is either for a short period only (due a to temporary overallocation) or for pages that are very rarely (if ever) used again (chunks of code & data that get loaded by then only certain configurations ever touch again), then you want to keep the fastest part of the drive for things that are going to be assessed regularly (your root partition for instance) in normal operation. For the occasional write & read of swap it makes little difference, and once you are properly thrashing pages to & from swap the time cost of head movements completely dwarfs any difference made by the actual location of the swap area (the heads will be spending most of their time in/near it anyway in such circumstances).
If you were relying on swap for general operations because the amount of RAM you'd need otherwise was just far too expensive, then you have a workload that warrants custom partitioning, to put it elsewhere but the end or ideally on another drive if you could afford a second.
If speed is an issue then you want it near the most commonly accessed data. Back when I used to have to think about these things much at all my general default arrangement was “boot, LVM” and within LVM “root, var, swap, homes, other data”. Swap being in the middle makes resizing in-place something I wouldn't generally consider, but if I needed more temporarily the extra would be created as a swap file (with lower priority than the partition) instead and/or better on a different drive (with higher priority, moving the main swapping load off the system drive).
Another, though less commonly useful, reason might be because it is easier to resize that way: if you need more than shrink the filesystem and add an extra swap area in the newly freed space.
> that also made moving an existing installation to a larger disk more complicated, since you couldn't just resize the os partition, you had to delete and then recreate swap
That isn't really a significant issue though, you shouldn't need swap while performing that operation (unless you are somehow moving the root filesystem around live) so stopping swap isn't going to be a problem (and a user capable of safely performing such a move at all will be able to handle the three extra commands needed). Assuming that you move everything first then resize, my preference would instead to be to move and resize individual partitions instead of moving everything so swap doesn't need to be moved and resized at all.
Yes. You expect the seek time to dominate performance.
The reason that the swap was faster when placed at the beginning is likely because the filesystem is mostly empty and so the allocated portion is at the beginning of the partition.
If the filesystem was near capacity and the files are distributed throughout, then you would expect the performance of the swap at the end and the swap at the beginning to start to converge.
If you think about it extending a filesystem is pretty easy: you just have to write in the filesystem control structures that you have more blocks available to store data than what originally planed. The problem of course is shrinking, since you have to relocate the blocks that goes beyond the new partition size.
Probably not an issue for a desktop since no one would want to use it under heavy swap all the time, but for a server no one pays much attention to... maybe.
SWAP on SSD or NVMe is still miles better than HDD, you can notice the difference when the swap is being used.
Turns out that the app grew huge over time and the machine would swap like crazy and would eventually slow to a crawl. The machine was already maxed out on RAM, so we added a service to restart the app twice a week. Finance said it took hours off their month-end work, they thought the app was just slow.
You can also monitor swapping activity in iotop. If need be, this can also be written on third party tools, the interfaces are exposed by the kernel after all.
Oh and you can use the modern PSI monitoring of the kernel to measure how much pressure a subsystem is experiencing, so you can restart services way before you'd even notice the swapping on other tools.
At least in my experience, it's pretty hard to actually gauge memory use, but swap use makes a reasonable gauge most of the time. There are certainly many use cases where the swap use ends up not being a useful gauge though.
Which means that CDs and DVDs are always read at the same speed, no matter where the laser / read head is.
Only those who really worked with hard drives noticed the speed increase at the inner ring.
This is only true for "slow" drives, CD drives faster than 12x typically use CAV and DVD drives >= 8x use CAV or Z-CLV (sometimes P-CAV).
Where's the intuitive start or end of the disk? I knew the answer was the tracks furthest from the center. Whether that was the beginning or end, I couldn't tell you.
If you divide the platter into concentric circles of equal width, you will notice there is more area available on the outer circles... for this reason the number of sectors per track are greater the further the track is from the centre of the platter. Yet the head will pass over the entire track in the same amount of time... i.e more data in the same time.
It makes sense that the logical volume would be arranged from the outer edge to take advantage of the speed as soon as possible.
an additional consequence is that for the same amount of data it takes lesser number of tracks thus making for faster/shorter seeks inside that data.
Is this a something like natural theory, or de facto standard? Maybe the start can be most inner edge on HDD in another planet?
Older HDDs did not make this optimization, they had a constant number of sectors per cylinder (track), and they also didn't come with HDD controllers, or came with more minimal controllers, and much like old floppy drives they exposed a lot of the physical layout and properties to the host system. This is why the CHS format is a historical part of OS and partitioning, modern HDDs only expose LBA.
Anyway, my point is that with the older drives that lacked ZBR there is no obvious "natural law" dictating that you shouldn't create a logical layout from the inner edge - and since they exposed CHS to the OS I wonder if it was possible to chose the layout direction.
[0] https://en.wikipedia.org/wiki/Zone_bit_recording
[edit]
I'm trying way too hard to entertain your idea, but now I thought it I gotta write it: the one way that ZBR could exist at the same time as it making sense to start the logical volume at the inner track (or more correctly making no difference), is if the angular velocity was not constant... modern HDDs have a constant angular velocity (constant RPM), and assuming the head's maximum read capability is the linear speed at the outer cylinder, then in theory the drive could spin faster when the head is closer to the center to provide the same linear speed, at which point the sequential performance is almost constant (cylinders sector density is segmented as the area passes a threshold, so there will be a subtle periodic difference).
This is how CDs work (ignoring the fact that they are a spiral), they have a constant linear speed rather than a constant angular speed and they start from the center, so the motor has to continuously change RPM to maintain it.
I suspect the reason this is not used is that HDDs spin much faster than CDs, and rather than single spiral track there are cylinders that the head has to move between, the head can move between cylinders much faster than the drive motor can accurately make the subtle changes required to achieve a constant linear speed.
The end result is probably very poor seek performance (i.e the seek performance would be limited by the motor rather than the head - which is very fast)... I don't know much about CDs but I suspect they also have poor seek performance, but by choice - that may have more to do with being one giant spiral track rather than motor performance.
Looks like the old timers knew what they were doing :)
This arrangement has advantages even in the age of VMs and SSDs. If I want to change to a larger disk or array (or resize the virtual disk), I can simply extend the last partition where the extra space is most likely to be needed. If swap was last, it would get in the way. On the other hand, if I needed more swap, I could just add a swapfile somewhere.
I haven't really had to worry about partitioning on any Linux machine I manage for over a decade thanks to LVM; I just create volumes based on what makes sense for the applications hosted on the servers.
One of the things that screws up HDD performance much worse than placement of files on disk is randomness in the usage pattern. The mechanical nature of a HDD means that when you read and write lots of small files in different sectors, the head spends more time moving around than reading or writing. Back when we used to defragment Windows filesystems, we doing a bunch of up-front disk optimization to organise files into continuous chunks so they could be read back quickly when needed.
The biggest problem I have seen with these situations is that you don't have direct control over the order of operations that the disk will be asked to perform. You think that because your file is written contiguously that it will be read that way. But depending on how busy the system is, that might not be the case. Where many processes are contending for disk access, and especially when the kernel is doing a lot of swapping to the same device, that head might be racing back and forth regardless of your file placement, and your disk performance goes straight into the toilet.
Typically I saw 30% to 100% performance improvements on ext4 by deleting and restoring database directories.
You can see disk fragmentation on linux with the filefrag and other commands.
So, is the an efficient way to leverage the speed improvement for other than swap -- like binary caching of executables of some form?
Sure. Like the site I mentioned in other comment, more or less "out tracks are faster". But this applies just to HDD drives. It's mostly useless for modern infrastructure, like SAN (even HDD based), all kinds of SSD and so on.
A curiosity - last time I saw a cool optimization of HDD usage was on "old gen" consoles, like ps4 and xbox one. Most games duplicated assets multiple times. Games took much more GB then needed, but the drive did not had to jump between many HDD tracks so much and it mattered for instance in big open world games.
I don't know the details of (Linux) caching, though. On my (32 GB) system, there are a few completely unused GB, it seems.
This could be an artifact of the particular kind of workload that the author used. Maybe it causes large numbers of adjacent blocks to be swapped in and out at the same time?
I originally ran the benchmark with 1GB RAM instead of the final 2GB, but the start-of-disk test did not finish in the 9 hours I let it run. With 0GB, I don't doubt that you'd see the expected 1,000,000x latency difference between disk and DRAM.
165s (2:45) — RAM only
451s (7:31) — NVMe SSD
Good argument for when the uninformed state that "NVMe might as well be RAM"(though, I guess this doesn’t give us any latency info, just throughput. I’d expect RAM latency to still be faster)
The spinning disk result is only 10x slower than RAM. But a spinning disk's throughput is 100-1000x less than current RAM, and for latency it's even worse.
Similarly, the other factors in the benchmark graph are way off their hardware factors.
This benchmark is measuring how one specific program (the Haskell Compiler compiling ShellCheck) scales with faster memory, and the answer is "not very well".
I originally tried running the test with only 1GB RAM, but killed the job after 9 hours of churning.
If your ops aren't latency sensitive, then NVMe might as well be RAM, if they are latency sensitive, then NVMe is not RAM (yet)
How well does NVMe scale to multiple devices, that is, how many GB/s can you practically get today out of a server packed with NVMe until you hit a bottleneck (e.g. running out of PCIe lanes)?
I created a ramdisk as follows:
~$ sudo mount -t tmpfs -o size=32g tmpfs ~/ramdisk/
~$ cp -r Downloads/linux-5.14-rc3 ramdisk/
~/ramdisk$ cp /boot/config-5.13.5-100.fc33.x86_64 linux-5.14-rc3/.config
~/ramdisk$ cd linux-5.14-rc3/
~/ramdisk/linux-5.14-rc3$ time make -j 32
My compiler invocation was: ~/ramdisk/linux-5.14-rc3$ time make -j 32
And got the following results Kernel: arch/x86/boot/bzImage is ready (#3)
real 6m2.575s
user 143m42.402s
sys 21m8.122s
When I compiled straight from the SSD I got a surprisingly similar number: Kernel: arch/x86/boot/bzImage is ready (#1)
real 6m23.194s
user 154m24.760s
sys 23m26.304s
I drew the conclusion that for compiling Linux, NVMe might as well be RAM, though if I did something wrong I'd be happy to hear about it!does anybody have hints or a link to some page explaining how to set up Linux so that it uses swap reaaally only if there is almost no free RAM available?
I have a few private servers & VMs, all having swap enabled, and all start using swap if I do a lot of I/O even if I have e.g. more than 20GBs free out of 36 being available. Usually swap is not being used just after having booted the server or VM, but after a few hours or days of doing reads & writes to disk the kernel will start writing stuff to swap - it's very little (few KBs being written every few seconds), but that accumulates and after a few days I end up having GBs of swap used.
On one hand I just personally hate seeing that happening, on the other hand some of my workloads are irregular so when the workload changes the swap is emptied (at least partially) and the whole thing starts over again.
So far I played with the values of "/proc/sys/vm/swappiness" (tried to set there 0, 1, 60, 100) and "/proc/sys/vm/vfs_cache_pressure" (tried to set there 50, 100, 200), but when doing a lot of I/O the OS always ended up using swap.
I would like to have swap available/enabled to cover potential extreme cases without having the programs crash (e.g. I might set memory limits of SW that might rarely run concurrently too high, or some database might suddenly allocate more than expected, etc...) => seeing that swap is being/was used would tell me that something is NOK in relation to the total RAM being used by my SW.. .
Nowadays I generally don't use any swap at all and find it annoying when distros/Windows create swap anyway. I mean if my 128GB+ or even 32GB of primary memory runs out, is it really going to help to swap 2GB to disk? And any larger swap than that is too slow to be usable.
I used to spread my swap out across all my disks on my system. When I had 2 disks. I put /boot, / and /var on one disk and /home on the other. When I had more disks, I moved /var onto its own disk, and had an extra drive that I symlinked into /home.
I put swap first on all the partitions. It's not like I did any benchmarking, there was just lore that swap should be close to the middle, followed by frequently accessed user data. At some point I got enough RAM that the swap wasn't really important, but I always provisioned it.
Now everything is SSD, and I feel like the whole idea of filesystem that you have to mount and keep consistent is kind old fashioned, but we have so much stuff built on the filesystem it will be with us a long time.
It is "a Linux kernel feature that provides a compressed write-back cache for swapped pages, as a form of virtual memory compression. Instead of moving memory pages to a swap device when they are to be swapped out, zswap performs their compression and then stores them into a memory pool dynamically allocated in the system RAM"
Also, firmware can "remap" some (bad) sectors into reserve area without kernel knowing.
There is also a difference between remapping a sector and reallocating a sector. Remapping simply means the sector was moved for operational reasons, reallocating means a sector has produced some read errors but did read fine.
A disk can operate fine even with 10s of thousands of reallocated sectors (by experience). The dangerous part is SMART reporting you pending and offline sectors, doubly if pending sectors does not go below offline sectors. That is data loss.
But simpy put; on modern disks the logical block address has no relation to the position of the head on the platter.
WD kind of tried that with device managed SMR devices and they show absolutely horrible re-silvering performance.
Without a relatively strong relations of linear write/read commands and their physical locations also being mostly linear, spinning rust performance is not on a usable level.
The issue with SMR is that because a write can have insane latencies, normal access gets problems.
CMR doesn't have those write latencies, so you won't face resilvering taking forever.
It also helps if you run a newer ZFS, which has sequential resilvers that do in fact run fine on an SMR disk.
I will also point out that wear leveling on a DMR disk tries to achieve maximum linear write/read performance by organizing commonly read sectors closer to eachother.
So if you place the swap near the rest of the files the hdd arm will not need to move so much.
Given that this was pretty much a clean Linux install I would assume that most files where at the start of the disk close to the best swap location.
1. choose the fastest HDD device
2. Use direct partition, no LVM
3. partition in middle of spinning platter’ busiest region of hard drive
4. single swap partition only
5. keep swap and hibernate storage separate.
6. encrypt swap (only downside)