The NOVA filesystem
lwn.net
lwn.net
due to the per-CPU inode table
structure, it is impossible to
move a NOVA filesystem from one
system to another if the two
machines do not have the same
number of CPUs.
seems like a dealbreaker. What use is a FS if it isn't portable? At least it sounds like they're very aware of this issue.I'd be very interested to see what they end up doing to make this behave better prior to upstream consideration. I wonder if a linked list journal of inode table changes (with space drawn from per-CPU freelists) would be safe/fast enough. That could be periodically coalesced and remapped to the NOVA device's non-CPU aware inode table.
Changing a single CPU-related flag in the BIOS should under no circumstances render a filesystem unmountable.
You don't have external hard drives or flash drives...?
The entire complaint itself is the fact that this file system isn't able to do what people expect from most file system images (read: pretty much every one, except this one). Citing the fact that this particular one fails that pattern isn't a justification, it's circular logic.
Wouldn't a Server hardware cluster potentially prove problematic?
> The origin of the CPU-count dependence is that NOVA divides PMEM into per-CPU allocation regions. We use the current CPU ID as a hint about which region to use and avoid contention on the locks that protect it.
> So moving from a smaller number of CPUs to a larger number of CPUs just means more contention for the locks. Moving from a larger number to a smaller number is no problem at all. So, our current plan is to set the CPU count very high (like 256) when the file system is created.
This is required, I assume (after 5 min scan so can be wrong), that in order to speed things up, information about the CPUs are embedded in the filesystem.
So is this not a classic "hardcode for speed" vs "generic but slow" trade off?
I actually wrote a comment in response to loeg elsewhere in the thread and lost it by refreshing (and then lost interest, hah!) on the subject of conversion. I think there are some interesting problems around introducing a new FS and having to build "migrations" into it from the start. Migration feels like an expensive (in time/effort) word and I'd be concerned about issues coming up when you need to rewrite chunks of your NOVA FS because the per-CPU region has to become larger and everything else needs to shift along. It wouldn't need to be a full rewrite since only some data would need to be moved but it still seems clunky.
As mentioned in another comment, it seems like the solution being put forward is to oversubscribed the per-CPU data region for 256 CPUs but I wonder how long until that turns out to be a bad assumption. Who knows, maybe it doesn't matter. Maybe it's just out of scope for NOVA. One FS doesn't need to rule the world after all!
"The origin of the CPU-count dependence is that NOVA divides PMEM into per-CPU allocation regions. We use the current CPU ID as a hint about which region to use and avoid contention on the locks that protect it. So moving from a smaller number of CPUs to a larger number of CPUs just means more contention for the locks. Moving from a larger number to a smaller number is no problem at all. So, our current plan is to set the CPU count very high (like 256) when the file system is created." - comment under the post by one of the designers of the fs
No one thinks twice about having to "do something" with their RAM disk to upgrade the CPU.
It's time we realized the new class of nonvolatile memory hardware is a new entry in the storage hierarchy. One shouldn't think of it as RAM, but it's much, much faster than traditional disk (or even SSDs) and should be thought of something different.
It's a mistake to put the same restrictions on the potential performance of this just to make it fit our preconceived notions of storage.
What type of usage am I missing where losing/having to do an offline rebuild of your data would be acceptable during a hardware upgrade? NOVA seems to be aiming to address the recognised issue by overprovisioning (which I disagree with but such is life). I don't really understand your argument.
Imagine a hardware architecture which considers NV-RAM as a very large and safe cache, and software which understands that.
This would be great for things like a (huge) RDMS which traditionally requires tables and especially indexes to be in RAM. Suddenly you can get adequate performance on multi-terrabyte indexes, and providing the software understands the storage hierarchy there is no explicit rebuild needed - the data is also stored on spinning rust and on a WAL maybe on a SSD, and will be reloaded into NV-RAM on read.
Interop is vital, though. I'm less concerned with being able to pull and seat a drive between machines[1] than sensible replication/backup, and a 'send/receive' type-thing shouldn't be too big of a deal.
[1] Thumb drives are an obvious exception.