HNHacker News
TopNewBestAskShowJobs

prakashsurya

125 karma · joined February 24, 2014

OpenZFS Developer at Delphix
submissionscomments
prakashsurya··on OpenZFS: Reducing ARC Lock Contention
Well, if the exact same semantics were required (i.e. always evict the oldest buffer), then it's really O(M*n) where M is the number of bytes to evict, and n being the number of sublists; each sublist would have to be checked for each buffer that evicted. But, I agree, if that was the only concern then the extra comparisons _probably_ wouldn't be enough to be concerned about.

The bigger issue is how to lock the sublists. It seems heavy handed to lock all sublists, find the oldest buffer, evict it, and then unlock all sublists. But if each is locked, checked, then unlocked; each sublists' oldest buffer can change while the eviction thread is still iterating the sublists to find the oldest buffer.

So, while the current solution is not very elegant, it is simple, requires minimal locking, and appears to work well enough.

prakashsurya··on OpenZFS: Reducing ARC Lock Contention
Currently, the number of sublists created is equal to the number of cores on the system.

While it'd be fairly trivial to extend the code to make it easy for an admin to define the number of sublists, I don't have any reason to believe that'd be a useful knob to export.

prakashsurya··on ZFS vs. Hammer
Could not have said it better myself. It's disappointing to see a post with such unfounded nonsense rated so high on HN.
prakashsurya··on Is ZFS a suitable replacement for other Linux filesystems?
yup, that's why I added "by default". a user won't get that functionality without explicitly asking for it.
prakashsurya··on Is ZFS a suitable replacement for other Linux filesystems?
Error detection can always be used, but error correction may or may not be available (it depends on the type of block). Metadata blocks are redundant even on a single drive pool; so if you just have a partial failure (e.g. overwrite a metadata block) it might be able to correct the block using another redundant copy on the same drive. Data blocks will require a redundant pool configuration, though, as these are not store redundantly by default (e.g. multiple drives in a raidz or mirror).
prakashsurya··on Is ZFS a suitable replacement for other Linux filesystems?
No problem. It's really a shame to have such a high bar for using dedup. It either fits a given workload extremely well, or can be extremely detrimental.

There's been talk in the developer community about ways to address the usability of dedup, but so far nothing has gone further than small prototypes.

prakashsurya··on Is ZFS a suitable replacement for other Linux filesystems?
There is work in progress to add the ability to remove a device from a pool. So, hopefully in the not too distant future, that work will land in illumos and migrate to all the downstream implementations (e.g. linux, mac, and freebsd).
prakashsurya··on Is ZFS a suitable replacement for other Linux filesystems?
No, that's flat out not true.

I've seen that metric thrown around when talking about the "dedup" feature of ZFS, but honestly, don't use dedup unless you know what you're doing. It's way to easy for things to go wrong otherwise.

prakashsurya··on Is ZFS a suitable replacement for other Linux filesystems?
Yes, I run ZFS on Linux on my single drive laptop currently.

ZFS doesn't _need_ multiple drives to work well, but it is capable of using them if they're available. One can also run a HW raid solution underneath ZFS, where the RAID engine just presents ZFS with a small number of LUNs. It's all up to the admin; and what sort of performance, redundancy, and maintenance guarantees are required.

prakashsurya··on The State of ZFS on Linux
FWIW, if there's anybody interested in learning about the ZFS code base, we'd love help porting patches from ZoL into Illumos and vice versa. That's a good way to get a new developer integrated with the code and process surrounding each platform.
prakashsurya··on The State of ZFS on Linux
I can't recall the exact detail from memory, but I believe #1 has to do with the fact that zfs creates/clones/snapshots/etc are done in "syncing" context. Thus, each command has to wait for a full pool sync to complete, limiting the rate at which these can be done.

This is a known problem, and likely to be fixed in the not too distant future.

prakashsurya··on Show HN: Sysdig, a tool for Linux system exploration
Is dtrace available on Linux? I know there's been work towards that goal, but I haven't payed much attention to it recently.