Paragon submits 27k-line NTFS driver to Linux kernel
theregister.com
theregister.com
But unlike most patches that draw a similar response it's not a narrowly useful patch that mostly serves the submitter. Proper NTFS support benefits a large proportion of Linux users (the jab from the article that there are more advanced file systems out there seems out of place; there are no signs that windows is about to switch its default FS to something else).
Additionally this code has been used in production for years now (e.g. my 2015 router runs the closed source version of this driver in order to support NTFS formatted external drives) so most likely a lot of quality issues have already been found and addressed.
So I feel it's a bit unreasonable to respond with so much negativity to this contribution.
Split your diff! and Fix your makefile! have to be one of the most benign and common pieces of diff feedback i've seen. I feel that you could make a media story about any submission to the Linux kernel based on there being comments in the review process.
> So how exactly do you expect someone to review this monstrosity ?
But it's not like splitting the feature into 100 patches of 250 lines each would make it any quicker to review. Or merging code that was known not to work, as it was only a fraction of what was needed for the functionality.
That would also be rejected, because the kernel maintainers aren't idiots and their standards aren't the stupid arbitrary rules you construe them to be. They generally want big changes to be broken up into logical, sensible chunks that each leave the tree in a usable state, so that git-bisect still works.
I guess one could start by merging a skeleton of the filesystem which supports mount/unmount but then returns an IO error on every operation? And then a patch to add directory traversal (you can view the files but not their contents), and then a patch to add file reading, and then a patch to add file writing, and then a patch to add mkdir/rmdir, and then a patch to add rename/delete of regular files.
Breaking down an existing filesystem into a sequence of patches like that, no doubt it is doable, but it is going to be a lot of work.
https://www.zdnet.com/article/linux-developer-who-took-on-li...
The entire dispute seems to be that minor question of style, nothing substantive. I don't think anyone's specially unhappy on either side. The controversy seems manufactured, perhaps by a reporter who noticed the gruff language but lacked the technical knowledge to understand what's actually going on.
Most people developing free software (probably including both the submitters and recipients of this patch) could make a lot more money elsewhere, but have chosen to instead to do work with considerable public benefit. That's thankless enough already without some reporter inventing drama for clicks.
git is capable of breaking down a large diff into manageable pieces (e.g., limiting a diff to a single file), but reviewing code in a mailing list means replying to the message that contains a patch and replying inline to certain parts to comment on it.
As for higher level software that could break down a large commit, what specifically do you have in mind? I can't think of any feature that other review tools like Git??b, gerrit, reviewboard, phabricator, etc. that would make something like this easy to review.
27 kLOC will be a big project to review no matter what, but I'd probably rather take them in a single commit--the files presumably depend on each other, and there's probably no order in which the files could be reviewed in isolation without reference to files not yet reviewed. (Obviously we try for hierarchical structure that would make that possible, but not usually with perfect success.)
That's a matter of personal preference, though, and people who want a project to merge their contributions should adhere to the maintainer's preferences. In any case, it seems Paragon intends to do exactly that. I doubt Paragon expected their reward for their contribution would be an article read by thousands of people that called it "half-baked" over this minor point, and I can't imagine such publicity encourages others to make similar contributions in future.
Github does allow you to filter the diff down to the commit or jump to a particular file within the diff. Commenting on a line in the diff isn't really any different than positioning one's comment inline below the relevant line(s) of code in an email reply. I know that in Github, it's also possible to comment directly on a commit (though those comments are not displayed with any context in the general PR view), unlike an email reply to a particular patch series.
Depending on one's email client, it's certainly possible to search for things like /^diff/ or /^@@/ to jump from file to file or hunk to hunk within the compose window.
> perhaps follow references into the full code faster than you could flipping between your mail client and your editor (and save you the effort of applying the patch to a local tree for that review).
For some, the email client doubles as an editor (i.e., gnus). And, at least in my experience, it's far faster to navigate code in an editor compared to the web interface that Git??b provides.
> 27 kLOC will be a big project to review no matter what, but I'd probably rather take them in a single commit--the files presumably depend on each other
While that's true, the dependency can be preserved when merging the branch of the series of commits in the mainline repository. Plus, many may find it easier to review declarations, definitions, and calls in that order.
I think most free software developers are normal corporate employees. I work on tons of free software as my job, like most of my peers, but that’s normal in the industry. I don’t consider myself a free software developer.
I meant independent volunteers or people working for free-software-focused companies (which I believe usually offer well below FAANG-level compensation, especially at the high end, though still enough to live quite well). Excluding hardware vendors porting Linux to their own products, I believe the core kernel developers tend to fall in to that last category. I have no specific knowledge of their individual compensation, but the technical leads responsible for closed-source projects of similar scope make incredible amounts of money.
Others noted it needs to pass the existing test suite and that it is close.
1) The Linux kernel already has an in-kernel read-only NTFS driver. What should be done about it? (There are a number of reasonable options here, including just getting rid of it and replacing it with Paragon's, but that requires at least some buy-in from the maintainers of the existing driver.)
2) The patch didn't actually build, which was a one-line Makefile fix, but raised some concern about how it wax tested/how the patch was generated.
Imagine if you could have a backup drive (with reasonable modern data protections) that you could just plug into different systems and save all your files to. Isn't it odd that such a simple thing isn't possible? I guess network attached storage has gotten pretty accessible at this point so there's no need for it?
That is fine and appropriate for a drive that will be connected to the system for the foreseeable future.
That kind of compatibility concern makes me squeamish about using ZFS for a drive that I want to share between different systems. If it's easy to make it incompatible between two releases for the same system, that smells like a waiting nightmare trying to keep it compatible between Linux, FreeBSD and Windows.
Since I didn’t have money for any more hard drives at the time, I couldn’t transfer the data to anything else. So then when I wanted to access that data I’d do so via a FreeBSD VM running in VirtualBox. The performance was... not great.
I took the data that I needed the most, and for the rest of the data I let it sit at rest.
This week I wanted to use the drive again, and in the end because I was doing general cleanup, I decided to install FreeBSD on my desktop temporarily.
I actually love FreeBSD but the reason that I prefer to have my desktop running Linux is in big part because I want software on the computer to be able to take advantage of CUDA with the GTX 1060 6GB graphics card that I have in it, and unfortunately only the Linux driver by Nvidia has CUDA, the FreeBSD driver by Nvidia does not.
I was actually looking at installing VMWare vSphere on the computer instead, so that I could easily jump between running Linux and running FreeBSD with what I understand will probably be good performance compared to VirtualBox at least. But the NIC in my machine is not supported and vSphere would not install. I found some old drivers, messed around with VMWare tooling which required PowerShell, and which turned out not to work with the open source version of PowerShell on any other operating system than Windows. So then I downloaded a VM image of Win 10 from Microsoft [0], and used that to try and make a vSphere installer with drivers for my NIC. No luck at first attempt unfortunately. A decade ago I probably would have kept trying to make that work, but at this point in my life I said ok fine fuck it. I ordered an Intel I350 NIC online second-hand for about $40 shipping included, and the guy I bought it from sent it the next day. It is expected to arrive tomorrow. Meanwhile, I installed FreeBSD on the desktop. When the NIC arrives I will do some benchmarking of vSphere to decide whether to use vSphere on the desktop or to stick to either FreeBSD for a while on that machine or to put it back to just Linux again.
Anyways, that’s a whole lot more about my life and the stuff that I spend my spare time on than anyone would probably care to know :p but the point that I was getting to is that, with OpenZFS 2.0 I will be able to use ZFS native encryption instead of GELI and I will be able to read and write to said HDD from both FreeBSD and Linux.
I still need to scrape together money for another drive first before I can switch from GELI + ZFS to ZFS with native encryption though XD
Oh, and one more thing, with the external drive I was having a lot of instability with the USB 3.0 connection on FreeBSD, leading to a bit of pain with transferring data because the drive would disconnect now and then and I’d have to start over. But yesterday I decided to shuck the drive – that is, to remove the enclosure and to connect the drive with SATA like you would any other regular internal drive. It worked out excellently, the WD Essentials enclosure was easier to pry open than I had feared, and a video on YouTube showed me how to do it [1]. As prying tools I used a couple of plastic rulers. As a bonus, it also looks like I/O performance is better with the direct SATA connection than what I was getting with the USB 3.0 connection.
Speaking of that, some people have reported finding that the drives in their WD Essentials external drives were WD Red HDDs. I didn’t have the same luck with mine; mine was WD Blue. But idk if WD Red is even common with the capacity that mine has anyways. Mine is “only” 5TB and I think the people that have been talking about finding WD Red drives in theirs has bought 8TB models often. Idk. The main thing for me anyways is just to have my data and someplace to store it ^^
[0]: https://developer.microsoft.com/en-us/microsoft-edge/tools/v...
While FAT/exFAT leave the possibility of a variety of different types of filesystem inconsistency, these seem to be fairly rare in actual practice, probably in good part due to Windows disabling writecaching on devices it thinks are portable. This kind of special handling of external devices is sort of distasteful, and leads to some real downsides on Windows (e.g. LDM handling USB devices weird), but using a newer file system doesn't really eliminate that problem - NTFS and Ext* external devices require special handling on mounting to avoid the problems that come from file permissions traveling from machine to machine, for example.
(Now I use NTFS and Tuxera’s commercial Mac driver, because I don’t know how else to have a cross-platform filesystem without a stupidly-low file-size limit.)
UDF is natively supported on all major OS.
I haven't used it in macOS but on paper macOS seems to be having even better support for it than Linux so you can give it a shot.
Lastly, you can use https://github.com/JElchison/format-udf for creating most compatible filesystem across different devices.
I'm using it on an external HDD for copying/watching video files between Linux and Windows boxes and haven't had any problems yet.
On the other hand 7-8 years ago I tried to use an UDF partition for sharing a common Thunderbird profile between Windows and Linux and had done strange errors on Windows side after a while. I didn't dig further so it may have been a non-udf os or tooling issue, or it may have been solved in the meantime.
What's distasteful is that on my linux machine with a lot of ram, when I copy a multiple gigabyte file to a USB key it "completes" the copy almost immediately, when actually all it has done is copy the file to a ram buffer. Then when I try to disconnect the drive, it will hang for ages while it actually finishes the write. IMO windows does it better here (although I never realised what exactly they did, nice to know.
What I end up doing is using a "watch" on some command I can't remember that shows overall dirty pages.
On really new kernels, it might work well to use io_uring to issue linked chains of read -> write -> fdatasync operations for everything the user wants to copy, and base the GUI's progress bar on the completion of those linked IO units. That will probably ensure the kernel has enough work enqueued to issue optimally large and aligned IOs to the underlying devices. (Also, any file management GUI really needs to be doing async IO to begin with, or at least on a separate thread. So adopting io_uring shouldn't be as big an issue as it would be for many other kinds of applications.)
But it hasn't and won't work for obvious reasons.
With journalling, I don't have to know or care about any of this: I restart and chances are it's all back to normal. This is desirable.
You can store your backups on stone tablets, with a machine that carves rock to write 1s and 0s and a conveyor belt that feeds new tablets. That is perfectly feasible. It is also not desirable.
Many reasons, the most important being patents and insistence by vendors to keep stuff proprietary (exfat, ntfs, HFS, ZFS).
Also, there are fundamental differences on the OS level:
- how users are handled: Unices generally use numerical UID/GID while Windows uses GUIDs
- Windows has/had a 255-char maximum length of the total path of a file
- Unices have the octal mode set for permissions, old Windows had for bits (write protected, archived, system, hidden) and that's it
- Windows, Linux and OS X have fundamentally different capabilities for ACLs - a cross platform filesystem needs to handle all of them and that in a sane way
- don't get me started on the mess of advanced fs features (sparse files, transparent compression, immutability of files, transparent encryption, checksums, snapshots, journals)
- Exotic stuff like OS X and NTFS additional metadata streams that doesn't even have representation ín Linux or most/all BSDs
And finally, embedded devices and bootloaders. Give me a can of beer and a couple weeks and I'll probably be able to hack together a FAT implementation from scratch. Everything else? No f...in' way. Stuff like journals is too advanced to handle in small embedded devices. The list goes on and on.
The actual path of file in Windows can be practically unlimited, but either requires using special network notation that can exceed 256 characters or relative addresses. Recent versions of Windows includes a setting that removes the limitation in the APIs because of development issues like node nested packages.
Which is why we have abstraction layers like Samba (etc) on top of networked drives. They're descendants of vintage cross-OS utilities like PIP which provide a minimal interface that supports folder trees and basic file operations.
But a lot of OS-specific options remain OS-specific, and there's literally no way to design a globally compatible file system that implements them all.
This isn't to say a common standard is impossible, but defining its features would be a huge battle. And including next-gen options - from more effective security and permissions, to content database support, to some form of literally global file ID system, to smart versioning - would be even more of a challenge.
[0] https://en.wikipedia.org/wiki/Universal_Disk_Format [1] https://github.com/JElchison/format-udf
Same thing applies to NTFS, but I've seen that more often than not Windows creates files on removable drives with extremely open permissions, and in my experience NTFS-3G just straight ignores NTFS ACLs on the drives it mounts, so more often than not it JustWorks™ in common use cases.
I think a journaled extFAT-like filesystem would be perfect for this task, but given how hard it was for exFAT to even start to displace FAT32, even if it actually existed I wouldn't expect it to succeed any time soon.
It doesn't build, but the person who pointed it out also supplied a diff to make it happen.
It also fails a few tests, but Paragon are more than happy to see if they can make it a bit more compliant.
UBSan finds a few potential bugs, but again, Paragon are more than happy to fix the problems.
There's some style guide suggestions, which Paragon seem to immediately take on board:
> The patch will be splitted in v2 file-wise. Wasn't clear initially which way will be more convenient to review.
[0] https://lore.kernel.org/lkml/2911ac5cd20b46e397be506268718d7...
It also seems a bit ignorant to call the submission "half-baked" just because it needed to be worked on and because someone pointed that out (also) in a slightly irreverent way.
All of those points you bring up about the feedback they got seem like just business as usual for a larger merge of code to a carefully developed FOSS project.
That said, I think the tension is more imagined than real, as the LKML doesn't really seem to have responded any more negatively than they do to most other patches.
example: https://arstechnica.com/information-technology/2020/03/the-e...
Its gonna take a while till this driver is mainline in Linux kernel, and till that Linux kernel is included in distributions (especially LTS).
I haven't read The Register article, but in the past it has come to my attention they dramatize their articles, and I don't want to read such media.
Earlier discussion: https://news.ycombinator.com/item?id=24170001
Don't look a gift horse in the mouth, especially when you need one.
I think their point is ‘man, how are we going to divide this up amongst the maintainers? Who gets to check which function or call?’
Paragons response (more or less will fix in v2) https://lore.kernel.org/linux-fsdevel/a8fa5b2b31b349f2858306...
We’ll all try really hard to ignore our individual biology to satisfy the most sensitive sensibilities.
Cause that expectation is not manipulative at all.
You needn't use your real name, of course, but for HN to be a community, users need some identity for other users to relate to. Otherwise we may as well have no usernames and no community, and that would be a different kind of forum. https://hn.algolia.com/?sort=byDate&dateRange=all&type=comme...
After all, we all lived long enough without their software, we can live without it a bit longer.
Yes it's a gift, but it's also code that has to be maintained and updated along with the kernel, so it's not really 100% a gift. If you then consider that they might keep on selling a proprietary version of the code (which - don't get me wrong - is 100% legit and fair) they might also get basically free labour: they could rebase onto the latest public gpl version, they might get notes of various issues and bugs...
Quite literally, it's free labour.
(As a mental exercise sometime, go to a pet store and figure out how much a “$12 hamster” costs once you get everything you need to set up and maintain a habitat.)
Puppies can bring a lot of joy, but they certainly bring obligations.
Really spot on.
I have a mid-junior level co-worker who submits PRs that are excessively large. As best I've been able to determine, he's not especially good at managing the dependencies in his code, and he doesn't want to submit a broken PR, so his default is to wait until he gets everything written instead of breaking it down into smaller pieces.
Sad as it may sound, free code doesn't come for free.
Hooray, free puppies. Now you just have to care for them for the next 10 years.
There's no drama happening here. Paragon guys are trying to give it "properly", and linux guys want it "properly", and the only thing happening is defining "proper" in this context
Additionally, the driver is mature, as is the FS.
And who needs it? There are alternatives to read or write the occasional file (e.g. if you happen to have forgotten the Admin's password) of a NT box, but I can't say that I missed the ability to create new files in a NTFS. And why now? Surely those who actually needed that functionality, needed it years ago and meanwhile found some other solution. Perhaps those who actually need it still volunteer to ready that driver for inclusion into the kernel or sponsor someone who can?
Not quite:
> CONFIG_NTFS_RW:
> This enables the partial, but safe, write support in the NTFS driver.
> The only supported operation is overwriting existing files, without changing the file length. No file or directory creation, deletion or renaming is possible. Note only non-resident files can be written to you may find that some very small files (<500 bytes or so) cannot be written to.
Interesting you choose that metaphor. I can think of a particular gift horse where you would have seen Greek soldiers if you have looked into its mouth. :)
It's very debatable that Linux "needs" more support for a proprietary 25-year-old filesystem that is, in many ways, obsolete.
In what kind of context? The Linux kernel?
Filesystem code is pretty tricky to begin with, and prone to very subtle bugs with very not-subtle consequences. And this isn't greenfield development of a new filesystem, but an implementation that needs to remain highly compatible with Microsoft's version. This FS driver has to be maintained to track changes to two operating systems. So this 27kloc can reasonably be expected to encompass a lot more complexity than your average 27kloc, and it requires a lot more review effort than something like 27kloc of GPU driver register definitions.
Not the Linux kernel, but a large embedded system.
> Filesystem code is pretty tricky to begin with, and prone
> to very subtle bugs with very not-subtle consequences.
This code has been running in the wild for quite a while now, it has had a trial by fire. And there's no way around testing, subtle ext4 bugs still crop up despite the maturity of the filesystem.
> And this isn't greenfield development of a new filesystem,
> but an implementation that needs to remain highly
> compatible with Microsoft's version.
No amount of code review will stop Microsoft from adapting their version. Also, I doubt Microsoft themselves will change too much about the filesystem given the compatibility they themselves have to maintain with cold storage NTFS drives.
> This FS driver has to be maintained to track changes to
> two operating systems.
You make it sound as if Microsoft have a hand in any of this. Also, have you seen the state of the current NTFS driver? It's a bit flakey (no disrespect to the maintainers).
Way to miss the point. Code review for the kernel isn't just about verifying that the code currently works. It's also about making sure the code is maintainable. Microsoft is relevant here because their actions will increase the maintenance burden of any Linux NTFS driver. Kernel developers rightly need to be concerned about how difficult it will be to extend the NTFS driver to handle new NTFS features that Microsoft introduces.
The point wasn't so clear, but I see what you're saying now. Maintainability is normal code review though.
> It's also about making sure the code is maintainable. [..]
> Kernel developers rightly need to be concerned about how
> difficult it will be to extend the NTFS driver to handle new
> NTFS features that Microsoft introduces.
Maintainability is one thing, extensibility is another. Preparing your code to implement some changes completely outside of your control seems like a waste of time and something that might bite you later on.
This is 27 kLOC.
It's going to take months to throughly assess the quality of this.
Hope I’m wrong but I think they’ll be fighting initial bad impression for a while despite good intention
Is this remotly true?
Will my grandmother not be using NTFS in 10 years?
So yes she will I guess.
Then it's vital Linux gets NTFS working well.
My partner is not going to use Linux things if every time they try and transfer the 8k 3D holographic photos of our CRISPR'ed dog learning to spell to my grandmother it doesn't work on her Holovision.
True story, last month I just lost about 1 in 10 of my media files on my Linux Share to NTFS issues. So now I run the box on Windows.
Were you seriously running a Linux-based NAS with NTFS as the underlying filesystem, or did you mean something very different? I can't imagine why anyone would ever think their choice of disk filesystem on a server—hiding behind a network filesystem—should be influenced by what disk filesystems are supported by client devices.
My setup was -
Windows and Ubuntu dual boot gaming PC with a NTFS 4G external hard disk, Windows network share. Default boot was to Ubuntu.
Torrents running off wireless laptop to the network share.
Amazon Fire stick with Kodi wireless running to network share.
Accepting a new driver which may in fact not work, breaks in ambiguous ways, interacts with other components poorly, or otherwise generates headaches for the kernel could be more trouble than it's worth. Ultimately this driver will eventually need to be modified by others, and ensuing it doesn't get a reputation as a nightmare right out of the gate is also worthwhile.
That's literally why having to review 27K lines of code is a pain. You're not seriously going to trust, sight unseen, that there's nothing in the code that slips in a backdoor or catastrophic data-loss bug, are you?
But more important than that is the maintenance burden. NTFS will be around for a long time, but it's also a moving target because Microsoft hasn't replaced it yet. Kernel developers have to keep in mind how this code will look in a few decades, after all the original developers are retired. If it's written in a very different style from other Linux filesystems, will there be anyone left who both knows enough about the workings of the Linux IO stack 20 years hence, and understands Paragon's code conventions?