WizTree isn't open-source like WinDirStat but "free as in beer" with optional donations.
There's also a fork of WinDirStat patched to read the MFT but I don't know anyone who's tried it: https://github.com/ariccio/altWinDirStat
WizTree isn't open-source like WinDirStat but "free as in beer" with optional donations.
There's also a fork of WinDirStat patched to read the MFT but I don't know anyone who's tried it: https://github.com/ariccio/altWinDirStat
One disadvantage is that you can't read the MFT of network shares or device emulators presenting "virtual drive letters" to the OS.
The typical (and slower) Win32 API functions FindFirstFile()/FindNextFile() used to iterate through the files structure work at a higher level of abstraction so they work on more targets that don't have an NTFS MFT. Indeed, if you point WizTree to a SMB network share, it will be a lot slower because it can't directly read the MFT.
It's conceivable that Microsoft developers could have programmed Windows Explorer differently to have an optimized code path of reading MFT for local disks and then fall back to slower FindFirstFile()/FindNextFile() for non-MFT disks. Maybe that adds too much complexity and weird bugs. I notice that most of the 3rd-party "Win Explorer replacement" utilities also don't read MFT.
Surely this would have been worth doing, even if it meant flushing out bugs elsewhere.
Would that mean that there's no way to "scope" the MFTs?
Edit: That also makes sense, since if I got it right they aren't necessarily supposed to be consumed by userspace programs?
I guess that's why those tools always ask for admin access and basically all perms to the FS.
It's a bit sad that the user gets exposed to a much slower search and FS experience even if the system underneath has the potential to be as fast as it gets. And I don't think ReFS is intended to replace NTFS (not that it's necessarily more performant anyways)
Linux has device mappers (dm-crypt, dm-raid and friends). But those sit below the file system, emulating a device. Window's file system filter drivers sit above the file system, intercepting API calls to and from the file system. That's super useful if you want to check file contents on access, track where files are going, keep an audit log of who accessed a file, transparently encrypt single files instead of whole volumes, etc. But you pay the price for all that flexibility in performance.
https://man7.org/linux/man-pages/man7/fanotify.7.html
https://lwn.net/Articles/339399/
It even lets you block the access until the scan/decision is made.
Or if you just want to generally make the filesystem so slow that everyone has to invent their own pack files just to avoid file system api calls as much as possible.
Maybe stating the obvious, but if the security can be violated that easily, it's not very secure.
Even if you want more control over your system, I still think technically capable people would be better served by having a separate administrator account from your normal day-to-day account which you have to explicitly log into (so no UAC prompts, you need to go onto that other account and then you get the UAC prompt). Unfortunately, I think most Desktop OSes are still too unusable with this sort of workflow due to how much software insists on admin for installation.
I really like your idea of logging in separately, such that is isn't something you're going to do cavalierly. That seems like a great compromise to me! I fully agree that we way overuse admin and really don't need it for the majority of things.
The main use case for filter drivers is antivirus, and that is primarily about file contents not file metadata - so if MFT access bypassed filter drivers, that might not be a major issue. I think most non-antivirus use cases are also primarily about data not metadata.
If necessary, one could even devise a design in which MFT access is combined with filter drivers - MFT scanning to find matching files, then for each matched file access its metadata via standard APIs (to ensure filter drivers are invoked) before returning to client. That would be slower than a pure MFT scan but still faster than a scan done purely with standard APIs. A registry key could turn this on/off so sites can decide for themselves where to place the performance versus security tradeoff
> and would also ignore any permissions (ACLs) on who can see those files
They could expose an API which enables MFT scanning with some degree of ACL checking added.
If you do the ACL check as late as possible in processing the query, it would give much better performance than standard APIs that evaluate ACLs on every access. For example, suppose I want to scan a volume for all files with the extension ‘*.exe’. The API would only have to do an ACL check on each matching entry, not on every entry it considers.
There also might be reasonable situations in which ACL checking could be bypassed. For example, if I am requesting a search for files of which I am the owner, just assume the owner should have the right to read the file’s metadata. Or, if I have read permission on a directory, assume I am allowed aggregate information on the count and total size of files in that directory and its recursive subdirectories. These “bypasses” could be controlled by system settings (registry entries / group policy), so customers with higher security needs could disable them at the cost of reduced performance.
Rather than putting this in the OS kernel, it could be a privileged system service which exports an API over LPC/COM/etc. Actually with that design it isn’t even necessary to wait for Microsoft to implement this, it could always be implemented as an open source project, if someone felt sufficiently motivated to do so. (Or even as a proprietary product, although I suspect that would limit its adoption, and the risk is if it takes off, Microsoft would just implement the same thing as a standard part of Windows.)
The only workaround (used by Everything by VoidTools) is to install a service which would run with a needed rights and communicate with it in the GUI.
You do similar things even with more modern stacks - assign a permission to an application and grant permissions to the application to the user.
The only real concern is that Windows NT permissions are not as granular as they could be.
For objects, Windows NT permissions are ridiculously granular; e.g. GENERIC_WRITE can be mapped to a half-dozen separately settable type-specific flags, depending on the object type (file, named pipe, etc.). It’s too granular for even an administrator to make sense of, arguably, and the documentation is somewhere between bad and nonexistent. (The UI varies from decent, like the ACL editor you can access from e.g. Explorer, to “you can’t make this shit up”, like SDDL[1].)
For subjects, the situation is not good, like on every other conventional OS. You could deal with that by introducing a “user” for each app, as on Android. But I’m not aware of any attempts to do that (that would expose this mechanism in a user-visible way).
(Then there’s the UWP sandbox, which as far as I tell is build with complete disregard of the fundamental concepts above. I don’t think it’s worth taking seriously at this time.)
[1] https://learn.microsoft.com/en-us/windows/win32/secauthz/sec...
I’ve had to work with SDDL before to setup granular permissions for WMI monitoring on a whole lot of computers and my god, did it make me love the Cloud and Linux. I can’t emphasize enough how unintuitive setting these permissions is creates systemic over privileging.
It does not say anything like that in FAQ and i don't remember it being fast.
https://github.com/seanofw/spacemonger1/blob/6a41c012534b170...
One possible reason is that it isn't a published part of the filesystem's external interface, and the format is not guaranteed to be static between versions or even point releases (though in reality, while the behaviours may be officially undefined that are unlikely to change significantly).
Also, it requires admin elevation to access. Anything running elevated is a potential security concern as it can access much else too.
> Why doesn't Microsoft do it in file explorer
Not sure, but it could be because that would be seen as an unfair advantage so to avoid anti-trust allegations they would have to publish the format and make stability guarantees for it, so others could use it as easily/safely. That, and the reasons above & below too.
> and why wouldn't every tool use it instead of walking through the file system?
Largely because walking the filesystem works for all filesystems, local and remote, so you cover everything with one tree walk implementation. Implementing a tree-walk over the MFT data where available is extra work to implement and support for one filesystem, and not many care enough, or are not aware of the potential speed benefit at all, for it to be a huge selling point such that all toolmakers feel compelled to bother.
I am not going to pull every document, but the MFT structure is documented and published. I am uncertain what you mean by "external interface".
"About 9,810 results (0.04 sec)"
https://scholar.google.com/scholar?hl=en&as_sdt=0%2C11&q=mft...
You can find a lot of articles talking about SQL Server's DBCC IND and DBCC PAGE, but that isn't official documentation – they are essentially internal functions and not supported and could change or go away entirely despite having been around for many versions, as they have in Azure). Similarly there articles talking about sys.dm_db_database_page_allocations which sort-of does the job of DBCC IND, but again this is not officially documented & supported.
> I am uncertain what you mean by "external interface".
I meant the published interface. Maybe "supported API" would have been a better phrase to use?
Though as pointed out below, there is at least some official documentation on the MFT structure.
It doesn't even bloody support network drives so there's no such reason.
Do that and it’s alarmingly fast and responsive except for the minute or two right after launch.
I believe version 3.38 was the last version that is completely "free as in beer" with optional donations.
Which is enough for me to not use it because WinDirStat still only takes a minute. Cool software though.
Fact is, WDS community must be kind of abandoned, or else it would be doing the same trick. It's SO much faster that it becomes a genuine quality of life improvement. I need it, and don't mind using a non free tool until the OSS solution has the capability.
The thing is, FastWinDirStat uses a licensed propietary component. No problem for me, but the author did have some back and forth with another user on GitHub.
Seems FastWinDirStat license don't match with using a closed source library, or something...
As for its actual functioning, it does as it says. Works much faster than WinDirStat
(They could have been clear-ish (with caveats) by distributing only the source code and let the users do the compiling and linking, similarly to how you could download ZFS and build it into Linux. But you mustn't distribute the result further.)
https://github.com/qarmin/czkawka
I don't think it uses the MFT, but it's the fastest and most flexible open source dup finder as of last year or so
(This might seem obvious, but it took me a long time to realize, hence why I’m passing the tip on.)
It's not just an improvement, it's freaking astonishing.
I sincerely hope that the author open sources it one day, or that mtf-based solutions come to open source.
It's just life changing and will change when you want to do a scan. It removed any sense of hesitation or time waste.
I find myself much more willing to pop open wiz tree to get a quick view of my system or a particular storage folder.