The weird world of Windows file paths
fileside.app
fileside.app
All I can say is, this article is the tip of the ice berg on Windows I/O weirdness. You really don't realize how strange it is until you are actively comparing it to an equivalent implementation on the two other competing operating systems day-to-day.
Each of them has their quirks, but I think Windows takes the cake for "out there" hacks you can do to get things running. I sometimes like to ponder what the business case behind all of them was, then I Google it and find the real reasons are wilder than most ideas I can imagine.
Fun stuff!
PowerShell is also quite a powerful alternative to Bash/Mingw, although it came out much later.
Windows might do some things differently than UNIX-like OSs, but it does them really well.
Sometimes I'll do a "blank slate" and delete all my accumulated COM ports in Device Manager (need to enable "Show Hidden Devices").
But it's a DOS relic. Actually, Windows has a "Win32 Device Namespace" under "\\.\", which is something like /dev/ in Unixoids. COM1 is actually \\.\COM1: https://learn.microsoft.com/en-us/windows/win32/fileio/namin...
I'm glad our new releases work on Linux and we don't have to deal with that crap in 99.99% of cases now.
Your comment is very confusing to me. The serial ports are abstracted to a file on Windows just like on unixes - the file is actually discussed in the above article: \COM1
Maybe you're talking about the old days where you would just outb 0x3f8? The modern interfaces are actually fairly similar.
I spent a lot of time reading the disassmbly listing in the back of the manual to see what happens when I jump to the monitor.
COM9 = CreateFile("COM9", ...) => Nice!
COM10 = CreateFile("\\\\.\\COM10", ...) => NOT nice!
Under the hood this is because the USB devices do not have the (optional) unique serial number (or in some cases they all get the same serial number).
https://devblogs.microsoft.com/oldnewthing/20041110-00/?p=37...
It used to be completely predictable when I was working with drivers on 1994 (patching the code), then less predictable when hardware for more diverse, and predictable again (or at least "always the same") with UUIDs.
It was always amateur/hobby dev or sysadmin so I may have had the wrong impression.
As a side note: on modern Windows NT implementations the so called "registry bloat" is non-issue (anyone who tells your otherwise is trying to sell you something), but keeping list of every device that was ever plugged into the computer in there is somewhat ridiculous.
It's really a confusing thing to me that the script I use to change sound output and leveling suddenly didn't work after a bios/mobo software/whatever windows update and noticing the device have an appended (2).
Also, always locking an open file is repulsive. Other OSs allow for renaming an open file. Not Windows! Thumbs.db being locked because File Explorer keeps the file open / locked preventing deleting an empty folder and wastes so much time waiting for Windows to unlock the file.
You have to pay me to use Windows!
https://devblogs.microsoft.com/oldnewthing/20041110-00/?p=37...
How modern? I manage Windows 7 (transitioning to 10) machines that are used for QC in a hardware manufacturing environment that enumerate hundreds of devices (with mostly identical VID/PID) every week. We find that if we don't clear old devices out of the registry every so often, the enumeration time slows to a crawl.
In the times when it was a real issue (I would hazard a guess that that means “before XP”) the reason was that the original registry on disk format made every registry access more or less O(n) in bunch of things like the overall on disk hive size, total number of keys, number of subkeys in each of the keys along the path…
In one way it is beautiful. "Laid cards lies", you know. Don't mess with the user for your conception of Agile Clean Extreme Code (tm). Each stupid design decision is forever.
Windows .bat files win the awfulness and quirkyness contest with sh with a razor thin margin. And both are awesome.
Seems worth investing a lot time into given Microsoft’s history of not rug pulling developers.
Not that it’s a particularly compelling feature on Linux with the standard offering, but it’s a good option for cross platform scripts at times, particularly running in docker.
I think a lot of it gets down to ergonomics/aesthetics to decide if you find a useful niche for PowerShell for yourself. Python's os module is powerful and lets you run/chain almost any native commands and shell operations you want to spawn, but it is still a very different API and abstraction with different ergonomics and aesthetics than shell-style pipes and redirects.
PowerShell gives you that focus on shell-like pipes/redirects, but then gives you some Python-like power on top of that to also work with the outputs of some commands as objects in a scripting environment. There's a lot of interesting value to comparing/contrasting PowerShell and Python and if you are happy with Python maybe there isn't a big reason to learn PowerShell. PowerShell is there for when you are doing a lot of shell-like processing pipelines and want to write them as such, but have some of that power of a language like Python behind it. It's a lot more powerful than posix sh and it is similarly but differently powerful to Python but it starts from a REPL that looks/acts more like posix sh. I don't know if you have a need for that niche yourself, but I find it useful for that.
"Well yes, kind of..."
Because I learned computers when DOS was a thing, I will always be able to write a .bat or use CMD when necessary, but having been on the UNIX/Linux side since 2003, I didn't learn C# or PowerShell but rather bash, php, ruby. So while I'm friendly with modern Windows now that they closed up the "is it a stable OS" gap with Apple, I don't really know what to do in PowerShell and am more likely to use WSL!
Because not only it is a proper programming language, it is integrated into .NET, COM/DLLs as well, so not only you can script the OS, any application automation library is exposed as well.
Nowadays, it is possible to automate anything on Windows via PowerShell, the same OS APIs exposed by GUIs are also accessible to PowerShell.
On UNIX side there are things like Fish shell that also offer these capabilities, but they aren't widely adopted as PowerShell on Windows.
To give just one example: its automatic boxing and unboxing of arrays disqualifies it as a programming language. Try to return a one-element array from a Powershell function and you'll see what I mean.
The _beauty_ of bash is you can learn the basics of the language very easily, call out to external tools, and _take that knowledge with you_. Those tools exist independently.
It ends up with manual error handling to handle the case of bat, ps1 and exe's
`$PSNativeCommandUseErrorActionPreference = $true` is an experimental flag [1] as of PowerShell 7.3 that applies the same $ErrorActionPreference to BAT/EXEs (native commands), stopping (if $ErrorActionPreference is "Stop") on any write to stderr or any non-zero return value.
[1] https://learn.microsoft.com/en-us/powershell/module/microsof...
Why?
Because of backwards Compatibilytytytyy
5.1 was the last "Windows-specific"/"Windows-only" PowerShell (and is still branded "Windows PowerShell" more than "PowerShell") before it went full cross-platform (and open source). It's an easy install for PowerShell 7+ and absolutely worth installing. If you are using tools like the modern Windows Terminal and VS Code they automatically pick up PowerShell 7+ installations (and switch to them as default over the bundled "Windows PowerShell"), so the above command line really is the one and only step.
Winget isn't bundled with windows 10 either, (but I think it is with 11), and it's not on windows server.
If I need to install a package manager _and_ a shell, I might as well just install WSL and be done with it.
I can't speak to your usage of Windows Server, but provisioning winget and PowerShell 7+ are standard bootstrapping steps in VM images at places I work, because those are generally assumed to be basic equipment at this point.
I wasn't aware of $PSNativeCommandUseErrorActionPreference though, seems like it's very new. How does that work with the helpful windows tools that decide to not return 0 on success (hello, robocopy)
I had it randomly throwing exceptions the other day that a path was too long
(it was only about 300 characters...)
The user-mode stuff is kind of a mess. The kernel-mode stuff is comparatively orthogonal.
[1]: https://learn.microsoft.com/en-us/windows/win32/api/fileapi/...
[2]: https://learn.microsoft.com/en-us/windows/win32/fileio/namin...
* for %%a in ("*.mp4") do ffmpeg -i "%%a" -vcodec libx265 -crf 26 -tune animation "%%~na.mkv"
gci -filter "*.mp4" | foreach { ffmpeg -i $_ -vcodec libx265 -crf 26 -tune animation ($_ -replace '.mp4','.mkv') }
Files with spaces just work.My main complaint about sh and Unix is that environment variables are mostly inherited from parent processes, and that can be awkward.
(And Windows at least, these days, is seeing a lot more PowerShell adoption.)
Ye agreed. Separation of data and code always was a mistake.
I wonder if this feature of bat files was like a thing once as a "best practice"? Practically, you should only append lines I guess. When I close my eyes, I can see a DOS batch file doing actual batch job processing, appending to itself and becoming a intertwined log and execution script.
Meanwhile macOS appears to just change an internal link to the data that’s already written on disk. As such, it’s usually so very fast compared to Windows.
One other thing that can be an issue particularly on NTFS with ACLs is that moving files typically retains their ownership and permissions, while copying files typically inherits the ownership and permissions of the destination. This can bite you if as an administrator you're moving data from one user's account to another because a move will leave the original owner still owning the files.
If you need to seriously move/copy lots of files or lots of data in Windows it is generally a good idea to use shell commands. Robocopy [1], especially, is one of the strongest tools you can learn on Windows. (It gets very close to being Windows' native rsync.)
[1] https://learn.microsoft.com/en-us/windows-server/administrat...
Windows Explorer is also slow for an unknown reason.
Doing file operations through the API with Real-Time Protection turned off is several orders of magnitude faster in the case of small files. It's crazy stuff.
Well, then, is there a more detailed summary than this one that's accessible?
This one looks very useful and I'll use it, but to make the point about more info it'd be nice to know how they differ across Windows versions.
For example, I've never been sure about path length of 260 and file length of 255. I seem to recall these were a little different in earlier versions, f/l being 254 for instance. Can anyone clear that up for me?
Incidentally, I hit the 255/260 limit regularly, it's damn nuisance when copying stops because the path is say 296 or 320, or more.
Many apps do something like
char path[MAX_PATH]
In that case, no amount of prefixing will help you, if random app enforces the limit.
I've noticed that, it's partially the reason for my confusion (I didn't wake up for quite a while as I put it down to the different versions of Windows I was running on various machines). Other pains are caused by apps that still don't run Unicode and crash or stop copying when they encounter a non-ASCII character.
"...(this value is commonly 255 characters)."
I think the word 'commonly' in those notes confirms my point in that it's changed slightly over the years. Also, I recall the internal processing was once 16k and not 32k, come to think of it this may have been with the previous version of NTFS (can't remember which version of Win that was). My interest is now piqued so I'll search it out.
That we're even discussing such matters confirms the thrust of the article.
Programs, e.g. python scripts, can then use those paths, but only after explorer has resolved them. Before that, the path will be treated as not existing.
>Windows I/O weirdness. You really don't realize how strange it is until you
"can I get a wut... WUT?" - Developers, Developers, Developers
====
it's just hard for people today to understand what a revolution stdin and stdout was, full 8 bit but sticking with ASCII as much as possible. There was nothing about it that limited Unices from having whatever performant I/O underneath, but it gave programmers at the terminal the ability to get a lot done right from the command line.
The web itself is an extension of stdin and stdout, and the subsequent encrusting of simple HTTP/HTML/et al standards with layer upon layer of goop that calls for binary blobs can be seen as the invasion of Cutlerian-like mentalities. It's sad that linux so successfully took over that all the people who we were happy to let use IIS and ASP and ActiveX had to come over to this side with their ideas. No idea of which is bad, but which together are incoherent.
client server FTW, bring it back
$ create/directory dm0:[foobar]
$ set default dm0:[foobar]
$ copy DM0:[SYSMGR]SYSTARTUP.COM example.txt
$ dir/full
DIRECTORY DM0:[FOOBAR]
21-APR-23 14:42
EXAMPLE.TXT;1 (615,2) 0./0. 21-APR-23 14:42 [1,4] [RWED,RWED,RWED,RE]
TOTAL OF 0./0. BLOCKS IN 1. FILE1. unix doesn't expose devices in the file tree, and thank the [deity] for that
2. directory part: everything to the last /
3. filename: everything afte the last /
4. TEX_ROOT=/foo/bar; cd $TEX_ROOT/MF
Could it be any easier?
I mean, they are exposed as files (e.g. /dev/sda) within the file tree, but they aren’t exposed by a file path.
I have a vague notion that there might have been the ability to have a virtual device span multiple directories, but as I think about it, this seems unlikely since it would make it ambiguous where something would be created if a directory exists in both root1 and root2, so perhaps not. It’s been 23 years since I last used VMS and 30 since it was my daily driver though, so it’s hard for me to say too much.
Distributed FS sounds like something with a huge design space. It's better to do that in userspace. Was Plan 9 really the first time virtual file systems could be implemented in userspace? It seems like such an obviously useful idea in retrospect.
https://en.wikipedia.org/wiki/Diff3
Can you really call what we see in VMS full revision control when it lacks this capability?
Right, my complaint about filename/path length being exceeded in my earlier post often occurs when a web page is saved by a browser (some web pages have outrageously long filenames).
Incidentally, many years ago I did a tour of Microsoft's operation in Seattle around the time Microsoft introduced subdirectories into MSDOS and the tour guide (can't recall his name but he was responsible for the development of MS's Flight Simulator) gave a considerable spiel about why Microsoft decided to run with the backslash instead of the forward one as per Unix. Even then, I thought 'oh no, here comes confusion', and others with me thought the same. When we challenged him about it, he said we (Microsoft), want to clearly differentiate ourselves from Unix (there was an arrogance about his answer that I well remember).
I'm sure you heard what you thought you heard, arrogance and all, but iirc it was backslash because IBM insisted the slash be the switch character. Anybody remember $SWITCHAR
It's too long ago for me to associate the name 'Alan Boyd' with the person in question but I do remember that he had a loud, penetrating self-assured voice. (Incidentally, he spent considerable time demonstrating Flight Simulator's new features).
You're right, IBM was a large part of the discussion as back then it was the principal client for MSDOS. However, I came away from the visit with the understanding that MS was in full agreement with IBM's decision despite MS's dabblings with Unix.
I had a particular interest at the time as I had a S-100 Godbout CompuPro computer, and in addition to CP/M, I ended up putting Seattle Computer Products' DOS (SB86 from Lifeboat Associates) on it which meant that I had compatibility with MSDOS.
I can understand why MS would have wanted to differentiate MSDOS and the backslash being one way, what I'm still not clear about is why IBM would have wanted to make such a distinction.
Re $SWITCHAR, very vaguely, but from my point it added confusion. Like some other commands its implementation and architecture appeared to be the result of afterthought rather than good design. I've forgotten much of that stuff.
Filesystem paths used to all be weird in the sense that there was more OS diversity. I'm sure some people here remember that classic MacOS paths used colon as the separator:
Hard Drive:My Folder:My Document
VMS (designed by the same person as Windows NT by the way) had paths that looked like this (per Wikipedia): NODE"accountname password"::device:[directory.subdirectory]filename.type;ver
RSTS/E had [project,user] in the filename.Multics paths:
>dir1>dir2>dir3>filename
Apple Lisa used dashes as the separator. Etc.One fun surprise is that because of codepage reasons, the Windows will use ¥ as a path separator in Japanese. In Korean, it's ₩. These characters represent U+005C, which is \ in Latin-compatible character sets.
First floppy drive was A, Second B, and when internal Hard Drives came along they defaulted to C to be compatible with computers that had at least 2 disk drives.
https://en.wikipedia.org/wiki/Drive_letter_assignment?wprov=...
Starting from A and iterating on through Z makes sense, for an OS that's designed for two drives at most. /dev/sda and /dev/sdb are no less arbitrary than A: and B:.
One major difference was that Unix was used on big servers and couldn't fit itself onto a single disk, so /usr had to be created. DOS and Windows never needed a second drive to boot, so they didn't need to embed their resources into the drive hierarchy.
Of course, you can mount NTFS volumes at any directory you wish since at least somewhere in the early 2000s. Very few people do it, but you can!
For example:
$Disk = Get-Disk 2
$Partition = Get-Partition -DiskNumber $Disk.Number
$Partition | Add-PartitionAccessPath -AccessPath "G:\Folder01"HPFS had extended attributes, but not substreams. You are thinking about HFS; substreams were added to NTFS to support storing resource forks on network shares used by Macs.
$dir1.dir2.dir3.filename ADFS::IDEDisc4.$.Games.!Repton.Arctic
ADFS is the filesystem, IDEDisc4 is the disc name, $ is the root directory, Games is a subdirectory, !Repton is an application directory (since it begins with !) and Arctic is a file within the application directory, not normally referenced by users. Resources:$.Apps.!Edit
is the application !Edit from the built-in ROM.In modern macOS (previously OS X), you’ll eventually bump into those if you need to work with paths in AppleScript. You have to specify when you’re using a POSIX path so it is properly converted. Example:
$ osascript -e 'POSIX file "/System/Applications/Mail.app/Contents/MacOS/Mail"'
=> file Macintosh HD:System:Applications:Mail.app:Contents:MacOS:MailNow change it to have a forward slash. macOS will happily abide.
Finally, look at that file’s path in a terminal. Where the Finder shows a forward slash, the terminal will show a colon.
Redoing the AppleScript example:
$ osascript -e 'POSIX file "/tmp/file with : forward slash.txt"'
=> file Macintosh HD:private:tmp:file with / forward slash.txt(sort of amazing the original premise, and the exceptions and workarounds you gradually accumulate and take for granted)
The `/` root dir is quirky. You can’t just do `dirPath.split(‘/‘)`. You have to handle it as a special case. Would be easier if it had a special name. Like `$/dir1/dir2`.
Or am I missing something.
Interesting. When you think about it, that doesn't look all that different from:
scheme://user:password@host:port/directory/filename.type?key=value&key=value#fragment
which is arguably the most common kind of "file path" in use today.
The volume id string (what you get with mountvol) is - at least up to Windows 10[1] a UUID version 1 according to RFC 4122, i.e. time and node based:
https://www.famkruithof.net/guid-uuid-make.html
https://www.famkruithof.net/guid-uuid-timebased.html
Since windows creates the UUID the first time it "sees" a volume, and - usually - uses the network card MAC as node, by decoding the UUID you can get the MAC address of the PC and the time the volume was seen (this can be useful for forensics, expecially with removable devices and to verify there has been no manipulation of the MountedDevices in the Registry).
[1]possibly windows 11 changed that, or at least the UUID's shown in the article are type 4
> UNC paths can also be used to access local drives in a similar way:
> \\127.0.0.1\C$\Users\Alan Wilder
> UNC paths have a peculiar way of indicating the drive letter, we must use $ instead of :.
This is actually incorrect... He's actually accessing some random share that has no real connection to a drive. Yes, sometimes (quite often), the C$ share corresponds to the C: drive's root, but this is by no means given, as one can easily either delete the C$ share, or have it pointing to somewhere else entirely
> When the current directory is accessed via a UNC path, a current drive-relative path is interpreted relative to the current root share, say \\Earth\Asia.
This is also wrong. There is no "current directory" on an UNC share (which can easily be shown by trying to open a command prompt on a UNC share, it will show an error and start you somewhere on C:\users), and the example he gives just tries to access the share "Asia" on the server "Earth"
> Less commonly used, paths specifying a drive without a backslash, e.g. E:Kreuzberg, are interpreted relative to the current directory of that drive. This really only makes sense in the context of the command line shell, which keeps track of a current working directory for each drive.
Also wrong, it's not the command line shell that keeps track of the current directories, it's the Windows kernel itself. But I agree that such a scenario is quite useless as you can never be quite sure on what CWD you are on a given drive
> For the most part, : is also banned. However, there is an exotic exception in the form of NTFS alternate data streams.
Yeah, well, surprise: the ":" is not part of the file name, it's just a separator between filename and stream name. This is like saying that "you cannot have \ characters in a file name, but in directory names it is allowed". No, it's not. It's a separator
SetCurrentDirectory allows setting the current directory to a UNC share. https://learn.microsoft.com/en-us/windows/win32/api/winbase/...
> Also wrong, it's not the command line shell that keeps track of the current directories, it's the Windows kernel itself. But I agree that such a scenario is quite useless as you can never be quite sure on what CWD you are on a given drive
Not for a long time. It's set as a special (hidden) environment variable like `=C:=C:\current\directory`. https://devblogs.microsoft.com/oldnewthing/20100506-00/?p=14...
Exactly. The cmd prompt not setting UNC paths as current directory was introduced around Windows 2000 (or maybe post-XP, it's been a while) to help legacy batch files being run from a share and then getting confused by being on a UNC path instead of one beginning with a drive letter.
This was also why, when you do a pushd \\server\share cmd.exe puts you on a mapped drive instead of directly on a UNC path.
If you use the Windows native version of tcsh, for example, you can happily use UNC paths as current directories and run commands (provided they don't try to parse drive letters from their CWD)
Eh, I just thought the $ in a Windows NAS share was to ensure the share was hidden from browsing. Microsoft used to have documentation on that, but seems to be missing from their site after they removed old articles.
It should also be noted that while the single driver letter ones are automatically created, the "$" at the end just marks them as hidden. You can create your own hidden shares if you ever want to.
In fact this probably isn't the worst thing - it's even worse than this. Because you first need to strip off the prefix that represent the volume (like C:\) before you can look for a colon. But the prefix can be something like \\.\C:\ or \\.\HarddiskVolume2\, or even \\?\GLOBALROOT\DosDevices\HarddiskVolume2\. Or it can be any mount point inside another volume! (Remember that feature inside Disk Management?)
Moreover you can't even assume the colon and alternate data streams are even a thing on the file system - it's an NTFS feature. So you gotta query the file system name first. And if the file system is something else with its own special syntax you don't know, then in general you can't find the file name and strip the last component at all.
All of which I think means it's impossible to figure out the prefix length without performing syscalls on the target system, and that the answer might vary if the mounts change at run time.
Data stream is basically the file content and on NTFS a file can have more than one. In practice it is comparable to extended attributes in the Linux world but somewhat superior.
But like extended attributes it doesn't seem to have too much real world use. The only use case for alternate data streams I can remember are the "this file was downloaded from the internet, do you really want to run it" warnings. In such cases the browser attached a standardized marker as alternate data stream to the file.
You seem to talk about a specific command line argument of the compact command with a Windows typical (and IMO ugly) option style with '/' instead of '--' as option marker and ':' instead of '=' as option value separator.
But that would not be directly related to ADS and I cannot imagine a good use case where the compact command should use ADS.
In context of ADS the first thing I imagined was storing the compressed and uncompressed file alongside. (which is rather silly, why compress at all)
This use case is also kinda strange. Have the compressed content as ADS, the primary contend filled with 0 as sparse and fill it when needed/accessed. :/
In general, there is a tool which comes with the SysInternals suite that allows you to see which files have streams and their size:
https://learn.microsoft.com/en-us/sysinternals/downloads/str...
macOS stores fonts in resource forks? I'm confused, what use does this have and what happens when you accidentally miss them?
Classic MacOS considers fonts to be a type of resource, and hence stores them in the resource fork. Contemporary macOS fonts are just ordinary files with a data fork only. I think grandparent is talking about the 1990s, although some of those machines remained in active use through the first few years of this century.
Windows originally considered fonts to be a type of resource too – the original bitmap fonts used with Windows 1.x-3.x are stored as a resource–except unlike MacOS it embeds resources into EXE/DLL file data instead of putting them in a fork. In fact, a .FON file containing a Windows bitmap font is just an EXE with no code, only resources. Nobody really uses this any more, everything is TrueType now and TrueType uses its own file format not resources, but Windows still supports the old bitmap fonts for any legacy apps which still use them.
Anyway... so macOS fonts themselves were made of resource forks and therefore trying to transfer fonts themselves across a non-resource-fork-supporting network share will fail? As in, the resource forks were needed in order to use the font file?
Not me. ajcoll5 made the statement, you expressed confusion with it, I tried to explain what (I assume) ajcoll5 meant.
> Anyway... so macOS fonts themselves were made of resource forks and therefore trying to transfer fonts themselves across a non-resource-fork-supporting network share will fail? As in, the resource forks were needed in order to use the font file?
On Classic MacOS, some files, all the actual contents is in the resource fork, and the data fork is ignored and can be empty. So you copy such a file to a filesystem which doesn't support resource forks, you can end up with an empty file.
A good example of this is executables. 68k Mac executables, all the code is stored in the resource fork (as code resources), and the data fork is ignored and can be empty. So you copy a 68k Mac executable to a forkless filesystem, you can end up with an empty file.
By contrast, PPC Classic MacOS executables, the code is in the data fork, and the resource fork only contained actual resources such as icons or strings, not the code. If you lost the resource fork, you'd still have the code of the executable. But it probably won't run without the icons/strings/etc it expected.
This was how Apple's original (1994) implementation of "fat binaries" worked. The data fork contained the PowerPC binary and the resource fork contained the 68K binary. PPC Macs would load and run the PPC code from the data fork, 68K Macs would ignore the data fork and load and run the code from the resource fork. If you only needed PPC support, you could shrink the executable by deleting all the 68K code resources from its resource fork.
The core resources of Classic MacOS were originally stored in a single file, the "System suitcase". Originally, each installed font was a separate resource in the resource fork of that file; its data fork was unused, except to store an easter egg text message. Fonts were distributed as resources in separate suitcase files, and the "Font/DA Mover" copied them from the distribution suitcases into the system suitcase. So yes, a suitcase file used to distribute a classic MacOS font, the actual font data would be in the resource fork, and the data fork could be empty. In System 7.1, Apple introduced a separate folder called "Fonts". In some MacOS versions (not sure when it was introduced, but definitely was there by System 7.0), Finder displays suitcases as if they were folders, even though they are actually resource forks.
Contemporary macOS doesn't really use any of this stuff. It supports resource forks for backward compatibility, but modern applications don't use them. The "Font Book" app can import Classic MacOS fonts (not bitmap ones, but TrueType and Type 1) from the resource fork of a suitcase file. But once imported, the fonts are stored in ordinary files (with a data fork only) on the filesytstem.
Eh, whatever. I originally thought whoever meant. can't edit the comment now.
> On Classic MacOS, some files, all the actual contents is in the resource fork, and the data fork is ignored and can be empty. So you copy such a file to a filesystem which doesn't support resource forks, you can end up with an empty file.
Yeah, that's about what I thought. That makes sense, thank you~
You think I jest ? Look up the leaked source code for the US government spy tooling. They hide data to be exfiltrated in an ADS on the root directory of the share :-).
I finally realized ADS were the mother of bad ideas when Ted Tso responded to me asking why I couldn't have them in Linux for the umpteenth time by showing me a Windows task manager screenshot of Myfile.txt as an actively running process.
If the ADS ends in .exe then Windows will happily run it :-).
If the :stream syntax is not FS-specific then you can parse the data stream name out statically in almost every case. Yes, you have to work out the prefix, but you can mostly do that statically too, I think:
> In fact this might not even be the worst thing - it's even worse than this because you first need to strip off the prefix that represent the volume (like C:\) before you can look for a colon. But the prefix can be something like \\.\C:\ or \\.\HarddiskVolume2\, or even \\?\GLOBALROOT\DosDevices\HarddiskVolume2\. Or it can be any mount point inside another volume! Which I think means it's impossible to figure out the prefix length without performing syscalls on the target system, and that the answer might vary if the mounts change at run time.
The prefix of `\\.\C:\Foo:Bar` is `\\.\C:` as `C:` couldn't be a file name. The prefix of `\\.\HarddiskVolume2\Foo:Bar` is `\\.\HarddiskVolume2` because the volume name ends at the backslash. The prefix of `\\?\GLOBALROOT\DosDevices\HarddiskVolume2\Foo:Bar`... can be harder to determine but it doesn't matter because clearly there is no letter drive name in sight since a letter drive name would be... a single letter, but if the volume name were a single letter then it might require using system calls to resolve it (`\\?\GLOBALROOT\DosDevices\X\Y:Z\A:B` is harder to parse because X might be the volume name, or maybe Y: might be the letter drive and X might be part of the path prefix).
It is, I believe, as I alluded to in the comment.
> `\\?\GLOBALROOT\DosDevices\X\Y:Z\A:B` is harder to parse
As in, this is impossible to do statically in the general case - those names aren't guaranteed to look like that. See the note I had added about mount points. Remember C:\mnt can itself be the mount point of a volume instead of a drive letter. (Junctions present a similar problem, but at least for those, you can make an argument that they're intended to look like physical folders, and treat them similarly. With mount points, you might not have that intention - you might be just trying to go over 26 drive letters.)
The FILE_STANDARD_INFORMATION_EX structure alludes to a common handling of alternateStream. Winbtrfs is a great resource on this, since it implements many bells and whistles from NTFS in an open way -- you just grep for a keyword and you will be close. The code exercising the Windows API for testing is src/tests /streams.cpp.
Grep on FILE_STREAM_INFORMATION in the source should provide more useful hits on the source, but phone browsers are clumsy.
> All Unicode characters are legal in a streamname component except the following:
> * The characters \ / :
> * Control character 0x00.
> * A streamname MUST be no more than 255 characters in length.
>
> A zero-length streamname denotes the default stream.
https://learn.microsoft.com/en-us/openspecs/windows_protocol...
Though ironically that still doesn't help you strip the last component, since it could still be a volume mount point. Like you don't want C:\mnt\..\foo to suddenly become C:\foo, just like how you don't want \\.\Server\Share1\..\Share2 to become \\.\Server\Share2, or for \\.\C:\..\HarddiskVolume1 to become \\.\HarddiskVolume1, etc.
NTFS would accept almost anything. The Windows API (I think of the old Win32 one) would apply most restrictions the article mentions.
But for example not the normalization part. A filename can end with a space, no problem. That lead me once to a minor bug in .NET Framework. One of the path related functions (I think it was Directory.move) did not correctly apply this normalization and could produce directories with trailing whitespace. Good luck removing/fixing those in Windows Explorer.
So for the longest time Adobe software had random bugs where it would create a series of folders name "Application Data" repeating recursively 3000+ characters deep.
Yea, that was fun to try to delete.
In practice I think the biggest problem with using forward slashes on Windows is confusing programs which expect "/" to indicate program switches. The non-uniformity of shell parsing is also a big unix/win design difference.
Windows supports it, CMD doesn't. programs that you run from a CMD prompt support other options flag syntaxes, so it's just a cmd.exe feature.
CMD.exe is its own thing with its own backwards compatibility requirements and the case could be made that cmd.exe is "Windows" as much as anything else is, so I get it.
As another example, you can’t use forward slashes in the File Open dialog of Visual Studio: https://developercommunity.visualstudio.com/t/allow-forward-...
In fact, you can use forward slashes across the entire file API on Windows. That's the point.
I don't believe that's true, I am almost positive they're SMB shares, just like any other, but are created by the system, which is why "accessing drives in this way will only work if you’re logged in as an administrator."
you are correct that they are just SMB shares like any other. They can be removed, though many management processes across different applications assume that those shares will be present
(In version 10.0.19041.985 of cscsvc.dll in Windows 10 I'm seeing the string "If you hit this breakpoint, send debugger remote to BrianAu." Presumably that's "Brian Aust", referenced in a chat[0] re: Offline Files.)
[0] https://techcommunity.microsoft.com/t5/storage-at-microsoft/...
Lots of these (eg: the COM/LPT stuff) could be dropped and wouldn't affect most people either way, but for those things depending on it, it would be a profoundly breaking change.
'echo foo > COM1' returns 'The system cannot find the file specified.' on Windows 11. (Machine doesn't have a COM1; if this wasn't being redirected to the port, it'd have gone into a file of that name.)
The Win32 paths are like an emulation layer. They parse the given path and produce a kernel path. Win32 implements all the weird history you know and love as well as things like `.` and `..`. You can use the `\\?\` prefix to escape this parsing and pass paths to the kernel.
The NT kernel has paths like `\Device\HarddiskVolume2\path\to\file`. NT paths are much simpler. There are no restrictions on paths except that they can't contain empty components (notably they can contain nul). At this layer, `.` and `..` are legit filenames.
However, it's the filesystem driver ultimately says what's a valid filename and what isn't.
Oh no. No. Windows allows files to be named `..`?!
Maybe not?
Under unix, if you create a symlink to a directory, e.g. `~/syslogs` is a symlink to `/var/log`, then `..` can be used to traverse the "true" parent directory. So `~/syslogs/../lib` will traverse `/var/log/..` and refer to `/var/lib`, not to `~/lib`.
However, a "normalising" path interpreter will just take something like `~/syslogs/../lib` and change it to `~/lib` without consulting the filesystem.
Given that (AIUI) Windows has supported symlinks for a while now (?), it's possible that files called `..` aren't actually allowed, but the ability to access `..` is still necessary.
(Notably, the article does point out that filenames ending in `.` are disallowed - which should exclude `..` as a name one can give a file.)
You can put `..` earlier in the file name, though.
https://msfn.org/board/topic/131103-win_nt~bt-can-be-omitted...
But seriously: no, at least not on NTFS. This filename does have trailing space. Though it is enough to defeat Explorer, you cannot move or delete it and properties window is broken.
I also routinely use single extended unicode characters as root folder names and identifiers for various purposes.
Using a search programme 'Everything", it's a lot easier to find things if I use something like pilcrow symbol as the root folder for any directory dedicated to text documents, when the alternative is to wade through results for 'documents', 'text', 'reading' or any combination of those words.
For the same reason, I find I can make much more memorable associations. It helps me harness things relationally. I can preserve uncertainty and avoid the frustration and negativity of trying to make shades of grey and rose fit black and white patterns. It does sound a bit new age, but there's no doubt in my mind, flat heirachical alphanumeric patterns are restrictive, prescriptive, insufficient. For example, a lot of artists actively work to defy pidgeon holing. I still need identifiers.
I mean, even if I wasn't into 'bleeding edge' culture, restrictions, problems and frustrations are the normal experience. I think this is illustrated by the unsatisfactory experiences that people find when they try to make id3 tagging "work".
It's as close as I can get to banishing the pervasive 'what-if' heartbreak of WinFS being cancelled. Sadly it doesn't help at all make up for what 'Semantic Web' promised. But that's probably why I'm a believer in GPT and the like.
Is it just me that can't help thinking they are products that have arisen from the need to make non-semantic computing useful again?
Windows explorer could not delete the file. You have to specify the \\?\ path to get the delete call to work, but that didn't work well with cmd.exe's `del` command.
I've since used these files to create directories that can't be deleted by automated cleanups and such, like a special folder in %TEMP% that one program needed but didn't create on its own.
That was necessary to support the use case where an older OS tried to read the disk (could happen because the user rebooted into an old DOS, for example, or if an external disk was moved to a different computer)
https://en.wikipedia.org/wiki/8.3_filename#VFAT_and_computer...:
“VFAT, a variant of FAT with an extended directory format, was introduced in Windows 95 and Windows NT 3.5. It allowed mixed-case Unicode long filenames (LFNs) in addition to classic 8.3 names by using multiple 32-byte directory entry records for long filenames (in such a way that only one will be recognised by old 8.3 system software as a valid directory entry).
To maintain backward-compatibility with legacy applications (on DOS and Windows 3.1), on FAT and VFAT filesystems an 8.3 filename is automatically generated for every LFN, through which the file can still be renamed, deleted or opened, although the generated name (e.g. OVI3KV~N) may show little similarity to the original. On NTFS filesystems the generation of 8.3 filenames can be turned off. The 8.3 filename can be obtained using the Kernel32.dll function GetShortPathName“
The other direction worked great, though, DOS filenames always worked on the Mac side of the network.
I realized at some point that there is a discrepancy between what's allowed on the file system, and "Windows" itself (or, more exactly, the programs running on Windows and using its APIs to communicate with said file system.
In this case, NTFS, totally allows for "illegal" characters such as < > : " | ? * etc... pretty much everything except / and \, and \0, I think.
This makes for funny situations, where sometimes Windows programs cannot deal with that. At best, they can't read, write or rename them... at worse they'll crash, which is always fun.
I think this is an interesting space with room for innovation.
Fileside starts out with a grid of four directories: Home, Documents, Desktop and Downloads. You can customize and name new grid layouts that are shown in a sidebar for quick switching. This seems like a neat idea for specific recurring manual workflows.
It's doesn't seem to be targetted to the minimalist crowd. Directory entries beginning with a dot are visible (but greyed out). Full Unix-style permissions are shown for each entry, etc.
It looks like it's Electron-based and implemented in a javascript SPA framework. It doesn't use the default system font (SF Pro) on Mac. A bunch of other things also don't look or behave as you expect.
The font weight in the size column maps to each file's relative size. All the way from very thin to very bold. Kind of cute.
The path completion seems pretty good - as could be guessed from the blog post.
I think this app sometimes confuses power with details/verbosity. There are some gold nuggets in there though.
Here's an alias from my Cygwin .bashrc (which took me way too long to figure out) where both the *nix and Windows style paths are invoked:
alias ms='/cygdrive/c/Windows/System32/OpenSSH/ssh.exe -A -i 'C:\Users\me\.ssh\mm-id_rsa' me@myserver'
$ tasklist /V
ERROR: argument/option invalid - "C:/msys64/V"
or the time I wanted to use xmlstarlet and any xpath expression was interpreted as Windows path :(At least for MSYS2 the environment variable MSYS2_ARG_CONV_EXC can be used to prevent conversion. https://www.msys2.org/docs/filesystem-paths/#process-argumen...
(I forgot how this topic is called though and on what layer it takes places.)
Windows XP already could do it, but didn't for the most part. I remember, like ~18 years ago, I was a sysadmin. One user out of 50 got his "My Doc/Pictures/etc." in English but for me on the File-Server it was all in German. Very confusing.
Kind of? Surely unix file path conventions are of a similar age (if not older), yet somehow they seem to exhibit much less weirdness..
The “and this is why fileside exists” transition at the end - perfect. No “sign up” or anything. Awesome!
To go up to the parent directory you had to use an additional slash. So /xxx was the equivalent of Windows ..\xxx, and you could add more slashes to go further up in the directory tree: Work:a/b/c/d////file was the same as Work:a//file.
Drives could be "virtual" ones, similar to "bind" mount points in Unix, that could be associated with multiple positions. E.g. you could assign both System:Libs and Data:MyLibs to the virtual drive LIBS: so that LIBS:xxx would match a file called xxx in either directory.
Files used to have a comment field to store kind of extended attributes, but was seldom used IIRC.
Wildcards were quite unique, I think ? was like regexp . meaning any char, and # like * but prefixed meaning any number of the _following_ char, so that #? would match anything.
I'm sure there were other niceties I can't remember right now.
MSDOS included some compatibility for CP\M and everything since has maintained compatibility with the version prior, except for a few exceptions.
so even today we have compatibility built into Windows for things that don't exist anymore in any real capacity.
Microsoft is very serious about backwards compatibility.
programs written using ANY of those technologies run on Windows 11 unmodified, and they will for Windows 12, too.
backwards compatibility means new OS versions can run programs written for older OS versions.
backwards compatibility promises do not prevent you from coming up with new ways to write programs.
I am on the Microsoft ecosystem since MS-DOS 3.3.
backwards compatibility is not about keeping all features once supported in visual studio in all future versions. that is forwards compatibility. Microsoft does not do that.
we are talking about backwards compatibility: the ability of new operating systems to run software unmodified which ran on old versions of the same operating system.
echo > \\?\C:\path\to\file.
and echo > "\\?\C:\path\to\file "
Similarly, files with such names can be created with Cygwin.Then, the author goes on to explain that every path on Windows starts with "\\" instead. :)
File.join(too_long_identifier, some_other_long_identifier).gsub('/', '\\')
is often useless and nauseating to read when "#{dir}/#{filename}"
is portable enough for most of Ruby's File APIs.> < > ” / \ | ? * Never allowed
A friend once somehow crated a file called <HTML>.
I don't know how, but he also couldn't delete or do anything with it.
Connecting from another operating system that allowed names like that to be corrected to a Windows share.
There are a few other possibilities where you boot to Linux using a FS driver for NTFS that allows you to create illegal file names. And/or odd things like WSL/Cygwin.
Turns out you could create files with "illegal" (and invisible) characters in the filename. The standard OS utilities would not allow them, but the underlying file system did not care. So you could write a short program to do it.
I had to write a utility just to delete it.
"Broad software compatibility was initially achieved with support for several API 'personalities', including Windows API, POSIX, and OS/2 APIs – the latter two were phased out starting with Windows XP."
When I made carefulwords.com, which I made because I wanted a thesaurus where you could just write eg https://carefulwords.com/book for "book" and get the results, I found out the hard way that you cannot make a file named "con" on Windows. Or "con.html", or any file extension. You can try to force this, make it via a script, but then programs like Git will hang when they come across it. So in my thesaurus the actual page is /con-word.html and I just have it rewrite to /con
Gonna give it another try this weekend.
Your life will end up as a series of awful MYPATH~1\ kludges. You have been warned.
EDIT: There seems to be different camps here and perhaps a generational divide. Maybe the kids haven't (or never will be) burned with this one, but i've seen too much, wasted too many hours, wrote too many workarounds, and will forever remain #TeamNoWhiteSpace.
$ cat > my.patch <<PATCH
diff --git a/alpha beta/charlie b/alpha beta/charlie
new file mode 100644
index 0000000..3b18e51
--- /dev/null
+++ b/alpha beta/charlie
@@ -0,0 +1 @@
+hello world
PATCH
$ patch -p1 < my.patch
patching file 'alpha beta/charlie'
$ patch --version
GNU patch 2.7.6...admittedly i'm not the best with Powershell. There are probably workarounds or things that i did wrong.
Man, it's never that simple. How about generating Powershell scripts with quotes? Sometimes you are dealing with escape characters \", sometimes you aren't. Why bother with all that? Hell, versions of software and compilers change such behaviors either on purpose or by accident all the time. It just gets messy dude. C:\NEVER_~1\AGAIN_~1.NOP
Perhaps i'm a dinosaur still scarred by the golden olden days, but, there is no convincing me that whitespace in filenames or URLs is ever a good idea.
Man, i've had config files break because i saved them in UTF-8 instead of ANSI. I hope to god you never have to experience the horror... You also have a nice day.
Isn't this because random Windows tools add garbage bytes to the start of the file (a "byte order marker") and then don't tell you they did this?
It's not the encoding that's the problem, it's the garbage bytes.
P.S. I do know of exactly one program that only accepts UTF-16 for its config file, and won't accept UTF-8 or ANSI.
I knew the guy who tested the feature in PowerShell and he was deeply frustrated over the fact that there was no way to escape some sequences. For example, you can loosely match "[abc]" with something like "?abc?" or "[[]abc?" because the [[] indicates exactly one opening square bracket but there's no way (as far as I know) to say "this section ends with a square bracket".
The PM really wanted that feature, though, even though you could probably count on one hand how many times it's been used in real life.
Edit: That guy down there corrected my example.
> ni '[abc]'
Directory: C:\Users\<user>
Mode LastWriteTime Length Name
---- ------------- ------ ----
-a--- 21/04/2023 12:23 0 [abc]
> gi '[abc]'
# no output - would have matched files named a, b, or c
> gi '`[abc`]'
Directory: C:\Users\<user>
Mode LastWriteTime Length Name
---- ------------- ------ ----
-a--- 21/04/2023 12:23 0 [abc]
> gi -LiteralPath '[abc]'
Directory: C:\Users\<user>
Mode LastWriteTime Length Name
---- ------------- ------ ----
-a--- 21/04/2023 12:23 0 [abc]
> gi -lp '[abc]'
Directory: C:\Users\<user>
Mode LastWriteTime Length Name
---- ------------- ------ ----
-a--- 21/04/2023 12:23 0 [abc]
As far as I know that's been in powershell forever. gi "`[abc`]"
...because the `[ will be processed before the string gets sent to globbing, meaning the back tick will be removed. This would work with double quotes: gi "``[abc``]"
...because at the command line level this will evaluate to `[abc`] and the globbing will know that the square brackets are literal. So I will concede and downgrade my complaint from "impossible" to merely "overly complicated for the layman".Not necessarily. There are all sorts of scripting languages where single and double quotes are identical.
It is, however, super common for there to be escaping differences in shell scripting languages.
When writing software I always make sure that all my test data has spaces - and not just the normal one but weird unicode ones too - in paths and filenames.
Windows broke. Programs broke. The registry broke. Special characters work well enough, but never use them in standard directory paths.
If you want your folder to be written correctly in your own language, you can use desktop.ini to correct the name of the folder in Windows Explorer:
[.ShellClassInfo]
LocalizedResourceName=Fàñćŷ Näm̀é
Put that into a desktop.ini file, assign it the right attributes (`attrib +S +H desktop.ini`) and your folder will look right in any program that uses the shell API without any risk of breaking programs.I grew up in days of DOS 3 and onward, so I am basically physically incapable of using whitespace in filenames. And frequently I'll replace white spaces with underscores on files shared with me.
Fascinatingly though (and sometimes irritatingly), several of my mentors have cautioned me to drop this habit as I get promoted. At management level, they all use white space and dots haphazardly, and they apparently perceive underscored filenames negatively. This goes up drastically at executive level.
And most programs are able to handle that just fine.
Really? You posted a divisive and absolute statement and people pointed out how it's not accurate. Considering "Program Files", a very common Windows path, was introduced back in Windows 95 (https://devblogs.microsoft.com/oldnewthing/20120307-00/?p=81...), almost 30 years ago, it's pretty hard to blame it on "the kids".
That doesn't sound like "those damn kids and their newfangled apps that can handle spaces, they'll never know the pain I went through!!!" and more like "I said something that was clearly not accurate and people pushed back".
The whole thing reeks of https://www.mouser.com/blog/Portals/11/mrb-singularity-f1.pn...
> I said something that was clearly not accurate and people pushed back
Is my life experience invalid? It's a needless error that i never want to deal with again. I cover my mouth when i cough, i use my turn signal when changing lanes, and i don't put whitespace in a filepath or a URL. It's that simple.
What i cannot fathom for the life of me, on HN of all places, is vehemently defending a practice that is not guaranteed to work 100% of the time: "Well iiiive never grazed an oven coil pulling a potroast out of the oven so obviously this guy is an idiot for advocating oven mitts." Give me a break dude.
It’s almost as if people in 2023 get by fine with spaces in their filenames whereas you seem to be stuck squarely in the 1980s.
I know, it’s a crazy idea. Those kids and their insanity. /s
You know, my main hobby is writing ASM for classic video game consoles, and my opinions and experience involves lots of janky / homemade / antiquated programs, so honestly you're not wrong =p
Feast your eyes on this abomination doubters:
C:\PROGRA~1\MICROS~4\2022\COMMUN~1\VC\Tools\MSVC\1435~1.322\bin\Hostx64\x64\ml64.exe
This is using a Microsoft development program with a Windows Environment Variable, mind you...