What it takes to pass a file path to a Windows API in C++
mastodon.gamedev.place
mastodon.gamedev.place
- Forget about long paths unless you really need to care. You need to care if you're accepting a file path from another app, for example. You don't need to care if you're using your own files in your install directory; too much of Windows doesn't support long files (including explorer, as mentioned, not to mention third party apps), so it's very unlikely that you'll need to deal with them if you're just handling your own data.
- Work in utf-16 when dealing with file paths. Convert back and forth at the point of entry / exit to / from your app. It's unsurprising that you need to call the conversion function twice in order to find out how much memory to allocate, that's just how that works.
Don't change random registry settings or code pages on the user's computer, that's mad.
But that's not compatible with other OSes, so you can't write native multiplatform C++ code just because windows keeps insisting on their worst-of-both-worlds UTF-16
(worst of both worlds because it's not single value per unicode character like UTF-32, and not ASCII compatible like UTF-8)
However, the common denominator nowadays is UTF-8, which has been a blessing overall getting rid of most of the aforementioned mess for international multi-user systems. And there is the C.UTF-8 locale which is slowly gaining traction.
I forgot what Japanese emulator I tried to run when I found all of this out, ut sufficed to say I didn't enjoy the experience.
Characters can consist of multiple code points.
I agree that utf-8 is preferable to utf-16 though in most cases where the root script is English/Latin.
I enjoy being "that guy", and so, to be unfathomably pedantic, I will point out that the Unicode standard actually uses the word "character" to refer to _encoded characters_, i.e. the mapping from a code point to an abstract character [0]. So 1 code point = 1 character for Unicode. Of course, in real life, character is a silly made-up word that means whatever you think it means and using it to refer to an extended grapheme cluster or glyph is probably closer to how people really think of them anyway.
[0]: See section 3.4 of https://www.unicode.org/versions/Unicode15.0.0/UnicodeStanda...
Unicode does not define "character" because that term is far too imprecise for a deterministic standard. And the Unicode standard is extremely clear that an abstract character may be represented by multiple code points.
"A single abstract character may also be represented by a sequence of code points—for example, latin capital letter g with acute may be represented by the sequence <U+0047 latin capital letter g, U+0301 combining acute accent>, rather than being mapped to a single code point."
> Of course, in real life, character is a silly made-up word that means whatever you think it means and using it to refer to an extended grapheme cluster or glyph is probably closer to how people really think of them anyway.
Human language and writing are the defining inventions of the known universe. Unicode just happens to be one way of representing them. Referring to them as "silly and made up" in comparison to Unicode is nothing short of extreme mental illness.
> Unless specified otherwise for clarity, in the text of the Unicode Standard the term character alone designates an encoded character.
Reading is free. I even gave you the section number.
> A single abstract character may also be represented by a sequence of code points
An "abstract character" is not the same thing as a character. The whole point of my comment was that the general word "character" is a general class of definitions but means nothing by itself. Much like the word "number" - you will notice that mathematicians define many kinds of numbers, but never the word itself.
> Human language and writing are the defining inventions of the known universe. Unicode just happens to be one way of representing them. Referring to them as "silly and made up" in comparison to Unicode is nothing short of extreme mental illness.
All words are made up. I'm sorry you had to find out this way.
It's so, so interesting that you believe this to be true.
Do you believe all things made by humans are made up? Is HN "made up"? Are other people "made up"?
C++ also supports wstring, but its implementation is dependent on the platform for god knows what reason.
utf-16 is also supported in C11:
https://en.wikipedia.org/wiki/C11_(C_standard_revision)#Chan...
> Improved Unicode support based on the C Unicode Technical Report ISO/IEC TR 19769:2004 (char16_t and char32_t types for storing UTF-16/UTF-32 encoded data, including conversion functions in <uchar.h> and the corresponding u and U string literal prefixes, as well as the u8 prefix for UTF-8 encoded literals).
You can use uchar.h for converting instead of the Win32 functions.
Here's a helper I wrote to wrap these string types in std::string and std::wstring:
bool mbStrToWideChar(string const& in, wstring& out)
{
bool status = true;
i32 bufferSize = MultiByteToWideChar(
CP_ACP, 0, in.c_str(), -1, nullptr, 0
);
WCHAR* buffer = new WCHAR[bufferSize];
i32 conversionResult = MultiByteToWideChar(
CP_ACP, 0,
in.c_str(), -1,
buffer, bufferSize
);
if (!conversionResult)
{
cerr << "Failed to convert ANSI string [" << in << "] to Wide string!" << endl;
status = false;
goto Delete;
}
out = wstring(buffer);
Delete:
if (buffer)
{
delete[] buffer;
}
return status;
}
and bool wideStrToMbStr(wstring const& in, string& out)
{
// Get buffer size
i32 bufferSize = WideCharToMultiByte(
CP_UTF8, 0,
in.c_str(), (i32)in.size(),
nullptr, 0,
nullptr, nullptr
);
// Set outstring to new buffer size
out = string(bufferSize, '\0');
//Convert
i32 res = WideCharToMultiByte(
CP_UTF8, 0,
in.c_str(), (i32)in.size(),
&out[0], (i32)out.size(),
nullptr, nullptr
);
if (!res)
{
wcerr << L"Failed to convert Wide string [" << in << L"] to ANSI string!" << endl;
return false;
}
return true;
}Platform differences unfortunately go way deeper than just using UTF-16. Lots of gnarly details. No platform is worse than another - they just have different kinds of good and bad in different places.
The first thing to realize to make it work is that paths are not strings, so you actually don't want to treat them like they are.
Edit to clarify: many systems which deal with non-path UTF-16 strings, nowadays are so intertwined with the UTF-8 world, that they get the validation "for free" so to speak, here and there, wherever they interact with the rest of the world.
Paths though, remain a common little corner case, working fine until it suddenly doesn't, for some weird combination.
A million times this.
Multiple Valve games set the microphone volume in Windows to 100% whenever you launch the game, which, for my setup, makes me sound unbearably loud. Unless your app has a valid reason to change my settings, like if I've set it up to do that on a schedule or with some condition, sure, otherwise don't mess with my settings.
When the USB audio device class was created [1], it had support for AVR-like audio controls (volume, balance, tone etc.), and, like would be reasonable, made all that work with decibels. But, as far as I know, that part of the standard got ignored by basically everyone, including some of whom wrote it (Microsoft), by presenting e.g. the on-device volume control as a 0-100 scale, like Windows and ALSA still do. Except... some hardware devices actually implemented the spec, including the PCM29xx family of audio codecs. These, and clones of them, have been used in quite a few USB audio cards and some very low-end audio interface over the decades. And for those, "100" in Windows means something like "+30 dBfs" [2] - so you'll get extreme distortion with basically any signal.
[1] 1.0 is actually impressively complex and basically models the entire stack of audio devices someone could have possibly had in their TV cabinet in the 1995, including stuff like dolby and surround decoders.
[2] This is defined around 5.2.2.2.3 and/or 5.2.2.4.3.2 in https://www.usb.org/sites/default/files/audio10.pdf
In particular, paths must be a different type than strings.
As a bonus: it becomes easier to implement niceties like path concatenation and realpath.
Applications might react differently to input events generated by these APIs. For example, in Windows 11's display settings, the position of screens can not be dragged if input events are coming from SetCursorPos, but works fine if using SendInput. Microsoft's own PowerToys uses a mixture of both under certain (complex) conditions, but I never found out the actual difference between them.
I was writing an application that sends mouse input from a Linux machine to a Windows one (similar to Synergy), and I originally received mouse movement events that are accelerated (with user or system-defined acceleration factors), and I found out that there are no Windows API (three of them) that accepts relative movements without further acceleration (i.e. Windows will always apply further acceleration, making the mouse hard to use). I ended up directly hook into evdev to get raw mouse movements and let Windows accelerate them.
[1] https://learn.microsoft.com/en-us/windows/win32/api/winuser/... [2] https://learn.microsoft.com/en-us/windows/win32/api/winuser/... [3] https://learn.microsoft.com/en-us/windows/win32/api/winuser/...
Synergy types of applications don't have that freedom because the host mouse cursor don't extend to the other devices display.
I don't think the difference between the two is all that strange. One sets the position of the cursor, the other interacts with the system like a normal mouse. The mouse and the cursor are separate things, and they're handled at different levels in the API stack, like XSendEvent and sending data to libinput.
https://www.kernel.org/doc/html/latest/input/event-codes.htm...
This is very straightforward (EV_REL) and requires a very small amount of code. There can be different problems to deal with when working at this level, but in my experience, everything works as expected with keyboards, mice, and gamepads.
Both approaches are reasonable and both are implemented in desktop operating systems for this reason.
Also, mouse_event is just a wrapper around what's basically SendInput(mouse)
::SetWindowTextW(widen(someStdString).c_str());
Implementation is straightforward, relying on `WideCharToMultiByte`/`MultiByteToWideChar` to do the conversion:If this cost really matters (and practically speaking it never does), then, as the other commenter said, the correct solution is to just use OS-native encoding for all file system paths and names used by the program, hidden behind an abstraction layer if needs be. UTF16 for Windows, UTF8 elsewhere.
The conversion overhead is really negligible: https://utf8everywhere.org/#faq.cvt.perf
(note: the two api calls per conversion is because how those specific functions work, first call to get the size to allocate, second to do the actual conversion, but you can always use another library in the implementation for the utf8<->utf16 conversion that might be more optimized than those windows api functions)
https://lemire.me/blog/2023/09/13/transcoding-unicode-string...
Someone is suggesting a way of making it less tedious, and your response is "performance?!" even though in both scenarios you're running the same code and it is likely the compiler in release would remove the intermediary.
No, that's not what I said.
I haven't tried it yet but with that you can just use the -A variants of the winapi with utf-8 strings. No conversion necessary.
Doing a string validation check in every single API call would waste cycles for no good reason.
Interpreting a string as encoding text in particular character set is, as far as the kernel is concerned, a problem for userspace.
Windows does make claims of "supporting" UTF-16.
Windows supports UTF-16, it doesn't guarantee UTF-16 correctness. The native methods annotated with the W suffix all take UTF-16, so unless you want to render your own fonts, you're going to need to provide it with either that or ASCII.
No language or environment that uses UTF-16 validates it, ever. Not a single one that I know of. But that doesn’t make it UCS-2, it just makes it potentially-ill-formed UTF-16.
Your points about normalization only matter in so far as Windows is generally setup to be case insensitive (which can be disabled).
All UTF strings face normalization issues for a whole host of reasons.
Wouldn't that break spectacularly? The whole "the fs is case insensitive" assumption has kind of been a given on Windows for about 40 years now.
I remember accidentally creating an NTFS flash drive with two names with difference cases in Linux. Windows programs sure didn't like that.
The "260 character limit" is a good one; mostly seen that one in "I can do anything to this file but delete it!" complaints.
Usually that's foot-shooting, but sometimes you can do that in HKCU as a low privileged user to a value that gets read by a higher privileged process and causes it to "misbehave".
Admittedly, it's (probably) harder to pull off than on Windows, but still...
> A file system race is the condition that occurs when multiple threads, processes, or computers interleave access and modification of the same object within a file system. Behavior is undefined if calls to functions provided by subclause [filesystems] introduce a file system race.
It's not just implementation-defined behavior, but full UB! You're utterly at the mercy of your implementation to do something reasonable when it encounters a TOCTOU issue, or, for that matter, any kind of concurrent modification to a file or directory. And C++ has a long history of implementations being unreliable in their behavior when UB is encountered.
Anyway, the problem described in the post (scanning the string twice when converting) is as race prone, or not, as the the filesystem API.
Sure, there are opportunities for races, but POSIX and the Windows API limit the possible outcomes of concurrent modification, and they also provide tools (fixed file handles, etc.) for well-written programs to prevent TOCTOU bugs. Meanwhile, the std::filesystem API just throws its hands in the air and says that filesystem races are UB, period, making it unusable if the program isn't in complete control of the directory tree it operates on.
> Anyway, the problem described in the post (scanning the string twice when converting) is as race prone, or not, as the the filesystem API.
I don't see what you mean? The typical scenario here is, you receive some absolute or relative path from an external source, encoded in ASCII or UTF-8 or some other non-UTF-16 encoding, and you want to operate on the file or directory at that path. Your internal path string is never going to change, only the filesystem can change. So there aren't any races from scanning your internal string twice to re-encode it; races can only come from using the re-encoded path in multiple calls to the filesystem API.
Unfortunately by default it doesn't, because the constructors which take `const char *`/`std::string` will default to the local codepage, instead of just using UTF-8. So you have to use `u8path` or C++20's `u8` string literals to get sane portable behavior.
So yes, you technically don't have to pay attention to the format, but as it is typical of many features in C++ the default is wrong, and you have to manually make sure you don't write code that's subtly broken.
(Or alternatively use a linter which would warn about this, but does such linter exist?)
[citation needed]
UTF-8 is used for storage and serialization, yes, but most mainstream programming languages store unicode strings in memory as UTF-16. The notable exception is Rust, which does indeed use UTF-8 even in memory.
Uh. What languages do you have in mind? C/C++ don't have any preference for UTF-16. Python3 doesn't use UTF-16. As you mentioned, Rust doesn't.
C/C++ don't have a preference but, for example, if you're writing for Windows, you'd probably want to use wide characters. I'm not sure what kind of strings Linux GUIs use but I suspect it's wide characters as well.
This is increasingly frequently false. For stupid historical reasons, far too many languages adopted potentially-ill-formed UTF-16 semantics (never well-formed, sadly), but this makes for such a bad representation that they’ve mostly abandoned it as a sole representation, in favour of more complicated arrangements. For example, Java strings can internally be single-byte Latin-1 or two-byte UTF-16. JavaScript engines do similar tricks these days, and although Python 3 just never did UTF-16 at all (it uses code point semantics, which is almost worse than UTF-16 semantics) it does a 1-byte/2-byte/4-byte hybrid representation thing too.
For one of the most extreme examples of becoming unglued from UTF-16, the Servo browser engine manages to use WTF-8 (UTF-8 plus lone surrogates, to allow representing ill-formed UTF-16) despite being forced by these historical reasons to expose UTF-16 semantics, and although the performance of some operations are harmed by it, others are improved, and in the balance it was a pretty convincing win when they tested it (which many people did not expect).
PyPy is also a great example of changing your encoding to one that has mismatched semantics: it applies UTF-8 (or I suppose it must actually be UTF-8 plus surrogates) to Python, and likewise they found it surprisingly good for performance.
These things give me hope that even environments like the JVM and .NET CLR might eventually manage to switch their internal string representations to UTF-8 or almost-UTF-8.
But also I emphasise that these are historical languages that are stuck with the horrendous 16-bit decisions of the early-to-mid-’90s. When you look at newer languages, they almost always choose something more sensible, and UTF-8 is almost always at the heart of it.
What's so bad about storing a unicode string as a series of codepoints? UTF-8 is also a series of codepoints in essence.
> When you look at newer languages, they almost always choose something more sensible, and UTF-8 is almost always at the heart of it.
I still think that UTF-8 for memory is a bad idea. For example, it wastes a lot of bits to make sure that a string can be decoded starting from an arbitrary byte offset. This is generally a desirable property for files and some networking applications, but it's an absolute waste of space when used in memory.
>>> '\udead'.encode('utf-8')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
UnicodeEncodeError: 'utf-8' codec can't encode character '\udead' in position 0: surrogates not allowed
>>> '\udead'.encode('utf-16')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
UnicodeEncodeError: 'utf-16' codec can't encode character '\udead' in position 0: surrogates not allowed
(UTF-8 is a sequence of 8-bit code units, representing a sequence of Unicode scalar values.)—⁂—
> For example, it wastes a lot of bits to make sure that a string can be decoded starting from an arbitrary byte offset.
If you ditch self-synchronisability and ASCII purity of extension by ditching the first two bits (“1-0”) on continuation bytes, you can encode some more briefly: roughly, U+0800–U+1FFF go from three-byte to two-byte, and U+10000–U+FFFFF go from four-byte to three-byte, leaving just U+100000–10FFFF as four-byte. But honestly the difference isn’t often that much (unlike UTF-16 → UTF-8, which roughly halves memory usage most of the time, commonly even saving when dealing with text that makes heavy use of other scripts that are more compact in UTF-16 than in UTF-8 since you commonly mix things in with ASCII markup), and if you’re really trying to shave memory, other techniques like compression are far more effective.
I love microoptimisation (and spent a large number of hours on my Casio GFX-9850GB PLUS, where having only 30KB of memory encouraged learning to shave single bytes here and there!), but there’s also a lot to be said for consistency. Yes, some purposes could have an encoding that is more efficient than UTF-8. But doing this complicates matters quite a bit, and I think we’re better off with just UTF-8.
Also I will note that the self-synchronising nature of UTF-8 is useful for some operations: you can scan through strings much more quickly that way. So it’s not pure waste as a memory representation.
Go also uses UTF-8 btw.
I think you're being needlessly pedantic (I know it's unusual for HN).
They even wrote a book about API design "Framework Design Guidelines: Conventions, Idioms, and Patterns for Reusable .Net Libraries"
I was actually secretly cheering for that project, because I thought it could be actually innovative
But, as I look back on that with my modern understanding of Microsoft's backward compatibility focus, I realize such a thing would never, ever fly
.WAD
.PAK
.PK3I've had a similar issue with a recursive folder and transferring it to another host in Linux. A megabyte on my machine filled up the server pretty quick.
There is no need to do this; you can just call the A variant by name. Wide char and ascii functions are distinguished with W or A at the end, such as CreateFileA or CreateFileW. CreateFile is a macro.
Additionally I have my own Windows header file which undefines all of the macros so my function names do not get mangled with A and W suffixes.
For example, imagine, that you wrote your Enterprise MS Tech Contoso Ltd.(R) authentication system, where user registers in UTF-8 on webpage, then on some layer it checks that user with the same name (in UTF-8 modern DB) does not exist, then it writes user basic information in UCS-2 encoded legacy "Enterprise DB". Aaand... Voila, Evil User overrides data of other user, because UTF-8 can't losslessly represent arbitrary sequences of 16-bit code units (should have used https://simonsapin.github.io/wtf-8/ instead, but your Enterprise tech was destroyed).
C#/.NET even inherit this misnomer and call UTF-16LE “Unicode”: System.Text.Encoding.UnicodeEncoding
People had to deal with local language file names since Windows 3.1 at least and it became very common with Windows 95. Good luck if you wanted to deal with files named in more than 1 non-ASCII language (very common in Europe or the Middle East).
Passing a string as a file name in C++ with macOS or with Linux in my experience was simple. The permitted length of using ASCII characters is about 4 times as long (may god have mercy on your soul).
I am not here to shit on Windows but the Windows devs clearly have a very different set of priorities (e.g. backwards compatibility) than the other breeds of modern OS devs.
I guess to a large extent we all expect to be in the browser (gross) in some number of years but Windows seems so much harder from the perspective of someone that has programmed for Unixen and studied Windows as an OS.
SQL Server (2017?) breaks if you update it on a UTF-8 Windows because it runs a TSQL that doesn't work with that code page. That script is a mess. Some of it is indented using tabs, some space. Trailing whitespace. Yuck
Much of modern linting and commit hooking is dedicated to checking whitespace placement, variable naming and function lengths but the well-formatted newly rewritten code is still buggy as hell - it just looks pretty
There have been numerous bugs caused by incorrect code formatting, most notably an SSL security bug from 2014: https://dwheeler.com/essays/apple-goto-fail.html
Another reason for formatting is the “minimal diff” paradigm. If a formatting rule would not be followed, in the next commit hitting this code, the format would also be affected, causing a larger diff than necessary.
There are other reasons for simple format linting, but the reasons above are the most profound.
Lastly, formatting is part of a range of static code analysis tools. Generally, formatting inconsistencies are the easiest to detect and resolve, as opposed to more sophisticated tools.
It is kind of like "pattern matching" on the error patterns.
The inverse may not have the same correlation, since it can be automated.
You can write C# code dealing with reading/writing files once and compile it on Linux/Windows/Mac and it'll work pretty much the exact same.
First of all, games are apps, second even if apps unit keeps mostly ignoring WinUI/UWP (written in C++), whatever they do with Web widgets is mostly backed by C++ code, not C#.
On of the reasons why VSCode is mostly usable despite being Electron, is exactly the amount of external processes written in C++.
Applications being written in .NET is mostly on the Azure side.
According to wikipedia:
https://en.wikipedia.org/wiki/Universal_Coded_Character_Set
> The UCS has over 1.1 million possible code points available for use/allocation, but only the first 65,536, which is the Basic Multilingual Plane (BMP), had entered into common use before 2000. This situation began changing when the People's Republic of China (PRC) ruled in 2006 that all software sold in its jurisdiction would have to support GB 18030. This required software intended for sale in the PRC to move beyond the BMP.
You are of course, wrong about this. Most .Net/C# code is not Azure (yet anyway) -related; it is the billions of lines of enterprise application code across businesses around the world (for me, since 2001)…
In RAM-constrained world of the past, you would stack-allocate `char buff[MAX_PATH]` and do all your strcpy/strspn in there with no problems.
Now, if that app receives a long path into a too short buffer, it will instantly stack overflow and may cause exploitable problems.
``` wchar_t filename[MAX_PATH]; CreateFileW(...) ```
in both first part and third party Windows code, often in deep callstacks passing file names around. Changing the length requires fixing them all.
See comments in https://archives.miloush.net/michkap/archive/2006/12/13/1275...
But for functions outside of Windows itself, this is the exact reason why the long path feature is hidden behind an opt-in flag.
See for example https://learn.microsoft.com/en-us/windows/win32/api/fileapi/..., which also shows how they gradually made the input argument more flexible:
- “By default, the name is limited to MAX_PATH characters. To extend this limit to 32,767 wide characters, prepend "\\?\" to the path”
- “Starting with Windows 10, Version 1607, you can opt-in to remove the MAX_PATH limitation without prepending "\\?\"”
I also guess there’s lots of code that sees those paths (anti-virus software, device drivers)
Passing a string as a file name in C++ with macOS or with Linux in
my experience was simple. The permitted length of using ASCII characters
is about 4 times as long (may god have mercy on your soul).
Macs are simple enough too if you ignore the quirks. HFS (which was never seen on a modern MacOS) usually stores no information about what encoding was used for filenames. It's entirely dependent upon how the OS was configured when the file was named (although some code I've seen suggests that something in System 7 would save encoding info in the finderinfo blobs). So non-latin stuff gets mangled pretty easily if you're not careful. Filenames are pretty short (32 bytes) minus the one byte because (except for the volume name) they're Pascal strings with the length at the front.HFS+ (which is what you'll find on OSX volumes) uses UTF-16 but then mandates its own quirky normalization and either Unicode 2.1 or 3.2 decomposition depending… which can create headaches because most HFS+ volumes are case-insensitive. It's been so long since I've touched anything Cocoa, but I assume the file APIs will do the UTF-16 dance for you and the POSIX stuff is obviously OK with ASCII.
And, of course, let's not forget the heavily leveraged resource forks. Of course NTFS has forks but nobody seems to use them.
APFS standardized on Unicode 9 w/ UTF-8.
CDs? Microsoft's long filenames (Joliet) use big endian UTF-16 (via ISO escape sequences that theoretically could be used to offer UTF-8 support). Which sounds crazy until you realize their relative simplicity (a duplicate directory structure) compared to the alternative Rockridge extensions which store LFNs in the file's metadata with no defined or enforced encoding. UDF? Yeah that's more or less UTF-16 as well.
I think we're perhaps forgetting just how young UTF-8 is.
It strikes me how developer ergonomics have improved as computers have become cheaper/increased in power.
As to UTF-8, we may say it’s young but in 14 months it will be old enough to purchase and consume alcohol in the United States. From other comments it seems like Microsoft don’t think the tech debt is too great so long as they have good libraries in C#
Second, I don't see a risk in setting `LongPathsEnabled` on a machine as it's only 1 piece of the puzzle. You still need to opt-in at the application level:
https://learn.microsoft.com/en-us/windows/win32/fileio/maxim...
The longpath thing shouldn't be something you have to deal with each time you call an API. You should do what Git for windows and other software do which is enable long paths (assuming you actually need it) during setup/install.
>This document also recommends choosing UTF-8 for internal string representation in Windows applications, despite the fact that this standard is less popular there, both due to historical reasons and the lack of native UTF-8 support by the API. We believe that, even on this platform, the following arguments outweigh the lack of native support.
Most computers run Windows, by a rather big margin. Even more computers use Javascript. Are you running KDE? Don't be surprised if half the text in your screen is secretly UTF-16!
UTF-8 is definitely better, but if you're on Windows, you're making your own life harder by using UTF-8. Microsoft is migrating APIs to UTF-8 but it'll take years before that's finished. Just look at all the conversion code you need according to the page you linked.
Now, if you're writing a cross platform library, you'll have to decide what encoding(s) you use. I personally don't really see why you wouldn't make your library encoding agnostic, but if you want to stick to one single encoding, you'll have to decide.
If you're writing a program for Windows, technical superiority is a mediocre argument for making your own life that much harder. There are tons of formats and design decisions that may be technologically superior, but just end up wasting programmers' time, and this is one of them.
After a day or so of testing and trying it with various different SDKs across different platforms I settled on utf8. did everything needed, compatible with ASCII, all the functionality of 16 none of the _absolute insanity_.
looking back, these kind of issues are why its now some 20 years since I wrote any windows software (rather than porting linux/mac stuff or just using cross platform VMs like Javascript, Lua or Java).
MFC..... Oh my.... I actually feel sorry for the poor souls that still have that in their life.
[0]: https://en.wikipedia.org/wiki/Time-of-check_to_time-of-use
For any shipped game being played on a player's machine the data files should be some form of packed files, which are close to the executable and only a handful need to be opened.
All this can of course be fixed with another layer of indirection. But if you ever wanted to know why Windows file APIs and filesystems are so dog-slow, here is (part of) your answer.
[0] e.g. there are lower level APIs that can create and access files and filenames that higher level APIs will filter out or disallow.
Back in my day we just walked into the office of the man in charge and shook hands, we struck deals like real men. If we ask ask the NT kernel politely in a language it understands, we will get our filepath. Right?
http://undocumented.ntinternals.net/index.html?page=UserMode...
Describes Windows Forms vs everything else MS has been trying to do in the past 15 years.
Casey has also spoken out about the horrible state of affairs in Microsoft in the past https://news.ycombinator.com/item?id=27728177 so I'm sure it's not a one off phenomena.
Allegedly.
MaxPath is real though. Exceeding it not recommended.
So having a file system and other APIs support Unicode natively made a lot of sense at the time; it simplifies many things including deploying into organizations that have global users. This is now over 30 years ago. And the A vs the W suffixed APIs were not simple ascii; they were multi-byte, so to work with strings you still had to know the code page. Obviously UTF-8 is much simpler for that now, but that wasn’t the world when this was designed.