Some sanity for C and C++ development on Windows
nullprogram.com
nullprogram.com
But that's not what the article is about. What the article is about is that the C runtime shim that ships with Visual Studio defaults to using the ANSI API calls without supporting UTF-8, goes on to identify this as "almost certainly political, originally motivated by vendor lock-in" (which it's transparently not), and then talks how Windows makes it impossible to port Unix programs without doing something special.
I half empathize. I'd empathize more if (as the author notes) you couldn't just use MinGW for ports, which has the benefit that you can just use GCC the whole way down and not deal with VC++ differences, but I get that, when porting very small console programs from Unix, this can be annoying. But when it comes to VC++, the accusations of incompetence and whatnot are just odd to me. Microsoft robustly caters to backwards compatibility. This is why the app binaries I wrote for Windows 95 still run on my 2018 laptop. There are heavy trade-offs with that approach which in general have been endlessly debated, one of which is definitely how encodings work, but they're trade-offs. (Just like how Windows won't allow you to delete or move open files by default, which on the one hand often necessitates rebooting on upgrades, and on the other hand avoids entire classes of security issues that the Unix approach has.)
But on the proprietary interface discussion that comes up multiple times in this article? Windows supports file system transactions, supports opting in to a file being accessed by multiple processes rather than advisory opt-out, has different ideas on what's a valid filename than *nix, supports multiple data streams per file, has an entirely different permission model based around ACLs, etc., and that's to say nothing of how the Windows Console is a fundamentally different beast than a terminal. Of course those need APIs different from the Unix-centric C runtime, and it's entirely reasonable that you might need to look at them if you're targeting Windows.
we certainly have a different opinion on what "great" means. It takes less time to rebuild my whole toolchain from scratch on Linux (~15 minutes) than it takes to MSVC to download those friggin debug symbols it seems to require whenever I have to debug something (I sometimes have to wait 30-40 minutes and I'm on friggin 2GB fiber ! and that's seemingly every time I have to do something with that wretched MSVC !)
Thankfully now the clang / lld / libc++ / lldb ... toolchain works pretty well on Windows and allows a lot more sanity but still, it's pretty slow compared to Linux.
If something absolutely requires MSVC ABI compatibility, I use "clang-cl".
God bless LLVM developers.
If you suffer through this, you have an exe that works on your version of Ubuntu, maybe on other versions of Ubuntu or possibly other Debian-based distros. If you want it to also work on Fedora, it's back to tinkering.
Tbh i think the only sane-ish way of building to dockerize your build env with baked-in versions.
In contrast, you pick and SDK and compiler version for Windows, and as long as you install those versions, things will work.
Granted, writing an RPM is a special kind of hell, but at least you don’t have to package everything with your program. But actually you can still do that - I’ve done that plenty of times in embedded. You can always ship your program with its dependant libraries the way you always have to on Windows. In fact it’s a lot easier because most 3rd party libraries were originally coded on Linux and build more sanely on Linux. And RPATHs are pretty easy to figure out.
Linux gives you options.
IMHO, C++ dependency management kinda stinks, regardless of platform.
Indeed, and the fact that it's platform-specific in the first place certainly doesn't help!
> you have an exe that works on your version of Ubuntu
This is indeed a downside of the Linux approach, it's the price we pay for the significant flexibility that distros have, and the consequent differences between them. Windows has remarkably good binary compatibility, but it's a huge engineering burden.
> Tbh i think the only sane-ish way of building to dockerize your build env with baked-in versions.
This is an option, but bundling an entire userland for every application has downsides that the various Linux package-management systems aim to avoid: wasted storage, wasted memory, and less robust protection against inadvertently using insecure unpatched dependencies.
GTK is known to break their ABI across major versions (GTK1->GTK2, GTK2->GTK3, GTK3->GTK4) but as a C ABI it should be compatible between minor versions and everything can be assumed to have GTK2 and GTK3 available anyway. X11 as a protocol has always been backwards compatible and Xlib on Linux pretty much never broke the ABI since the 90s. Here is a screenshot with a toolkit i'm working on now and then running the exact same binary (built with dynamic linking - ie. it uses the system's C and X libraries) in 1997 Red Hat in a VM and 2018 Debian (i took that shot some time ago - btw the brown colors is because the VM runs in 4bit VGA mode and i haven't implemented colormap use - it also looks weird in modern X if you run your server at 30bpp/10bpc mode)[0].
Of course that doesn't mean other libraries wont be broken and what you need to do (at least the easy way to do it) is to build on the oldest version of Linux you plan on supporting so that any references are to those versions (there are ways around that), but you can stick with libraries that do not break their ABI. You can use ABI Laboratory's tracker to check that[1]. For example notice how the 3.x branch of Gtk+ was always compatible[2] (there is only a minor change marked as breaking from 3.4.4 to 3.6.0[3] but if you check the actual report it is because two internal functions - that you shouldn't have been using anyway - were removed).
[0] https://i.imgur.com/YxGNB7h.png
[1] https://abi-laboratory.pro/index.php?view=abi-tracker
[2] https://abi-laboratory.pro/index.php?view=timeline&l=gtk%2B
[3] https://abi-laboratory.pro/index.php?view=objects_report&l=g...
Off topic, but I wonder if anyone knows whether it's possible to use rustc with lld, instead of link.exe? I tend not to have Visual C++ on my home systems. Is it as simple as the Cargo equivalent of LD=lld-link.exe?
Fun fact: you can do the same in gdb and some distributions (e.g. openSUSE) have it enabled by default. Though you also get the source code too.
I was messing around with DRM/KMS the other day and had some weird issue with an error code, so i placed a breakpoint right before the call - gdb downloaded libdrm and libgbm source code (as well as some other stuff) and let me trace the call right into their code, which was super useful to figure out what was going on (and find a tiny bug in libgbm, which i patched and reported).
NixOS has had a similar thing for a while called "DwarfFS" where a FUSE filesystem instead resolves filenames back to the package that needs to be installed, which was around for a while before debuginfod, but very NixOS specific. I'm happy this is now so much more widely available as of recently.
I think you can just unclick the radio button in debug settings that requires that?
I wonder what these might be. You mean potential race conditions regarding file operations?
In summary, I can't recommend you continue using your Linux machine after an update without a restart, since you are open to an entire category of weird bugs.
and the Chrome thing sounds rather strange as I thought Chrome forked processes from a "zygote" (a prototype process) rather than re-exec'ing() the binary (which should retain the handle to the deleted library inodes)
not to mention the shared library naming scheme should prevent this sort of incompatible change from occurring
I hate the fucking "file is busy, try again?" dialog so fucking much.
This includes bugs too. I ran into an undocumented bug in select(1) IIRC that they couldn't fix since it would break backwards compatibility. I spent like a day trying to figure out why my program wouldn't work correctly on Windows.
Yes, and almost entirely in ways that are bad.
I think Microsoft have partially recognised that they're tied to compatibility with a set of choices that have lost the popularity wars and now look wrong. That's why they've produced the two different sorts of WSL, each of which has awkward tradeoffs of its own. And Windows Terminal to replace the console. But eventually I think they may be forced to:
- drop \ for /
- switch CRLF to LF as the default
- provide a pty interface
- provide a C environment that uses UTF-8 by default
It's been weird working with dotnet core and seeing the "cross platform, open source" side of Microsoft, who develop in a totally different style. It's like watching a new ecosystem being built in the ruins of the old.
API calls accept / as path separator (and interpret it correctly). Shell is a different beast though.
So you would also need to rewrite all command line utilities to use something like `ipconfig --all` instead of `ipconfig /all`.
MinGW GCC doesn't make a difference to the article. It uses the same C runtime library as VC++ and has the same problems with defaulting to ANSI codepages and text-mode streams. In fact, as far as I know, there is no native open-source alternative to the VC++ runtime that isn't a full alternative programming environment like Cygwin.
C++ is bearable (except the bloated piece of crap that is Visual Studio). but C is almost nonexistent. For many years their compiler lagged in standardized C features (this only somewhat improved recently) and you cannot use vast majority of system APIs, which use COM interfaces.
It gets worse with that old code if you try to share modules between windows and Linux applications.
Additional complications come from trying to support TCHAR to allow either type of char for libraries.
Anyway, I have ended up supporting monstrosities of wstring, string, CString, char *, TCHAR mushed together, constantly marshalled converted back and forth.
And more: https://docs.microsoft.com/en-us/cpp/text/how-to-convert-bet...
The reason for TCHAR is not "libs/tools/editors" not supporting Unicode, but the operating system itself. With TCHAR and related types, the same source code could target both Windows 95 and Windows NT, you just have to change a single #define (ok, IIRC there are actually three #defines: UNICODE, _UNICODE, and another one I can't recall at the moment) and recompile.
So, you might as well be insane? :D
I believe that the problem is about the definition of "elsewhere". The elsewhere is called POSIX. The portability is based on POSIX. So the article can be summed as "Windows is not POSIX compliant".
I'd like to add some historical background the UTF-8 problem mentioned there. UTF-8 was standardized in 1992. But the first Linux using UTF-8 as the default encoding was RedHat in 2002 [1]. 10 years later!
On the other hand, Windows started using Unicode before it was standardized, and back then there was no UTF-8. Or as Raymond Chen put it, "Windows adopted Unicode before the C language did" [2].
[1] https://www.cl.cam.ac.uk/~mgk25/unicode.html#linux
[2] https://devblogs.microsoft.com/oldnewthing/20190830-00/?p=10...
It's also worth noting that dealing with UTF-8 correctly is more than just declaring "the encoding of this set of characters is UTF-8"--it does require you rethink how string APIs are designed to work well. If the alternative is a 16-bit fixed-with character format, UTF-8's variable-width format doesn't necessarily look like a wiser idea. It's not until 1996, when Unicode moves from 16-bits to 25-bits, that UTF-8 actually looks like a good thing.
Probably not. The author of this thread's article has incorrect history about UTF-8 being widely accepted in 1993. At least 4 different systems independently chose UCS2/UCS4 instead of UTF-8 in the early 1990s:
- Microsoft Windows NT 3.1
- Sun Java in 1995
- Netscape Javascript in 1995
- Python 2.x
Why? Because before 1996, the Unicode Consortium initially thought 16-bit 65k characters was "more than enough" based on the CJK unification first recommended by language experts from Asia. The wikipedia page on Unicode revision history shows the explosion of extra CJK chars in Unicode 3.1 which exceeded count of 65k didn't happen until 2001: https://en.wikipedia.org/wiki/Unicode#Versions
The prevailing standard in 1993 was UCS-2 and not UTF-8. Considering that a huge revision of an operating system like Windows is a multi-year effort, they were really looking at what the standard was circa ~1991. So Rob Pike's napkin idea of UTF-8 in 1992 and then the later presentation at USENIX in January 1993 -- is really not relevant for a July 1993 release of Windows NT 3.1. The author writing in Dec 2021 has hindsight bias but MS/Java/Python just went with Unicode 1.x UCS2 which was logical at that time.
And then Microsoft spent the next 20 years explaining/complaining to everyone (and this includes Chen) why they couldn’t make UTF-8 a code page because it would mean revalidating all of those -A functions for 4-byte characters when they were originally intended for only 3-bytes?
This architecture made it possible to claim posix compatibility on paper, while making it clear microsoft would make your life hell if you tried to actually use it
https://en.m.wikipedia.org/wiki/Windows_Services_for_UNIX
However most people didn't care enough to use them.
Having WSL now in the box to do what many of us have been doing with VMware, is more an attack on Apple to cater to those that do not pay Linux OEMs to develop for Linux than anything else out of their good hearts.
But for non-trivial projects that's not much of a problem because it's quite trivial to ignore the C stdlib functions and call directly into OS APIs (via your own or 3rd-party cross-platform libraries).
I learned this approach from Handmade Here / Handmade Community and it's probably the deepest architectural lesson I learned anywhere, while it seems to be not widely known. I'm almost tempted to say that the conventional approach is wrong and misguided. However, an issue with the platform-calls-code approach is that the code usually can only communicate asynchronously with the platform (one of the issues with code-calls-platform is the one-shot synchronous nature of function calls). That might not work always, for example, I wanted RDTSC counters or Copy-Paste functionality, and I still do this by finding a platform abstraction that the code can call into directly.
I/O would be the major no-no.
But it gradually became clear that UTF8 predominates.
D best practice now is to do all processing in UTF8. When encountering UTF16 or UTF32, promptly convert it to UTF8, do the processing, and then convert to UTF16 or UTF32 on output. In practice, the only time this is needed is to interface with Windows.
All the D standard library code that interfaces with Windows converts to/from UTF8 to UTF16.
If I was doing a do-over for D, there wouldn't be core language support for UTF16 or UTF32, only library support.
I was accused of being an anglophile at one point for not doing everything in UTF-32 :-)
I do see how a native application generally behaves and performs better, but it's a whole lot of extra work vs. a less-than-native application. I'd even opt for a generic backend (say, Rust or Go based) and simply having the 'interface' be OS-native.
Factory automation at TSMC and the likes is mostly just Java, web and webbased interfaces all in a closed network. Same for most modern logistics factory automations.
Classic PLC-based HMI interfaces from Siemens and the likes do still come with Windows as a dependency, but those have normal HTTP interfaces as well. Besides, this is such a niche with so much money going around they hardly depend on the ability to work inside of Microsofts limits. Calls into private APIs galore, hence the limits on OS patches and upgrades that might break those undocumented interfaces...
The same attitude as yours has spawned an entire culture of repulsive inefficiency and egregious waste. Hundreds or even gigabytes of memory for a bloody chat client which still manages to lag noticeably on 2019-era hardware? How is that even acceptable?
...all because of some misguided and incredibly selfish notion that "developer experience" somehow trumps that of however many users there are (and because a lot of this is corporate shit, we have no choice but are forced to use it.)
https://www.qt.io/blog/2018/05/24/porting-from-qt-1-0
I'd do the same today, definitely !
The best GUI development experience I had on Windows aside from that was WTL/ATL. I kept hoping for a long time that this framework would reemerge as a first class GUI platform on Windows. I think Microsoft missed a tremendous opportunity by neglecting it. It actually made writing COM components enjoyable and appealing.
> "Until recently, Windows has emphasized "Unicode" -W variants over -A APIs. However, recent releases have used the ANSI code page and -A APIs as a means to introduce UTF-8 support to apps. If the ANSI code page is configured for UTF-8, -A APIs operate in UTF-8. This model has the benefit of supporting existing code built with -A APIs without any code changes."
https://docs.microsoft.com/en-us/windows/apps/design/globali...It seems not everyone is aware of this, too
But it does? In "How to mostly fix Unicode support" section.
Ah, but Cygwin doesn't actually change so much that we can't scale some of it back, to have sane C and C++ development on Windows.
https://www.kylheku.com/cygnal/
Build a Cygwin program, then do a switcheroo on the cygwin1.dll to deploy as a Windows program.
I developed Cygnal to have an easy and sane way of porting the TXR language to Windows.
You can control the Windows console with <termios.h> and ANSI/VT100 escapes. Literally not a line of code has to change from Linux!
In Cygnal, I scaled back various POSIXy things in Cygwin. For instance:
- normal paths with drive letter names work.
- the chdir() function understands the DOS/Windows concept of a logged drive: that chdir("D:") will change to the current directory remembered for the D drive.
- drive relative paths like "d:doc.txt" work (using the working directory associated with the D drive).
- PATH is semicolon separated, not colon.
- The HOME variable isn't /home/username, but is rewritten from the value of USERPROFILE. (I had problems with native Windows Vim launched from a Cygwin program, and then looking for a nonexistent /home due to the frobbed variable, so I was sure to address this issue).
- popen and system do not look for /bin/sh. They use cmd.exe!
- when exec detects that "/bin/sh" "-c" "<arg>" is being spawned, it rewrites this to cmd.exe. The benefit is that interpreter programs which expose command invocation functions which they implement from scratch by forking and execing /bin/sh -c command are thereby retargetted to windows.
- the /cygdrive directory is gone
- /dev is available as dev:/ so you can access devices.
Your SSL certificate is invalid and your server uses TLS 1.0 which is insecure and disabled by default in recent browsers.
That is correct. Basically, the OS upgrades kept breaking my highly customized setup, resulting in downtime and wasted time, so eventually I dropped out of the program. I need to rebuild a new server and migrate everything over. With two small kids now and full-time job, if I have any block of free time (basically late at night or early morning), I prioritize putting it into working on code. It's not just a web server setup and custom things around it, but Postgress databases, e-mail setup and other stuff. Every item could have some issue chewing up hours.
> which is insecure
Secure/insecure isn't a simple Boolean. It's less secure than newer versions of the protocol, but it's not "insecure" like, say, a plain password sent over telnet.
> Your SSL certificate is invalid
I don't believe that is the case. Firstly, I've never had an problem connecting to the site using https from any browser. Secondly, most online SSL checking tools do not report a problem other than remarking on it being self-signed.
For instance geocerts.com says "The hostname (www.kylheku.com) matches the certificate and the certificate is valid."
I found one which doesn't understand wildcard names: it reports that www.kylheku.com doesn't match *.kylheku.com. That's a bug in the checker. One other flags some unspecified issue with the Alternative Names property (which I didn't even specify when generating the cert).
Embrace C++ Builder, Visual C++ or Qt.
No sanity medicine required.
I also don't see how using VC++ for example circumvents API issues, since it's just a compiler. I can imagine that Qt might solve some issues, since it offers a very wide set of abstractions. As such, it can probably do things like convert string formats on-the-fly between getting it from the OS and passing it to your application.
People here are saying don't expect it to be Unix, and that's fine. When I write for Windows, everything is PWSTR and that's fine.
But I have seen a lot of people get tripped up on this. People who are not well versed in the issue have no idea that using fopen() or CreateFileA() will make their program unable to access the full range of possible filenames, varying with the language setting. I've absolutely seen bugs arise with this in the 21st century.
The old "ACP" is meant for compatibility with pre-unicode software, but that hasn't been a big concern for something like 20 years. It is a much bigger issue that people are unaware of this history and expect char* functions to just work.
Using API functions ending with ...A (for the ANSI version) like CreateFileA is actually rather lowlevel style and means that you know what you do (the same holds for the ...W versions) and explicitly want it this way.
Normally, you use CreateFile and Visual C++ by default sets the proper preprocessor symbols such that CreateFileW will be used. If you pass a char*, you will get compile errors.
In other words: the infrastructure is there to avoid this kind of errors - and it is even used by default. But if you make it explicit in your code that the ANSI version of an API function (CreateFileA) should be used instead of using CreateFile, don't complain that the compiler does obey and this turns out to be a bad idea.
The macros in question were always a bad idea, because they applied to all identifiers. So e.g. if you had a method named CreateFile in your class somewhere, and somebody did #include <windows.h> along with the header defining that class, you'd get a linker error.
Why do you think it's immaterial, though? It was back when A-funcs were considered legacy, and W-funcs were the way to go. But now the official stance is that A-funcs are for use with UTF-8. Being explicit about whether you're handling UTF-8 or UTF-16 strikes me like a good idea.
But I think of the canonical name as not having either A or W. Rather what you get for TCHAR is a per compilation unit setting, and the A or W in the linker name is sort of an implementation detail. That is to say I'd internalized the world the macros presented, and labelled that as "what makes sense". In particular it is distracting from the semantics of the APIs to have the string type show up in the name.
If they could do c++ style name mangling in plain C I feel like they might've and it'd be somewhat more "pure". The macro thing feels like sort of a hack to achieve this.
https://developercommunity.visualstudio.com/t/allow-redistri...
Also considered Rust, but that is so bloated/slow! Can't stop anyone from using Rust for the game in the future though, because any .so/.dll will be able to hot-deploy.
https://stackoverflow.com/questions/65807034/clang-llvm-zip-...
https://github.com/ossia/score/releases/tag/v3.0.0-rc7
you just need to build llvm/clang statically and target_link_libraries(<couple stuff>) in cmake (and ship the headers if you want to do useful things, this actually takes much more space uncompressed but it'd be the same whatever the compiler)
I'm surprised this is so hard to find, why is there no official redistributable compiler for Windows?! What is Microsoft so afraid of?
This is the only disadvantage Windows has compared to Linux!
I guess you can make this alot smaller if you remove the cross compiling parts?
struct std_string {
union {
char small[16];
struct { size_t len; char* buf; } big;
} value;
unsigned char is_big;
};
There is nothing it does regarding support for non-ASCII characters over what you get from buf and len. And for UTF-8, you don't even need len, plain old strlen from the 1980s works fine on valid UTF-8, so plain char* from C works just as well.And the union is just a performance optimization for small strings (is_big is probably not an additional field in good impls, but I separated it here), it's logically identical to just the buf and len.
Which is all just to say, go for it with tcc!
I don't even know what union does, but I can imagine you have many smaller chunks?
So ignoring that irrelevant optimization, std::string is basically:
struct std_string {
size_t len;
char* buf;
};
For storing valid UTF-8, the len is unnecessary, since a NUL byte is not valid UTF-8. You can still tell how many bytes of UTF-8 you have by using strlen, because when you find a 0 byte, it is never part of the string, it's always the NUL terminator. So the len is not strictly necessary for valid UTF-8 -- leaving us with just char*.And not trying to dodge your question about UTF-8 logic, but my point was you can dodge that whole question, because std::string provides the same amount of UTF-8 logic as char* -- that is, none at all. If you've been getting by on std::string, then you can get by on char*. If you only need to support UTF-8 input and output, and you don't need to manipulate strings (replace characters, truncate them, normalize them for use as keys in a data structure, etc.) or only need to do simple substring searches for ASCII characters, then you can just use char* or std::string. UTF-8 has a great design, which was consciously chosen to make all of that possible.
(I'm mostly thinking about Rust, which I think has to do some contortions on Windows. Probably also Go, to the extent it uses native APIs; arguably Python too.)
If the standard library dropped this requirement and only supported valid Unicode then it could simply use normal UTF-8 instead of WTF-8.
Long-term, though? Absolutely.
In general I think it’s best to stick with the native "W" functions and/or wchar_t on Windows.