Brian's Rules for Writing Cross Platform 'C' Code (2008)
ski-epic.com
ski-epic.com
Purely on the coding side of things, I would caution against extremism when fighting code duplication between platforms. There is much to be said for having a single C file per platform that is self contained, even if it entails a certain level of duplication. The alternative very easily leads to an increasingly bifurcated decision tree that tries to account for similarities and differences between platforms (autoconf syndrome).
Some things changed. Qt is pretty solid today. We use it as native Linux platform. We use native Windows and Mac front ends too.
Using Qt solves lots of issues very easy, like OpenGL drawing capabilities.
So in Windows and Mac we have two versions of our own software, one with native widgets, one with Qt ones. The Windows version using Qt is much faster and efficient, albeit a little ugly, because of OpenGL issues(instead of official DirectX) and font sizes or colors a little different from Windows and so on.
He found it was very good for cross platform work.
If you're just writing a utility of some sort then sure, you probably don't need the kind of overhead that Qt dumps on you. That said, I've written very performant code using Qt's threading and concurrency libraries and it t was surprisingly easy. Think industrial image processing that needs to be done at 150Hz on two camera streams simultaneously.
It's very, very hard to make any generalization about programming, and so I suspect that in his confident pronouncement that cross platform should never add more than a small amount of time to a project, there is something he's ignoring. The application they sell is for backup, and so it would probably be close to ideal for making cross-platform. I think he underestimates how complex GUI code can get when dealing with requirements less trivial than a backup application.
He also basically ignored Macintosh and ObjectiveC. I don't recall what the state of Macintosh was at the time this piece was written, but I'm pretty sure ObjectiveC was the standard language on Mac then as it is now.
I think there are sage principles here, though.
UCS-2 made a lot of sense back in the day when Unicode was supposed to be 16-bit and you'd just have one 16-bit "char" for each Unicode code point. Bliss! Then Unicode breached the BMP and this didn't work anymore. UTF-16 is just a way to let all that old code that assumes 16-bit Unicode to keep working. You get all the disadvantages of UTF-8 (no 1:1 correspondence between code units and code points) and none of the advantages (widespread compatibility with non-Unicode-aware tools, mostly-reliable discrimination between actual text and random binary data that might look like text).
One can kind of make a case for UTF-32 in some cases, but UTF-16 is almost never the right choice unless there is literally no other.
This is why we do not use wchar_t as a cross platform type. We like "char *" with NULL terminated Utf8 strings, which works really well on all platforms. Unfortunately on Windows, some of the core API still take Utf16 (2 bytes per char). Since it is a Windows call it must be in an #ifdef anyway, and so IMMEDIATELY BEFORE calling into the Windows call at the last possible moment we convert the string from Utf8 to Utf16. It sounds like a pain, but after you have wrapped all your OS specific calls (like File I/O) you're calling the higher level anyway, and it was completely worth it.
It sounds like you have it figured out, but just in case: UCS-2 is the original pure 16-bit encoding from the days when nobody could conceive of needing more than 65536 characters for anything. UTF-16 is the hacked up "oops, it turns out there are a lot more Chinese characters than we realized" version that tries to maintain some compatibility with software built for UCS-2.
Given that there was only a five-year period between the first release of Unicode and the introduction of UTF-16, it's kind of amazing how pervasive the 16-bit encoding got. It's thoroughly baked into Java, where "char" was supposed to be a unicode character, but ended up being a unicode character or a piece of a UTF-16 surrogate pair. Windows APIs and Cocoa's NSString went through the same process. JavaScript too, and the spec still allows implementations to only support UCS-2.
> He also basically ignored Macintosh and ObjectiveC
At Backblaze (my company) we write all lower level common libraries in 'C' and C++ so like a cross platform HashTable, and file I/O functions. It turns out ObjectiveC on the Macintosh is a PERFECT SUPERSET of 'C' and C++ so luckily all the Objective C code can call into our OpenSSL for encryption and can call into our base libraries for File I/O.
Translating DWORD to `unsigned long` is not correct! unsigned long is 32 bits on x86_64 Linux, but 64 bits on x86_64 Windows. Using long is essentially never correct for cross-platform code.
The right thing to do depends on what the value represents. If you want a 32 bit unsigned integer, use uint32_t. If you want something representing a pointer or size, use the corresponding type (uintptr_t, size_t). This doesn't just make porting easier, it adds useful hints for anyone reading your code.
"The most important place to apply this rule is in cross platform library code. You can't even call into a function that takes a DWORD argument on a Macintosh, so the argument should be passed in as an unsigned long."
If DWORD is declared, then your Mac code can use it fine. If it's not, then it won't even compile. At no point can you ever reach a point where "you can't even call into a function that takes a DWORD argument".
> If DWORD is declared, then your Mac code can use it fine. > If it's not, then it won't even compile.
Correct -> as a WINDOWS programmer you think it's fine because it compiles just fine. The point is it then won't compile on the Mac, or on Linux, or any other platform. The whole point is if you want a 32 bit integer, use a cross platform 32 bit integer from the beginning. When a Macintosh programmer goes to compile, it will compile. ONLY IF you need to call a Windows specific call, then type cast or convert into the Windows specific type, and then you do it inside an #ifdef. But most all of the time Windows programmers seem to love their DWORDS and they use them all over the place, even for public APIs that some other poor Macintosh programmer has to replace later. Why not start with the cross platform version? What does a DWORD get you that "int32_t" does not? (NOTE: we make sure "int32_t" is defined and a 32 bit integer on all of our platforms, then everybody uses the same thing.)
That's some funky C code you got there Brian.
1) 'C' is the most basic, it only has functions. A function call looks like "bzuiutil_pause_backup();"
2) 'C++' is a super set of 'C' and if you want to group several static functions together, you can put several static functions in one class. This defines scope, and also allows programmers to find where the function is written, in this case "BzUiUtil::PauseBackup()" means that there is a 'C++' class called BzUiUtil and it has a static function with EXACTLY the same contents as the 'C' version has inside "bzuiutil_pause_backup();"
At Backblaze, we tend to void writing 'C' code in favor of 'C' code in static C++ wrappers as described in point "2" above. There isn't any downside as long as all of your compilers on all platforms are C++ compilers anyway. I honestly didn't know a pure 'C' compiler still exists. Specifically, Microsoft Visual Studio does both seamlessly. On the Macintosh Xcode does both seamlessly. On Debian Linux (the flavor we use most and is also the slowest to adopt anything) the g++ compiler ships for free on every linux box that does both seamlessly.
I've encountered programmers that habitually refer to C++ and ObjectiveC as C. I suspect he has that habit.
No, we use 'C' and 'C++' on Windows PCs, and we use 'C' and 'C++' and 'Objective C' on the Macintosh. Objective C can call into 'C' libraries just fine on the Macintosh. Objective C can also call into C++ just fine on the Macintosh.
The distinctions are very important to us. :-) Objective C cannot compile on Linux or Windows, but 'C' and 'C++' can compile on all three platforms (Macintosh, Linux, and Windows) using the default development environments on all three platforms (GCC on Linux, Xcode on Macintosh, Visual Studio on Windows).
I stand by my guideline for applications.
Another way to think about it is if you want, use all your own custom compiler flags, then if you can manage to get your compiler flags set by a creative combination of BUILT IN compiler flags, that way your code will compile right out of the box correctly on all platforms. What I dislike is that out of the box the code compiles on zero platforms. Instead, you run a "configurator" program that sets all the cross platform flags correctly. If at all possible, write your code so that it self configures correctly with zero changes on all compilers and all platforms. Please notice the "if at all possible". I think if you are relatively careful and just use basic types this can be done easily. Is it really that amazingly hard to find a native "int32" and "int64" type using ifdefs on every platform/compiler?
You really, truly are best off making your own defines. (If you absolutely don't want to specify them in the build setup, you can usually autodetect many of your own defines by checking for _WIN32 and _M_IX86 and _MSC_VER and __GNUG__ and so on. Then there's at least only one place to get the logic wrong.)
After all the problems you encounter porting from Unix to Windows are different from what you hit porting across architectures. With cross-platform you're worrying over file IO functions or the lack of fork equivalent. With cross-architecture you must ensure the code is not abusing ints for storing pointers or assuming sizes of primitives.
You should ensure that anyway, even when you're working on a single architecture. Those are very bad habits to form.
So, for instance, if you want the lower 8 bits of an int don't declare a pointer to a character and then lift out the byte, but do (value & 0xff) instead.
That way if you move your program to a big endian machine it will still work. Storing pointers in ints assumes that pointers and ints are the same size, which is definitely not always the case.
So store pointers in a space sizeof(whatevertype*) rather than sizeof(int).
If you have decided that you do really need it, then use the standard C type for "integer type large enough to hold a pointer". Yes, there is one. Actually two. They're called intptr_t and uintptr_t depending on whether you want signed or unsigned.
Similarly, if you subtract two pointers and need to store the answer (this is less of a cause for reflection, but still not too common) then don't toss it into an int and hope that it works. Use the language-provided type for "difference of two pointers", ptrdiff_t.
If you have a value that needs 32 bits of storage, don't throw it in an int and say, hey, int is probably at least 32 bits. Use the language-provided types int32_t and uint32_t.
Understand what types are provided (Wikipedia has a decent summary at http://en.wikipedia.org/wiki/C_data_types), what they mean, and what the guarantees are. Don't fall into the trap of "it works for me, here, now, so it's fine". If you need a particular guarantee (such as the ability to hold a certain range of numbers), ensure that the type you're using provides it.
All that is a technicality, software has bugs. Heartbleed was a particularly bad security bug. This doesn't mean you throw out the baby with the bath water. As a society we all use HTTPS and it had a bug -> so we fix this bug and move on as best as we can.
As far as I know, something like 99 percent of routers and OS systems such as Linux, Macintosh, and Windows use the OpenSSL library and it has very little to do with Heartbleed. On the other hand, it is fairly important your system vendor stays on top of security patches and applies them.
OpenSSL is software, and therefore has bugs. But the OpenSSL encryption layer did not contain the Heartbleed bug. Or put differently, the OpenSSL encryption libraries that encrypt AES 128 and AES 256 had nothing to do with Heartbleed, and those are the parts we continue to use at our company completely for free, without paying any royalties to anybody, and they are a lot more secure than anything we could have hand-rolled ourselves even if we took years to do it.
I'm a strong believer in this approach.