CPython, C standards, and IEEE 754
lwn.net
lwn.net
> Multiplying infinity by any number is defined to be a NaN for IEEE 754
No, multiplying infinity by any number other than zero or NaN produces an infinity. Multiplying infinity by zero produces a NaN.
They don't normally mean different types of value. 4.0000000 and 4.0000001 are just different adjacent numbers.
> I guess I am not sure how this could be described as an optimization... an optimization relative to what? Some hypothetical floating-point format which has extra bits?
Using the exponent to distinguish, like the format uses for normal/subnormal/nonfinite. Or using extra bits, sure.
They took a bit pattern that would have been a NaN if not for Infinity, and used it to store Infinity.
- normal (exponent not minimum or maximum)
- zero (minimum exponent, mantissa == 0)
- subnormal (minimum exponent, mantissa != 0)
- infinity (maximum exponent, mantissa == 0)
- nan (maxmimum exponent, mantissa != 0)
NaN is only the last category, strictly. Finite is everything except infinity and nan. Subnormal and zero are collectively denormal. The distinction between "subnormal" and "denormal" is a bit esoteric.
> The calculation was using the HUGE_VAL constant, which is defined as an ISO C constant with a value of positive infinity [...].
Only if the floating point implementation has infinities, which ISO C does not require. C99 also defines INFINITY, but that one’s also not required to actually be an infinity, unfortunately.
More seriously:
> "Nowadays, outside museums, it's hard to find computers which don't implement IEEE 754."
AFAIU there are plenty of implementations that either run faster without or outright refuse to implement the “gradual underflow” semantics, which 754 technically requires with no equivocation. (Wikipedia tells me the Intel compiler turns on SSE “denormals as zeros” when optimizations are on, for example.)
Linux kernel developers want to use C11. CPython developers want to use C11. But does your compiler support it? If you were a C++ developer, would you request C++17/C++20 also?
CPython must be aligned with the manylinux project. manylinux build every python minor release from source. Historically manylinux only use CentOS. CentOS has devtoolset, which allows you use the latest GCC in an very old Linux release. Red hat has spent much time on it. Now you can't get it for free because CentOS is dead. So the manylinux community is switching to Ubuntu. Ubuntu can only provide GCC 6 to them. If anycode in CPython is not compatible with GCC6, then manylinux will have to drop ubuntu too. And they don't have any other alternative. As I said, for now Red Hat devtoolset is the only solution. When IBM discontinued the CentOS project, most people do not understand what it means to the OSS community.
(I work for Microsoft. It doesn't stop me to love Linux.)
See: PEP 600, PEP 656.
While the manylinux infrastructure project may use Alpine to produce "musllinux" binaries, Alpine will never be usable to produce "manylinux" binaries. So Alpine would only be useful in addition to another distro, not instead of another distro.
Alpine is a supplement. It can't replace the others. It hasn't been widely accepted by the community yet. I believe many python projects are not afforded to replace libc.
So? Who are those people building for servers they don't control and arbitrary old versions of OSes?
If you build your own wheels, of course no one’s going to stop you from building against anything.
(in other words: my understanding is that software languages and tools can continue to evolve while maintaining the ability to build backwards-compatible binaries)
yes absolutely. I think that it is wild that some people stay with older compilers solely because they happen to be on an old distro and don't want to update. Tying compiler versions to operating system versions is absolutely braindead. A compiler is just a program that takes text and outputs a binary which is supposed to work on anything with the expected ABI - you could even cross-compile from windows if you wanted ; at least I've done the opposite (cross-compile windows binaries from a linux host) a few times.
e.g. personally I mostly use clang-13, soon 14, for development with every possible C++20 goodies, and ship software that works back to windows 7, mac os 10.13 and until recently centos:7-era linux userland (recently upgraded to 8). There is zero reason to use an older compiler.
> I don't know why GCC-12 was in the picture, but I hope all of you can understand one thing: python's source code must be compatible with GCC 6. And no so called "Threading [Optional]" features. It is all because CentOS is dead. CentOS has been sold to IBM. There is no CentOS 8 anymore. The free lunch is ended. And there wouldn't be another OSS project to replace it.
the day centos:8 ended I replaced centos:8 with rockylinux:8 in my docker image for my builds and everything continued working fine with the latest GCC and Clang versions.
Honestly your pains regarding gcc-6 are entirely self-inflicted. Just ship a more recent GCC binary with the manylinux project or something, you can build it on centos 5 if you fancy in order to get it to work on older libcs
And while I'm keenly aware that Python and C++ are very different languages, if someone asked if I'd insist on using the "new" Python17 (aka 3.6), like that was wildly and unreasonably bleeding edge, I'd literally laugh at them.
For macOS and iOS, the problem is more complicated. Let's say I want to publish a package to pypi.org. The package contains some binaries compiled from C++ and it requires C++17. Then what is the lowest macOS version the binary can support?? It's very tricky that if the build machine which generates the package is macOS 11, then it can have Xcode 12.5. Otherwise it has to use XCode 12.4 or lower. And in XCode 12.4 it says std::optional, which is a C++17 feature, is only supported in macOS 10.14+. But if the build machine has macOS 11 and XCode 12.5, then the binary would be able to run on 10.13 too. And in whatever case, it can't support 10.12 or lower.
I believe you don't want to build tensorflow/pytorch packages from source. So the maintainers of these two OSS projects must consider the things above. If you wonder why they don't use C++17, this is the reason. For the same reason they want to avoid C11 optional features too.
No it's not. The oldest supported version of macOS your binary will run on will be the one you set in your CFLAGS with the -mmacosx-version-min=10.x flag when you build.
The OS you are running the build on and the Xcode version don't influence that, you should just run the most recent Xcode you can, build against the most recent macOS SDK available and set that flag. You can target 10.8 from Big Sur and Xcode 12 with afaik no issues.
Usually it doesn't. But in this special case it does. I have set the flag you said. When C++17 wasn't in the picture, it's enough. But now we are talking about how to enable developer using the new language features. New C++ features typically need new runtime. But the old systems do have it. Then one of the difficulty is how to figure out which macOS versions have it, which doesn't.
You can also link statically against libc++.
The result of this is that Apple really assumes you are going to use the copy of libc++ that ships on the system, and it might be missing exported systems that are assumed by the library. But this isn't related to the toolchain you are using. I do remember some weird corner case with std::optional and I think it might have been that Apple forgot to mark something correctly as unsupported in a previous toolchain build? The newer compiler is thereby actually giving you the more correct understanding of what can actually be presumed to work on older systems.
But like, the core issue here is that you really shouldn't use Apple's stupid toolchain setup, nor do you have to: just embed your own copy of libc++ and call it a day. I routinely use the latest versions of Xcode's copy of clang (though I prefer to compile for macOS now using the copy of clang that comes with the Android NDK for various reasons: I recommend avoiding Xcode except to get their system headers) to target ancient systems with modern C++ features as I'm not relying on the system copy of libc++ to even exist or work (10.7's doesn't) much less be complete.
At work we sidestep this by building the new C++ standard library implementation against an old libc, and ship it with our product. I can imagine that this could be problematic for other software though, especially ones that normally ship with the disto.
why would it be an issue ? that's how pretty much every software does on Windows and most certainly people can agree that things are much more sane there for the end-users
"Compilers and Runtimes" > "CentOS sysroot for linux-* Platforms" https://conda-forge.org/docs/maintainer/infrastructure.html#...
From https://conda-forge.org/docs/user/announcements.html :
>> 2021-10-13: GCC 10 and clang 12 as default compilers for Linux and macOS
>> These compilers will become the default for building packages in conda-forge.
Conda-forge specifies enough to solve for CentOS sysroot compatibility and newer GCC is already specified.
CentOS (RHEL SRPMs with RH trademarks removed + EPEL) lives on as {Rocky Linux, Alma Linux, CentOS Stream,} and SUSE is still RHEL-compatible.
https://github.com/pypa/cibuildwheel :
>> Build Python wheels for all the platforms on CI with minimal configuration.
>> Python wheels are great. Building them across Mac, Linux, Windows, on multiple versions of Python, is not.
>> cibuildwheel is here to help. cibuildwheel runs on your CI server - currently it supports GitHub Actions, Azure Pipelines, Travis CI, AppVeyor, CircleCI, and GitLab CI - and it builds and tests your wheels across all of your platforms.
Yes, ImportC is a C11 compiler!
https://dlang.org/spec/importc.html
`_Static_assert` and `_Generic`, too.
No C compiler is complete without extensions. ImportC is no exception. It has modules, forward references, and compile time function execution.
C is broken from C11 to C26 by committee. GCC was broken from 9 to 11. Python has at least the option to use GitHub, which can eventually detect PR's with unidentifiable identifiers. Linux will need to use linters to detect homoglyphs or bidi attacks, because reviewing emailed patches is impossible.
Most of the time float is just an implementation detail when you really want a decimal. I think the literals should be a decimal, and you could explicitly cast float when you actually need to do floating point math, or as an optimization. But I know it's too late for that.
Anything with a REST api or JSON doesn’t have native support for Decimals, so to use them you have to represent them in component form as an object or as a string.
Ints just make that easer at the slight cost at display time. Stripe is a good example of this.
> Otherwise you're applying that exponent everywhere you use the value, whether you display it or use it in a mathematical expression.
Slightly disagree, I think it’s only when displaying you have to convert to a “human readable” decimal form. They don’t need converting to Decimal for processing.
It’s also worth noting there are actually two none decimal currency’s still in use:
“Today, only two countries have non-decimal currencies: Mauritania, where 1 ouguiya = 5 khoums, and Madagascar, where 1 ariary = 5 iraimbilanja.”
This is a very slight nitpick (that might be wrong) but I think it's perfectly in line with the JSON specification to interpret JSON numbers as decimals.
"JSON is agnostic about the semantics of numbers. In any programming language, there can be a variety of number types of various capacities and complements, fixed or floating, binary or decimal. That can make interchange between different programming languages difficult. JSON instead offers only the representation of numbers that humans use: a sequence of digits. All programming languages know how to make sense of digit sequences even if they disagree on internal representations. That is enough to allow interchange."
https://www.ecma-international.org/wp-content/uploads/ECMA-4...
https://datatracker.ietf.org/doc/html/rfc7159
And most JSON parsers allow you to extend and change how types are handled. I think however from a developer point of view by standardising on ints you reduce the risk of making a mistake in your implementation.
I think the point is though that as long as your system doesn't need to handle fractions of a cent/pence/etc then using them as the base representation and only doing integer math makes the system simpler. It moves the complexity of currency to the display layer away from the business layer. Obviously there are times when you need to be able to handle smaller units and a decimal type would be correct there.
⇒ If you think the “sensible default behavior is to preserve precision“ I think you would have to call for (at least) rational as the default non-integer integral type. (If you think of √2 as a literal, it gets more complicated. In the end, you might have to use a representation that’s similar to that used in computer algebra systems)
Edit: you could also require constants that do have an exact representation in your floating point format. That would be annoying, too, though, but maybe not too annoying if you had a way to specify “the number closest to this one”, say by requiring one to write ~2.1~ instead of 2.1
IMO decimal rounding of those fractions, while not ideal, is a lot more understandable to a typical person. If 1/3 * 3 = 0.99999999999999999 that's annoying, but not crazy in the same way that the floating-point equivalent is.
(Also, "correctness" seems like the wrong word: making rationals the default type would give you perfectly precise literals, but you'd immediately lose that once you perform almost any operation on them.)
A default number format that is much slower than built-in floating point but still subject to unexpected deviation from real arithmetic numbers doesn't seem like a win for users.
Almost nobody thinks about float rounding, they just use something precise enough that it doesn't matter, or use integers.
Why break that standard? Everyone expects computer numbers to be just a bit off. If we didn't we would do a lot of stuff differently in incompatible ways.
Seems like it would hurt the Python ecosystem if people wrote stuff that assumed floats were perfect decimals, and encouraged code that depends on that.
You'd get stuff like serializing to JSON, and reading in some other language with a JSON parser that doesn't know the numbers need to be precise.
Plus, it's just faster.
Maybe off-topic, but in which language of theese two do you find a joy to program in?
As for me, Python's features are godsend. Only, sometimes I want that procedural simplicity of C.
I feel dirty trying to emulate C coding style in Python, even when it makes sense.
There are much better modern alternatives to get comparable speed for new projects. Many projects that used to be written in C are being rewritten in either C++ or Rust for better memory safety and reasoning about the memory model.
Disclaimer: I did my PhD (several years ago) in a systems and networking lab and most of the code I wrote was in C. Now for example modern kernel modules tend to be written in Rust if possible, while such an option was not available while I was working there
Not actually at the bottom, I'd choose it over Forth and ASM, but.... I haven't written it for anything but a microcontroller in years.
I don't really do much outside of Python and JS in general these days. Python's performance is basically the same as C, because... There's already a C extension for everything!
I'm really not even a fan of compiled languages in general especially not for open source work. Things like plugin architectures are a lot easier in dynamic everything is a dict type languages.
But C sure is good at being ubiquitous! I appreciate the heavy standardization.
(see also other prolific pseudonymous Linux contributor "George Spelvin <linux@horizon.net>")