Python 3.x: RCE in Python applications that accept floats as untrusted input
cve.mitre.org
cve.mitre.org
3.6.13 and 3.7.10 have already been released with the fix: https://www.python.org/downloads/release/python-3613/ https://www.python.org/downloads/release/python-3710/
The 3.8.8 and 3.9.2 release candidates with the fix will be promoted on Monday March 1st: https://discuss.python.org/t/python-3-9-2rc1-and-3-8-8rc1-ar...
If you're on 3.5 or lower, these versions are no longer receiving security fixes and will not be patched: https://python-release-cycle.glitch.me/
Case in point: The Debian security tracker, see their notes section referencing each commit.
https://github.com/docker-library/python/blob/master/3.8/bus...
>>> import ctypes
>>> x = ctypes.c_double.from_param(1e300)
>>> repr(x)
Segmentation fault
This happens when getting the string representation of a foreign function float. It doesn't affect the standard builtin float type. So it's not extremely common.But I can imagine this cropping up if you e.g. find a way to cause an exception that formats a value. Or if there's aggressive logging.
>>> repr(x)
"<cparam 'd' (1e+300)>"ctypes has never been considered even remotely secure, it can call any library function, including, drumroll, sprintf!
Their build system is interesting. One of the quirks is you define your application, and every library/package that you depend on. They don't depend on system libraries for their application code. The idea is to get as consistent an operating environment where possible. It's both generally amazing, and an absolute pain in the arse (usually when you least want it to be, because some upstream package changed their dependencies and you end up with version conflicts to unpick).
When Heartbleed came out, the patches landed in the Amazon build system for the OpenSSL package something like midnight. By the time I got in to the office in the morning, almost every service had been fully rebuilt with patches, and services that do CI/CD had already had the patches deployed. Services that didn't have CI/CD were already kicking off deployments of their front end fleets. IIRC they were paging teams when relevant packages were complete to make sure deployments got kicked off ASAP.
So in this case, someone will have patched Python packages for relevant versions, built an updated version, and everything that depends on it will have immediately recompiled, and from there all the packages that depend on those, and so on down the line. Given how python is used a lot for operations etc. I wouldn't be surprised to find a significant chunk of Amazon got rebuilt today, even in cases where Python wasn't being exposed to external users, or even used by the service directly. That's a lot of components, and probably left zero capacity left for anything unrelated, and no doubt there will be quibbles about the ordering in which things got built.
On a system without these protection mechanisms it would be a easy win.
Reference: https://haxx.in/posts/numeric-shellcode/
The sprintf is to a temporary buffer that's converted to a PyUnicode object before returning, so subsequent portions of the string are written elsewhere.
https://docs.google.com/presentation/d/19K7SK1L49reoFgjEPKCF...
I found one used with the return value from alloca (inside FindAddress in _ctypes.c). It checks alloca for NULL (does alloca ever return NULL?), but I could imagine it might expoitable. FindAddress is a static function that can be called during DLL loading. I imagine that there is very little code that accepts untrusted arguments to DLL loading though (if so, there are bigger problems...).
There is also a lot of use of fixed 32-byte buffers, for up-to-64 bit numbers, which is fine, but hopefully if in the future they can be 128-bit, people remember to fix it!
In other cases, like getnameinfo, the stack-allocated buffer is way too large! %d can never be 512 bytes, so it's just thrashing your cache for no reason.
Anyway surely this has been carefully audited, since grepping for sprintf is so simple!
That is rarely a safe assumption.
Why? I think it’s because incorrect statements can be corrected in a reply. OTOH correct statements with which one disagrees or otherwise dislikes cannot be corrected but they can be downvoted.
>>> import ctypes;x = ctypes.c_double.from_param(1e300);print(type(x))
<class 'CArgObject'>
When repr'd, an sprintf call with a statically sized buffer of 256 bytes is used to produce a string: https://github.com/python/cpython/commit/d9b8f138b7df3b455b54653ca59f491b4840d6fa#diff-4e23b3237d0aa08bf4c434d75fab19200a80837bd147051fefccd98b7f2480faL500-L507
With 1e300, the resulting string doesn't fit into 256 bytes, thus overflowing the buffer. Exploitation might be interesting, as you (probably?) can only use numbers (ascii range 0x30-0x39): >>> len("%f" % 1e300)
308As far as I know, ALL bindings in CPython are written with the C APIs and not ctypes, including the JSON library. (I can't guarantee that but it shouldn't be that hard to audit.)
The official page is not any more clear on this:
https://python-security.readthedocs.io/vuln/ctypes-buffer-ov...
I use one ctypes binding for convenience (a single function for CommonMark) but I prefer to use C APIs for this reason ... it's a lot of complexity and unsafety. Using ctypes wrong can crash your process due to invoking undefined behavior.
It would be nice to have a list of bindings created with ctypes, since I think most consumers probably have no idea. I thought that PyOpenGL is done with ctypes but don't quote me on that ...
- `uuid` used ctypes until 3.9 to get information like the IP address (without floats, so it's not vulnerable)
- `platform` uses ctypes in 2.7 to get the Windows version (without floats)
- `multiprocessing` uses ctypes for interoperability with ctypes
That's everything, on the versions I checked. `json` for example uses a native Python module `_json`, so it doesn't use ctypes.
Only issue is that the default string type requires heap allocations
I'm surprised noone ever looked at that code --- even when writing it --- and thought "how long can a %f get?" I've been writing C for a long time and that's just something which comes naturally, being ingrained into memory since the beginning. If I see a fixed-size buffer I will always question whether it's big enough (and also if it's perhaps even too big.)
The "psychology" around buffer overflows has always seemed strange to me; a real-world analogy is someone who has no idea how big his car is, finds a parking spot that "looks big enough", and just rams it in without a second thought, sometimes crashing into the surroundings. Not many people would do that in the real world. Yet countless programmers seemingly can't get something simple like this right?
Edit: downvoters, care to state your case?
It is, but it is an extremely difficult problem. I think “How to print floating-point numbers accurately” (https://dl.acm.org/doi/10.1145/93548.93559) was the first correct implementation. That is from 1990 and, according to its authors “was almost 20 years in the making” (http://kurtstephens.com/files/p372-steele.pdf)
(Faster versions have since been published)
I think it is unfortunately that %f, the easiest to remember type field for floating point number, is used for this behavior. Most of the time, you’ll want to use %g, which uses scientific notation if it is shorter.
I guess linters should warn about bare %f without length fields.
I would have expected this issue to be a solved problem too, and now we get extremely insecure Python 3 codebases as a result of this vuln.
So much for moving from Python 2 to 3, will now wait for Python 4.