Make CPython segfault in 5 lines of code
gist.github.com
gist.github.com
One I ran into in the wild recently is that in older versions of Lua, exceptions in GC finalizers (the `__gc` metamethod) can trigger a segfault. In those same versions of Lua, the bytecode format is notoriously dangerous to load.
I wonder whether this will be a large component of newer scripting language implementations. Do these safety issues warrant use of memory safe languages like Rust, or use of existing sandboxed VM implementations like WebAssembly?
Could you expand on this? Are the dangers such that they could be avoided with a bytecode verifier somewhat like Java's? Things like checking that the stack can never underflow, and that at the merge points of branches the stack always has the same depth.
5.1: https://www.lua.org/wshop11/Cawley.pdf
5.2: https://apocrypha.numin.it/talks/lua_bytecode_exploitation.p...
Lua used to have a built-in bytecode verifier as I understand it, but it never reached the point to where it was enough to safeguard the VM.
Here's a trivial example:
/usr/bin/gawk 'for (i = ) in steve kemp rocks'
I found fuzz-testing like this very very useful when writing my own BASIC interpreter, and playing with scripting languages though. It's almost magical how quickly problems are found!I don't know what the stance of other python runtimes are, but you should probably just use a sandbox at OS level which is likely to be tested far more thoroughly.
[0]: https://devblogs.microsoft.com/oldnewthing/20060508-22/?p=31...
> I now think that putting a sandbox directly in Python cannot be secure. To build a secure sandbox, the whole Python process must be put in an external sandbox.
https://mail.python.org/pipermail/python-dev/2013-November/1...
https://web.archive.org/web/20190113115213/https://blogs.msd...
CPython bytecode will segfault Python if it’s slightly incorrect. And that’s fine, it’s not a security risk and it’s not worth the performance overhead of validating wonky bytecode.
What about Javascript running on V8?
[0] https://github.com/nodejs/node/pulls?utf8=%E2%9C%93&q=is%3Ap...
[0] https://github.com/nodejs/node/pull/22273#issuecomment-41588...
It's a plausible approach to a fully-safe, near-native-speed plugin architecture.
A program fragment is safe if it does not cause untrapped errors to occur. Languages where all program fragments are safe are called safe languages. Therefore, safe languages rule out the most insidious form of execution errors: the ones that may go unnoticed.
...
It is useful to distinguish between two kinds of execution errors: the ones that cause the computation to stop immediately, and the ones that go unnoticed (for a while) and later cause arbitrary behavior. The former are called trapped errors, whereas the latter are untrapped errors.
Type Systems, Luca Cardelli
https://scholar.google.com/scholar?cluster=90442457768317510...
And practically speaking seg faults are easy to debug and fix.
I see a lot of abuse of the terms "safe" and "safe language" lately.
If the language only has trapped errors, which are indicated by seg faults, it's a safe language.
Every language has such errors. What does divide by zero do?
What does blowing the stack / infinite recursion do in Rust? It seg faults.
The seg fault is the safe behavior. If the stack overflow overwrote heap data structures or global data structures and the program kept running, that would be unsafe.
Returns a value, of course! (And that said, JavaScript does have some trapped errors, such as (1/0).foo.foo (and yes you need the second .foo…))
IMO, the execution "error" here (in this thread) is accessing memory illegally. Sometimes the runtime traps it, but sometimes it does not, and "sometimes" isn't always, so its effectively untrapped as we cannot depend on the trap. (Especially in adversarial circumstances.)
And further, the original quote talks about languages — the behavior of a language like C is that memory access is not necessarily trapped; the behavior is not well defined. Given the lack of a requirement in the C language for a trap, I think it is fair to call C "unsafe" given the above definition of safe/unsafe.
> What does blowing the stack / infinite recursion do in Rust? It seg faults.
Somewhat interestingly, it detects it and SIGABRTs, which technically isn't a segfault. And that's now some black magic that I'm curious about as I really thought it would have segfaulted.
It's a signal handler, of course: https://github.com/rust-lang/rust/blob/d8bdb3fdcbd88eb16e1a6...
I don't think anyone would disagree with this. What they're saying is that a segfault is safe and that's because a segfault is essentially your OS's version of an out-of-bounds error, one of the reasons that C is not safe is because an out-of-bounds access will not necessarily cause a segfault.
I don't understand the nuance here. In my Firefox developer tools, I can do the following:
> (1/0).foo.foo
---> TypeError: (intermediate value).foo is undefined
> (1/0).foo
---> undefined
> (1/1).foo.foo
---> TypeError: 1.foo is undefined
> Infinity.foo.foo
---> TypeError: Infinity.foo is undefined
> undefined.foo
---> TypeError: undefined has no properties
While these are all different errors, I don't really understand why the '(1/0).foo.foo' case is an example of a _trapped_ error, but the others are not.
In Python, Java, C, OCaml, and most other languages, 1 / 0 aborts the program. That is, division by zero is a trapped error.
In JavaScript, it's not. It keeps going and lets you do stuff like access nonexistent properties .foo on the result, which are also untrapped errors.
So JavaScript is unsafe in Cardelli's terminology. It gives you untrapped errors rather than trapped ones. The program keeps chugging along until you find out later and have to trace backwards to the bug.
Prelude> 1.0 / 0.0 :: Double
InfinityI’m sort of surprised by that since my memory is that infinity isn’t a real number, rational, etc. That is, does infinity in Haskell obey some algebraic laws?
However integer division by 0 isn’t. div 1 0 will fail.
This answer has an interesting way of looking at it. If you go on the theory that floating points are supposed to represent reals, then in floating point, you can't tell if a value is actually zero or just indistinguishably close to zero.
In the case of "indistinguishably close to zero", you're getting the wrong answer, and the program doesn't halt. It keeps on chugging doing bad math. So that's an untrapped error, and it's UNSAFE by Cardelli's definition.
https://cs.stackexchange.com/questions/82811/why-do-floating...
The key point is that "safe" sometimes means "crashes" and sometimes means "doesn't crash". It's an auto-antonym in that sense.
A broader definition is "errors are flagged as early as possible", including with seg faults / hardware exceptions.
Try it again with div for integer division.
WebAssembly koolaid is strong on HN, let's wait the first exploits that escapes the runtime to assess the "fully-safe" architecture.
The way they said it, though, makes it sound like WebAssembly is implemented with full process sandboxing or something, which is patently false. It works that way in neither Chrome nor Firefox, and there are no other browsers right now.
The WASM sandbox can call only set of specified host functions, but I expect so much functionality snowballing inside sandboxes that we'll have to allow everything including unsafe ones, anyway.
So it can segfault all the way up to the machine level. But more importantly, it's safe for it to 'segfault' out of its allocated memory, because it can't reach any other memory with 32 bit numbers.
But now the processes are behemoths with gigabytes of dynamically linked libraries that are too hard to secure and to restrict system access, we just enable everything. This will happen to wasm. I'm sure there are WASM blobs configured with unfettered access to DOM in the wild already.
Right, but look at what I'm replying to.
"The way they said it, though, makes it sound like WebAssembly is implemented with full process sandboxing or something, which is patently false."
Unlike with Javascript, the program hosting a WASM script is immune to corrupt pointers inside the script. That's equivalent to OS-level isolation, which is pretty good!
> But now the processes are behemoths with gigabytes of dynamically linked libraries that are too hard to secure and to restrict system access, we just enable everything. This will happen to wasm. I'm sure there are WASM blobs configured with unfettered access to DOM in the wild already.
There will inevitably be bugs in the code handling the html/css/dom/rendering. But WASM greatly reduces the attack surface of the scripting VM.
The types of bugs most likely to still exist with WASM are the types where isolating it in a separate process that communicates by pipes wouldn't help.
import sys, threading
def r():
sys.stdin.buffer.read(1)
t = threading.Thread(target=r, daemon=True)
t.start() import sys, threading, time
t = threading.Thread(target=sys.stdin.read, args=(1,))
t.start()
time.sleep(1)
sys.stdin.close()
Run it then after a few seconds press enter. It doesn't segfault in Python 3, but it still doesn't behave how I'd like, because I would like the close() to unblock the read(), but it doesn't unblock the read(), the read() still hangs until it gets some input.IronPython did that too, on .Net. It ran around one quarter the speed of CPython.
I wouldn't have bothered filing the bug.
If I jump on my bed enough, it'd probably break, but I'm not complaining to the manufacturer about the issue.
if I used my Snap-On(tm) wrench as a prybar (incorrect usage) and broke it, Snap-On would still replace it in exchange for the broken tool and knowledge of the situation that broke it.
To pretend that a language bug isn't worth reporting because you and your codebases will never encounter it seems short-sighted. Down the line, years from now, who knows what you'll have to do to get something to work. Maybe it'll be something this silly, and you'll be happy that the folks before you encountered it and remediated it.
All that said, from a practical standpoint I agree with you.
If you're doing something wacky, and it turns out as wacky as you thought it would, you're probably attacking the problem from the wrong angle, anyway. I just want to remind everyone that 'wacky' things are required and implemented daily in codebases around the world -- regardless of how bad they smell.
for x, x.__new__ in [(__import__('queue').Full, print)]: __import__('glob').iglob(0).throw(x)Someone commented below the gist with this one-liner:
(i for i in []).throw(type('E', (BaseException,), dict(__new__=lambda cls, *args: cls))())
I managed to golf it a bit down to this ;): n="__new__";(i for i in []).throw(type(n,(IOError,),{n:lambda c,*a:c})()) $ python3.7 -VV
Python 3.7.5 (default, Nov 7 2019, 10:50:52)
[GCC 8.3.0]
Some more golfing with yours: (i for i in[]).throw(type('',(IOError,),{'__new__':lambda a,*b:a})) import ctypes
ctypes.cast(1, ctypes.py_object)
Interestingly, this works: import ctypes, gc
x = 22
_id = id(x)
del x
gc.collect()
y = ctypes.cast(_id, ctypes.py_object).value
assert y == 22How about in Rust, then? [0]
Bugs happen in every language. When memory corruption occurs, you can segfault.
[0] https://users.rust-lang.org/t/rust-guarantees-no-segfaults-w...
At a guess, you want something with dependent types?
Like Idris, or Haskell. You already have a Haskell example. This [0] release of Idris fixed a segfault when concatenating strings.
Maybe you meant a language that is proven from the ground up. Like CakeML. You can find a segfault example here [1].
Maybe you meant a language with an algebraic type system like Ada. You can find a segfault example here [2].
Maybe you meant something like Dotty (Research for the next version of Scala). You can find a segfault example here [3].
In short: You'll need to describe what you believe to be a "proper" type system, and name the languages you think fit that description, or no one can have a conversation with you.
[0] https://www.idris-lang.org/idris-1-1-1-released/
[1] https://github.com/CakeML/cakeml/issues/438
[2] https://stackoverflow.com/questions/56227629/segmentation-fa...
Anyway, I said that this specific issue would not occur in a language with dependent types -- where incorrect code would cause the implementation to crash. Not that it is impossible to have a buggy compiler that at certain cases produces segfaults.
That's exactly what happened here, however. The instance check was missing from the interpreter.
Dependant types wouldn't have solved the underlying problem.
Apparently neither Rust, Haskell nor Go fit.