Reverse Debugging for Python
morepypy.blogspot.com
morepypy.blogspot.com
Completely serious: when can PyPy obsolete CPython? Why don't more CPython core devs work on PyPy? Besides perhaps some remaining gap for C-extensions support, what else is missing?
PyPy is also still several versions behind implementing the latest Python 3.
So while PyPy is definitely a great project, groundbreaking in many ways and a very useful for a lot of applications, it's not a replacement for all applications, for now at least.
Also while the actual PyPy Python interpreter is reasonably simple overall the project is incredibly complex. It at least appears as if there are parts of the PyPy project not even all PyPy developers understand. Making PyPy the reference implementation is probably questionable for that reason alone.
PyPy 5.3.0 is twice as slow as Python 3.5.0 for some work I regularly deal with (1.1s for PyPy, 0.6s for CPython). Another project it is ~40% slower (7.7s vs 5.6s).
Additionally, PyPy just released an alpha containing compatibility with Python 3.3 (which was released four years ago). There are some nice additions to the language in 3.4 and 3.5.
Is this because you are using ctypes or a C extension written against the CPython API? If not, have you filed a bug report?
Currently I don't have an interest in taking time to try and make PyPy less slow for my use case, for many reasons that I'm not going to get into here. (But I will note I'm primarily writing code that focuses on being easy to distribute. Most users aren't going to have PyPy installed anyway.)
Anyway, the main thrust of my post is just that Python isn't a monoculture in which PyPy is always a better solution.
That said, PyPy is awesome. :-)
This is a rather slow feedback cycle.
Secondly, while the syntax of RPython may be clean the necessary incantations to actually achieve something aren't at all obvious.
Greg
no python 3 as usual :( 99% of my production code is python 3.
Contribute work or funding[0] to pypy3?
[0] http://pypy.org/py3donate.html
[1] http://doc.pypy.org/en/latest/release-pypy3.3-v5.2-alpha1.ht...
I've done similar things before – in PANDA we don't strictly need to snapshot the full device state when creating a recording; it would be enough to just keep track of which memory regions are I/O and reconstitute that mapping. But QEMU's savevm/loadvm saves and restores that mapping as a side effect so it's easier to just let that happen.
Edit: also, instead of disabling ASLR system-wide, it might be better to just use "setarch `uname -m` -R <pypy>", which disables it for just a single process.
As a practical matter, ASLR should have little if any impact to nearly all programs design, so disabling it at development/debug time should not come at a big cost.
BTW this post states "There is no fundamental reason for [ASLR] restriction, but it is some work to fix."
EDIT: I see, your focus is specifically on recording. Disregard this comment, good point.
> Only works on Linux, and only with Address Space Layout Randomization (ASLR) disabled. There is no fundamental reason for either restriction, but it is some work to fix.
print id(object())In the example given here, you could get away with it since control flow doesn't depend on the value of id(object()). But for example, if you had:
if (id(object()) >> 12) & 1:
print "foo"
else:
print "bar"
You would need to guarantee that id(object()) always returned the same value so that the execution follows the same path.What I should have said is that the address space layout is a source of non-determinism that would otherwise need to be accounted for in the recording of a program.
Further, I hold to the view that given a program, knowledge of the thread execution schedule/scheduler and the results of external calls, a program's execution should be deterministic.
Thinking about this now, I realize disabling ASLR shouldn't be enough to fix this problem since recording and replaying are sufficiently different.
Ah yes, it needs to be deterministic in the context of a time-traveling debugger, I was thinking about the general case, sorry.
> Further, I hold to the view that given a program, knowledge of the thread execution schedule/scheduler and the results of external calls, a program's execution should be deterministic.
Surely it is, if you precisely know all timings and IO and can replay them there's nowhere for non-determinism to creep in.
That's not really an acceptable scale of required knowledge for a time-traveling debugger though.
This was my original point, the address space layout information leaks into python land via things like id.
> That's not really an acceptable scale of required knowledge for a time-traveling debugger though.
That's what recording captures (though, you don't need to know the times at which things happened).