PySnooper: Never use print for debugging again
github.com
github.com
Some codebases have built-in logging or tracing functionality: you just flick a switch in a module and begin to get ready-made prints from all the interesting parts. But I've found myself never using those: they are not my prints, they do not come from the context I'm in and they don't have a meaning.
Use what you want but please don't underestimate prints.
Screw it. print, run directly from shell, done.
>>> import logging
>>> logging.info('Hello, world!')
>>> # right, my log message went nowhere
Because by default, logging goes nowhere. And if you configure logging - using a most unintuitive config format (it's so weird that even the documentation about it can't be bothered to use it, but reverts to yaml to explain what it means!) - there's a good chance that loggers created before you got around to configure it (for instance if you, God forbid, made the mistake of adhering to PEP-8 and sticking your import statements at the top) won't use your configuration - and thus send their log messages, again, nowhere.Also, it's slow as hell.
Python's logging infrastructure is pretty bad but you fail to give any good, factual reason why it is. Instead you just vent your frustration on HN, making that platform all the more depressing to read.
Why is it this way, do you think? (Is it a reasoned stance? I would have expected the logger to send messages to stdout by default, so at the risk of getting a "Read the docs!" am I going to be equally surprised at the behavior of basicConfig?)
It gives you contextual debugging - so you can put fancy prints throughout a function but they are silent unless some context is true. It's useful for when you have a hot path that is executed a lot but you only want debug prints for one of those invocations.
Always logs to /tmp/q no matter what stdout redirects the app has set up. Syntax highlighting. Context dumping. Dumps large values to separate files. Etc.
In general, something like this has been around in IDEs for a very long time - e.g. Visual Studio proper added them back in 2005, except it calls them "tracepoints".
[1] https://django-extensions.readthedocs.io/en/latest/runserver...
[2] https://werkzeug.palletsprojects.com/en/0.15.x/debug/#using-...
import pdb; pdb.set_trace()
It pulls up console debugger and is a standard python package.But just in case if you from print cult, you will like https://github.com/gruns/icecream
Like, debugging should be considered part of programming, and a local dev environment that can’t be debugged should be viewed to be as broken as a codebase without e.g. a way to run the server in watch mode.
Also someone should make a debugger that supports the equivalent of print statements, e.g. set print breakpoint on a variable to print its value every time it’s run, instead of typing print everywhere.
If what you want is a breakpoint that just prints something out when it's hit, then these already exist, e.g.:
https://code.visualstudio.com/docs/editor/debugging#_logpoin...
But "set breakpoint on a variable" kinda sounds more like a data breakpoint (i.e. break/log when value changes)? These also exist in some places:
https://docs.microsoft.com/en-us/visualstudio/debugger/using...
but usually with many limitations. The problem with these is that it's hard to implement it in such a way that there's zero overhead when not debugging. If I remember correctly, native data breakpoints have some hardware support on Intel, but for high-level languages it can be difficult to map their data model to something like that.
Traditonal debuggers like gdb for compiled languages support this (breakpoint actions, memory write breakpoints, and variable displays). If something similar isn't already in your language's debugger, that might be a source of ideas for adding it.
The other day I was writing LPEG [1]. I had a rule that wasn't firing for some reason and I wanted to know why LPEG was skipping that part of the input. It was not a simple matter to fire up a debugger and put a breakpoint on the rule:
-- look for 1 or more ASCII characters minus the colon
local hdr_name = (R"\0\127" - P":")^1 -- [2]
I mean, I could put a breakpoint there, but that would trigger when the expression is being compiled, not run. LPEG has its own VM geared for parsing (it really is its own language) so yes, it is sometimes difficult to debug (especially when you end up with a large parsing expression to parse a document---I only found the bug when I was creating another document similar to my initial testing document that was just different enough).Fortunately, there is a way to hook into the parsing done by LPEG and I was able to insert literal print statements during parsing that showed what exactly was going on (my rule was a bit too promiscuous).
[1] Lua Parsing Expression Grammar http://www.inf.puc-rio.br/~roberto/lpeg/
[2] Yes, a regular expression for that is smaller than what I typed, but with LPEG, I can reuse that rule in other expressions. Also, that wasn't the actual rule used, but the effective rule I was using.
Since Python's built-in tracebacks are pretty minimal, the default crash logs don't offer much help other than a line number. I ended up writing a tool that prints tracebacks along with code context and local variables [i], sort of a souped up version of the built-in crash message. It's surprising how much that already helped in a few situations, makes me wonder why it isn't the default.
So, yes, more debuggers, but also abundant logging everywhere!
a) Crashes dump the program state to a file and then you load that in the debugger.
b) If the process is still running but broken, attach the debugger.
c) In my firmware I log bus fault addresses and restart. That allows me to see what code or memory access caused the error. 75% of the time it's fairly obvious what happened.
You can still do it, but you basically have to redo the whole thing from scratch. For example, instead of asking for a repr() of an object to display its value, you have to access the internal representation of it directly - and then accommodate all the differences in that between various Python versions. Something like this (note several different dicts): https://github.com/Microsoft/PTVS/tree/bcdfec4f211488e373fa2...
However, yours provides much nicer output. Thank you!
If you don’t want to go CLI-debugging both VSCode and Emacs have more integrated, in-editor options.
Many bugs can be isolated and reproduced in unit-tests. That’s just a click away from being debuggable inside a real debugger. Why use anything but that?
To me, not using a debugger to debug seems kinda crazy.
IMO it is the most value for money (i.e. time and convenience) debugging tool I've used till now. Simply write (TRACE function1 function2 ...) in the REPL and you will get a nicely formatted output of arguments passed and value(s) returned for each invocation of the given functions. Another nice feature is that a deeper an invocation is in the stack the more it is indented -- so recursive functions are fairly easy to debug too.
You can't use it for everything but its sufficient most of the time.
PySnooper looks good, but it is inconvenient in a couple of ways:
1. It prints every line of the traced function -- most of the time this is overkill and not what one needs. 2. To snoop on a function you need to modify the source file. Not a deal breaker but you still have to remember to revert this.
https://ipython.readthedocs.io/en/stable/interactive/magics....
``` import ipdb ipdb.set_trace() ```
It drops me into IPython, an interactive Python shell, with the interactive Python debugger. You can type variables and use `pdb` primitives like (c)ontinue, (u)p call stack, (n)ext line, etc. I really like it, but ofc YMMV.
edit: spelling
pdb and its variants (my favorite is PuDB) are generally difficult to use in complex, corporate projects. If you've got a multi-process, multi-thread Python project running on a remote host, you'll need a really full-featured debugger to work with it effectively. I recommend Wing IDE or PyCharm for that.
Certainly not a counter, as I'm less familiar with Wing and Pycharm's debugging features, but both of these have been helpful to me in multi-process, multi-thread python environments.
Using this tool will work immediately, while using pdb would be a bother (comparatively).
I'd add that discovering where to put the debugger and how to gate it when e.g. the issue needs warmup isn't trivial in large projects. Even if you know where to put the debugger (possibly a quest in and of itself) the exact callsite might get hit tens or hundreds of times before the issue shows itself.
And then, Python doesn't have a reverse / time traveling debugger, so hitting the callsite isn't always sufficient to understand the issue.
In all honesty, I think a more useful way to do this would be defining high-level tooling based on ebpf or dtrace, such that you can printf or snoop from the outside without having to perform any source edition.
And possibly combine that with a method to attach a debugger from the outside to a running program, using either PyDev.Debugger or a signal handler triggering a set_trace as a poor man's PDD.
[1] https://jupyter-contrib-nbextensions.readthedocs.io/en/lates...
> You can use it in your shitty, sprawling enterprise codebase without having to do any setup.
Debugging is a sensitive subject, particularly given how frustrating it can be. There’s a place for vulgarity somewhere, but I’d rather see your README provide authoritative info than crack jokes.
May I ask how the word "shitty" more truthful than "poorly written" or "complex"?
Evokes emotions people affected by a situation can sympathize with more effectively, which by virtue of establishing a shared emotional bond over a topic helps the developer convey not just the situation but the frustrations of the situation more effectively than one might expect "poorly written" or "complex" to do alone.
> Citing professionalism isn't a quip, it's shorthand for a code of conduct and long accepted practices of interaction with others.
The code of conduct isn't uniform, so it can't be used effectively as shorthand for such. But at this point we're in the weeds.
Again, no short-term revenue prospects, just a tool OP wants to socialize to make a few lives easier. If you have an objection over verbiage, that's fine, but it's an exhibition of professionalism from yourself to the OP to build a sound defense of your position as to how it would help the engineer to self-censor the description of a tool where the audience by-and-large may not care.
Up to you. My point is the engineer doesn't need to suppress who they are in this specific context, and my point to you is it shouldn't impact your usage of what looks to be an effective short-term debugging tool.
Your comment comes off as out of touch and old fashioned.
https://hackernoon.com/python-3-7s-new-builtin-breakpoint-a-...
My experience with python is limited but in my day job it's not uncommon to be tracking down stuff that happens infrequently. Debug cycles get brutally long.
Most bugs I create are for simple reasons and can be found by scanning the first error logs. If I add print statement debugging because I couldn't then they'll often be adapted into additional logging. If I use a debugger for this as my first tool and don't add logs, I'll have to do it again next time, too[1].
If the bug is not a simple one and is not a structural bug, there's a decent chance it's something debuggers deal with poorly: data races, program boundaries, non-determinism, memory errors. If it's something that can be found by calling a function with certain parameters, it's a missing test case.
So the times I find debuggers to be worth it are after I've already decided it's a difficult yet uncommon bug. So I use them with despair.
[1] If I fix it with a debugger and then add the logs, I still have to prove it gives the right output when it fails.
People create bugs by forgetting things all the time. If you believe their spin, FB snarfed millions of contact lists by accident because of that. I've troubleshot countless things that happened because of code that fell between the cracks. And on the other side of that, I could tell you about the time we went a month without logs in production because someone didn't properly test a library change, so everyone's carefully manicured logging broke.
There's nothing wrong with print(), at least that isn't wrong with a lot of other things.
Or you can be organized and competent at a higher level and just use tools. We're using computers for a reason :)
They're very visible in diffs then. And even if accidentally committed they're still easy to spot and eliminate later.
//#DEBUG 1
#include "debug.h"For example, dump this into a file and run it:
TracePoint.trace(:line) do |tp|
STDERR.puts "Ran: #{File.readlines(tp.path)[tp.lineno - 1]}"
end
a = 10
b = 5
puts a
a += b
puts a
puts b
It'll print out each line of code as it runs it. You could then parse each line for identifiers and print them out, etc. It'd be a project, but it's doable.I sometimes use a debugger to tackle with unfamiliar code, but I always prefer using trace/logging whenever possible, because 1) you can see the context and whole process that reached that point, and 2) the history of debugging can be checked into a VCS. I'd write an one-liner to scan the log file rather than setting up a conditional breakpoint. I particularly like doing this for a GUI application. A regression testing can be done by comparing logs.
With this package, it seems like I can just get my debugging via stderr.
Using debuggers tends to encourage people to fix problems without writing regression tests.
Honestly, I'd prefer "better support for print-line debugging" to "better debugger that you can set up" in most cases.
Can it be used as a context manager (`with pysnooper.snoop():`) for even finer targeting?
I was wrong. Upon skimming, it seems a huge plus PySnooper have over pdb is auto inspecting states, sparing whole lot of manual typing
I love auto-loggers like this where you can selectively capture interesting bits.
This is the basis of reverse debugging I.e capturing chronological snapshots. Would love to see a Vscode extension that allows to step forward/backwards through time when an interesting thing happens.
Both use gdb semantics which is great if that's what you're used to.
Presumably it works with Django? If yes then thrice thanks.
This project would justify an additional dedicated screen just for dumping function debug logs.
(occasionally, the low effort of print-debugging works, but if you keep having to print in more/different locations... it's a blunt tool IMO)
I have the impression that local program-specific debugging tools quickly evolve into something like functional tests that uncover issues that functional tests proper might not cover. For example, if I'm serving a ML model but I have a debug call that runs sanity checks on the data that are too expensive to run each time.
I just tried wrapping this around one of these functions in the django shell (actually shell_plus, but same thing). Just imported pysnooper, created a new function using `new_f = pysnooper.snoop()(original_f)` and called new_f on the realistic data and got a nice printout that included values and times.
Very useful.
~/anaconda3/lib/python3.7/site-packages/pysnooper/tracer.py in get_source_from_frame(frame) 75 pass 76 if source is None: ---> 77 raise NotImplementedError 78 79 # If we just read the source from a file, or if the loader did not
NotImplementedError:
You can set depth=2 or depth=3 and then you'll get a huge chunk of log. The deeper levels will be indented.
If the function is launched in a spawned process, it'll work, though you might have trouble getting the stderr, so you better include a log file as the first argument, like this `@snoop('/var/log/snoop.log')`
If the function launches new processes internally... I'm not sure.
If you try it and it doesn't work, open an issue: https://github.com/cool-RR/PySnooper/issues
@pysnooper.snoop(f"/path/to/file.{os.getpid()}.log")
Something like thatQuite apart from that, the tool described looks like something you'd use in an intensive debugging session, so it's hard for me to imagine how it would fit in with an alerting workflow.
In VSCode: built in. Press the play icon.
In Emacs: M-x realgud:pdb
Is that really more effort than this?
Why invent inferior solutions to solved problems?
cd project
pip install pysnooper # or pipenv or poetry
And it works.That's far superior to poking around in the dark trying to make an IDE see my project correctly.
If IDE authors ever figure out how to implement a test button that invokes `python -c "some stuff"` and shows me the results, I'd consider using them.
You mean like both Emacs and VSCode already does?