Hot-swapping Python 3 code
github.com
github.com
The only case where it does work is if you maintain a single reference to the module you wish to reload, _and_ the target module doesn’t load any C extensions, _and_ it doesn’t create state outside its module (very easy to do accidentally).
That’s why every autoreloader implementation restarts the whole process when any change is made, because how would you even detect that it was safe to hot reload something in the general case let alone actually hot reload it safely?
If you need to add an autoreloader to your project use Hupper[1]. If you’re weird like me and you’re interested in autoreloaders, check out the Django[2] and Flask[3] implementations or my talk[4] for more details.
1. https://pypi.org/project/hupper/
2. https://github.com/django/django/blob/master/django/utils/au...
3. https://github.com/pallets/werkzeug/blob/master/src/werkzeug...
The biggest caveat of these kinds of single file reloaders is they completely fall apart when you import between your modules. For example, if `a` does `from b import obj` and `b` does `from c import obj`, and you modify `c.py`, `a` and `b` will both have a stale reference to obj (which might be a primitive or even interned type like a string or an int, so it's not like you can mutate it as a solution).
Like, module A creates a class/function that is consumed by module B and C. If you change module A, how did you know to also reload B and C (recursively)?
I slightly break the mold of conventional Python imports - script folders don't have a single entry point, the entire user script directory tree is recursively imported (in alphabetical order) and autoreloaded on change.
Anytime I call into a module it is marked as the active module in its thread (a sort of call stack but at the module level and only for certain kinds of events).
"Call into" generally means the module is being imported, or I'm calling a callback that was registered by the module.
Anytime code creates a global object (such as opening a window, registering new voice commands, or registering a callback for some kind of event), cleanup for the object is registered on the active module. Additionally, all imports are tracked, so I have a graph of which modules import which other modules.
The import tracking allows me to reload dependent module chains (even long chains, trees, or import loops) at the correct time. The object tracking allows me to gracefully replace a module without it requiring any explicit cleanup.
This graph extends to more than just Python files. There are special functions for opening or reading files that mark those files as a module's dependencies as well, so you can modify something like a CSV and the script that uses it will get reloaded (and any module chains that depend on that script).
This all working well is tightly coupled to how the framework behaves. Scripts won't generally spawn threads or spin in long loops, they'll mostly register for event callbacks. For example, instead of a script spawning a thread or sleeping if it wants to run code later, there's a cron system that works something like javascript setTimeout/setInterval. Instead of storing something in a global, you can put it in the persistent storage layer if you want it to survive reloads (and restarts).
You can definitely do a few things that aren't reloadable, like spawning your own thread, but most things you want to do end up being reloadable, or have an equivalent or alternative API that is reloadable.
There are a few other nice touches and APIs that help glue all of this together. Generally it means you can write your scripts as short, flat, mostly declarative Python (with perhaps a few event handlers and setup code that don't need to go in a __main__ guard) and not worry about common cases for reloading. (There are also subsystems specifically watching out for edge cases, such as a user script running an infinite loop)
Honestly the most annoying part of all of this (that I do generally handle now) was reliably monitoring file changes when users put file or directory symlinks at arbitrary places in the tree (which it turns out a lot of users like to do).
I assume this still works in python 3, but I have not tried it. https://docs.python.org/3/library/gc.html#gc.get_referrers
While the docs say that this shouldn't be used for any purpose other than debugging, it never caused any trouble IME. Far more trouble was caused by classes from python's stdlib just blatantly not following the pickle protocol, not calling __init__, calling __init__ more than once, etc. I assume this is why people advised me to use composition rather than inheritance when making classes that were a list with extra behavior, a dict with extra behavior, and so on.
Edit: Well, the above primitive is not enough on its own, but if one gets the old module and the new module and explores outwardly from those to find the subset of the object graph that differs, then one can replace all references to the old things with references to the new things. This is still not enough in general, because you'll have e.g. a bunch of objects that never had a certain member set during __init__ because __init__ was different when they were created. You'll probably meet an additional can of worms if you want to use __slots__ and add or remove class members.
If you want autoreload for dev (e.g: like django dev server), better restart the entire process.
There is a 7 years old project that can help you with that called watchdog:
https://pythonhosted.org/watchdog/
It's a library that uses the fastest way available (e.g: inotify,kqueue...) to react to file changes (you can filter by types such as create, delete, write, etc) and triggers whatever code you want, passing you the metadata.If you just want something very simple, you don't have to code, you can install watchdog with its optional watchmedo dependancy:
pip install watchdog[watchmedo]
It will give you the watchmedo command, that lets you do something like this: watchmedo shell-command --patterns="*.py" --recursive \
--command='python your_script_to_rerun.py' .It is built specifically to restart things like servers as it was implemented for the Pylons Project (Pyramid web framework).
However it does not work for just watching a directory + running a command, it is specifically for watching Python code.
On the other hand, it's probably more performant than watchdog if you have a lot of files and need to fallback on manual watching.
Plus, if you are not a Python person, you may not want to use pip.
I've been writing my python code in a style that can be hotswapped. Roughly, I use global functions and pass around structs, like C. The function calls are made with `api.foo(data)` instead of `data.foo()`. Then I can just replace `api.foo`'s definition with whatever I want after attaching with a debugger. Example: https://github.com/shawwn/gpt-2/blob/e4868225382e8a475ef3c3d...
I assume there's a much better way to do this, but the class system seems to make it hard to swap things out. Doesn't "isinstance" break if you redefine the class?
EDIT: Oh, this isn't actual hot-reloading for python. Darn. I'm really hoping someone will tackle that... Cool project though!
EDIT2: Actually, this project does seem to reload the modules: https://github.com/say4n/hotreload/blob/0cf3a0b466466f99f8c7... ... so the original question stands.
By the way, you may want to use `import traceback; traceback.print_exc()` to log exceptions. I'm not sure how that plays with python's `logger` class though.
> Doesn't "isinstance" break if you redefine the class?
With importlib.reload, yes.
You might be interested in IPython's autoreload implementation, which does not break isinstance: https://github.com/ipython/ipython/blob/f8c9ea7db42d9830f163... However, you could still end up with semi-broken instances, for instance if you update
class A:
pass
to class A:
def __init__(self):
self.attr = "something"
def __str__(self):
return self.attr
> By the way, you may want to use `import traceback; traceback.print_exc()` to log exceptions. I'm not sure how that plays with python's `logger` class though.PSL logging has builtin integration with exception tracebacks. See Logger.exception [1] or the exc_info keyword param to Logger.debug/info/warning/error/critical/log [2].
[1] https://docs.python.org/3/library/logging.html#logging.Logge...
[2] https://docs.python.org/3/library/logging.html#logging.Logge...
I feel like this library is a bit misrenamed. It seems to be more like “autorun”.
Edit: here’s the ipython config I use to turn it on https://github.com/aidos/dotfiles/blob/master/_config/ipytho...
Unfortunately, it looks like it simply restarts Python each time the code changes: https://github.com/pallets/werkzeug/blob/75215972967c1c00c15... which isn't quite the same thing as reloading it.
Werkzeug keeps the socket open so you don’t observe the reload from the outside.
There are some gotchas with hot swapping, but it's designed and used in production for many years [2].
If somebody knows other languages/systems that do let me know.
1. https://elixir-lang.org/ 2. https://stackoverflow.com/questions/37368376/how-does-erlang...
I've used this technique in multiple games written in Lua, as well as a GUI app written with wxLua, and I can't imagine programming a long-running program in Lua again without it.
There are many caveats like stale module local state but they can be circumvented by designing the program from the start to put such state into a separate module which gets preserved. You do also have to modify your program's code to accommodate hot swapping in some cases but I would say the effort is repaid many times over in time saved not having to restart the program and go all the way back to part you're testing again every time you make a change.
I really wish it were easier to do in popular languages like Python but because of the way it interacts with imports I feel like your only chance to get it right is in the early stages of language design. One time I was wanting to program games in Typescript to take advantage of gradual typing but couldn't figure out an equivalent way to do hot reloading like in Lua, so I went with Lua instead. I feel like if more people knew how much benefit it brings from using a language that supports it we might see more mainstream support in new dynamic languages.
Clojure can do this via namespace live reload. ClojureScript is already doing this inside figwheel [1] environment. CommonLisp had this since dawn of time. Kawa [2] can do it, but because it aggressively optimize the code, you need to be careful. Racket can do it as well [3].
[1] https://github.com/bhauman/lein-figwheel
[2] https://www.gnu.org/software/kawa/
[3] https://github.com/tonyg/racket-reloadable/tree/master#readm...
EDIT: added Kawa and Racket.
Anyway it’s a very well studied and covered topic in the Java ecosystem as I understand it...
http://blog.ruwan.org/2012/12/dynamic-hot-swap-environment-i...
def print_name(d):
print(d['name'])
d = {'name' : 'hello'}
while True:
print_name(d)
Now you hot hot-swap it while the while-loop is running: def print_name(d):
print('hello', d['first_name'])
d = {'first_name' : 'there'}
while True:
print_name(d)
Erlang can do this. No other language that I'm aware of can. ;; define a map
(def myvar (atom {:foo 1}))
;; add watcher for 'myvar'; it will call anonymous
;; function and set old-state with previous value and new-state
;; with new value, when 'myvar' was changed
(add-watch myvar :dummy
(fn [_ _ old-state new-state]
(println "changed! - old: " old-state " new: " new-state)))
;; Run separate thread and and update :foo value to 2, after 6 seconds.
;; No scheduling is necessary, this is so it can be evaluated in REPL easily.
(future
(Thread/sleep 6000)
(swap! myvar assoc :foo 2))
;; Infinite loop. Sleep for 0.5 seconds just for printing purposes.
(loop []
(println @myvar)
(Thread/sleep 500)
(recur))
and you will see output like this: {:foo 1}
{:foo 1}
{:foo 1}
{:foo 1}
{:foo 1}
{:foo 1}
{:foo 1}
{:foo 1}
changed! - old: {:foo 1} new: {:foo 2}
{:foo 2}
{:foo 2}
{:foo 2}
{:foo 2}
{:foo 2}
{:foo 2}But, using myvar or your-like example (hot loading function) in the same thread works without any problems as well - put function in a loop, change function body and reload namespace where myvar/function exists - loop will pick that up instantly.
Let me give you a clearer example than my first one:
def fn():
d = {'name' : 'hello'}
while True:
print(d['name'])
threading.Thread(target=fn).run()
This causes a thread to be started which prints "hello" indefinitely. Clojure cannot, while the thread is running, change the execution to, say: def fn():
d = {'name' : 'there'}
while True:
print(d['name'])
time.sleep(0.5)
Clojure cannot change the value of the local variable d because it has no reference to it. Clojure cannot add the line time.sleep(0.5) to the loop because it has no facility for "upgrading" running threads.(going by https://stackoverflow.com/questions/1840717/achieving-code-s...)
But you don't have to use the gen_server framework to accomplish this. It's perfectly fine (but not recommended as there are lots of details you want to get right) to write your own code for handling hot-swapping. See http://erlang.org/documentation/doc-4.9.1/doc/design_princip... for some details on server principles.
Immutability in combination with Erlang's process centric view is what makes it possible.
There's an idea in there for a Clojure library :)
[1] http://webcache.googleusercontent.com/search?q=cache:ieBuArS...
there are also lesser known languages such as pike. although pike doesn't do hotswapping of actual objects but only allows the hotswapping of the compiled classes used to instantiate new objects. this works well enough on webservers where new objects are created for every request, and only live as long as the request is processed.
there is a platform called open-steam that emulates hotswapping by using proxy objects. when i reference an object, i actually get a proxy, that then references the real object. when i update the code, then the proxy gets updated to the new real object, and for me it looks like the object i am referencing has been hotswapped.
inotifywait -m myfile.py -e 'CLOSE_WRITE' |
while read line; do python3 myfile.py; done
It will be more responsive than a busy while-sleep loop.For those who run on linux, you can use the inotify tools to set up a watch on a file (or directory or recursive directories). inotify uses kernel events and won't require re-reading and re-hashing the file.
I found my solution superior in several ways:
- I can force an autoreload just by saving a file (no file changes needed to force an md5 diff)
- exec() doesn't suffer from the side effects of not restarting the python interpreter, clean start every time :)
- it is also quite portable and doesn't require extra dependencies like inotify/fswatch/etc. </ul>
Why do we only get hot-code-replacement and edit-code-and-continue-debugging from microsoft and sun(oracle)?
But for simple scripts being live-coded that need to be quickly "up" and hot-swapped on rapid modification this looks fine. I could always increase the delay to a minute or hour or more if the changes were only occasionally changed.
Erlang runs on a VM designed for hot reloading so the comparison is somewhat unfair to Python which was not designed for that.
https://github.com/bazelbuild/bazel-watcher
It's great, I use it all the time, especially with unit tests.
That's not hot-swapping at all.
How grimy.