Bringing a dynamic environment to C: My linker project
macoy.me
macoy.me
(defn foobar [] (undef-fn))
gives: 1. Caused by java.lang.RuntimeException
Unable to resolve symbol: undef-fn in this context
Attempting to call foobar: (foobar)
gives: 1. Unhandled java.lang.IllegalStateException
Attempting to call unbound fn: #'user/foobar
Environment info: ;; CIDER 1.3.0 (Ukraine), nREPL 0.9.0
;; Clojure 1.11.1, Java 18.0.2 (def undef-fn)
It should work after that.Evalutating this in CL (SBCL)
(undef-fn)
gives: Restarts:
0: [CONTINUE] Retry calling UNDEF-FN.
1: [USE-VALUE] Call specified function.
2: [RETURN-VALUE] Return specified values.
3: [RETURN-NOTHING] Return zero values.
4: [RETRY] Retry SLY interactive evaluation request.
5: [\*ABORT] Return to SLY's top level.
6: [ABORT] abort thread (#<THREAD "slynk-worker" RUNNING {100288C7E3}>)
[... as well as a clipped stack trace ...]
Pressing 1 will allow you to name an alternative function in a minibuffer prompt, 2 lets you set the return value of the undefined fn, etc.I guess the seamless continue without retry possibility could be useful if you are doing important side effects, but Clojure is quite functional unlike CL so it doesn't come up as much.
> In-place binary patching is based on a granularity of top-level declarations. Each global variable and function can be independently patched because the final binary is structured as a sequence of loosely coupled blocks. Another important characteristic is that all this information is kept in memory, so the compiler will stay open between compilations.
Pretty sure that feature is still very much WIP
Edit: He does mention some kind of data persistence system, but I think he vastly underestimates how complex and insidious this problem will be.
It's really not that different from doing a DB table update.
I haven’t checked these links in quite some time: https://gist.github.com/macintux/6349828
And if you don’t mind videos, here’s my talk from Midwest.io (RIP) on the philosophy behind Erlang & its VM: https://youtu.be/E18shi1qIHU
At Basho, we would load module fixes into Riak instances for customers...I hesitate to say "fairly regularly" because it's been long enough and I had limited visibility into it, but it certainly seemed like a routine operation.
The full "upgrade the entire application" hot loading definitely requires more planning than most companies will ever be willing to do.
That's how you'd handle it in Lisp. Smalltalk has something similar, but I'm less familiar with the internals for it. No restart needed, but it requires a runtime that is able to track data instances by type. You could build that onto C, though. Probably with a custom allocator that was passed some symbol indicating the structure type so it could report on instances for executing an update. Would be hard, though, to do this without invalidating pointers when the size changes. You'd probably end up with an extra level of indirection as a consequence to make it simpler (for the runtime author, not for the end user).
liballocs is an approach for providing high-level facilities such as changing structure layout to a a low-level Unix environment.
[0] https://www.erlang.org/doc/design_principles/release_handlin...
[1] https://www.erlang.org/doc/man/gen_server.html#Module:code_c...
It's... cumbersome, let's put it this way, and mind you, Erlang makes it happen at "safe" points, when the code being replaced is idle.
For Common Lisp, here's a toy example[1], and spec[2].
[1] https://malisper.me/debugging-lisp-part-3-redefining-classes... [2] http://www.lispworks.com/documentation/lw70/CLHS/Body/04_cf....
tcc does compile, link AND run the program directly on several platforms, including windows.
If you can live with C99 it is a very good fast-iteration system.
I think the codebase could also be used as a starting point for a fast linker-loader.
Saving/reloading state could be done at the framework level, and it can't work in all cases (typically when memory layout is changed between iterations).
But I think that we collectively realized how slow compilers had become only recently (something like five years ago) because it was a slow boil frog situation.
Unloading code while running is difficult. Dlclose doesn't really know whether anything is still holding onto pointers into the shared library.
What would probably work is compiling each object into a shared library and otherwise disallowing shared libraries. That would ensure the loader has a consistent model of which addresses have been patched to which shared library and give a relatively sane way to unload and replace at that granularity.
Moving symbols between object files during linking would be tricky.
There is overhead, but I don't think it would be too bad. You could also probably reduce the overhead by doing something fancier like having the daemon inspect the stack frames of other threads (which is expensive, but presumably updates are extremely rare compared to calling functions).
Put another way, assume that you solve low-level threadsafety/data race issues. You still have to design the library from the start to be unloadable; you can't unload code that isn't expecting to be unloaded without undesirable side effects.
See e.g. https://www.youtube.com/watch?v=i-inxFudrgI and earlier talks.
ORC is used by CERN's Cling c++ interpreter, Julia, PostgreSQL (for JIT database queries), the Clasp LISP VM, the LLDB debugger (for expression evaluation), and many other projects.
Or maybe more correctly the (currently) cling part?
The OP's project looks rather minimal in contrast, and very cool. I think I have had this thought too but I (semi-sadly) don't work with big enough compiled codebases to focus on it now.
Epic to see Andreas Fredriksson [1] mentioned as inspiration. I used to work with him in my distant gamedev past, and he certainly keeps knows a thing or two about performance and working with big codebases.
Edit: typo.
> It is difficult to get the boundary right between code which can and cannot be dynamically loaded. This is the same issue with using embedded dynamic scripting languages like Lua or Python—how much of your application should be written in them?
All of it.
> JIT compilation requires generating machine code, which is usually a complex process and a large maintenance burden. In practice this means shipping out to a 3rd party library for JIT, and the libraries are typically very large dependencies.
Is a c compiler not a large dependency and a maintenance burden?
Not in my experience, unless you feel that 577KB is a large dependency.
For size, ship TinyCC(100kb, also available as a `libcc.so` or `libcc.dll`) and a small stdlib (377kb for musl) with your app and you're done.
For the maintenance issue, it's not different from depending on any other library, except that TinyCC and/or Musl are incredibly more mature than any other library you are bound to use (with the exception of SQLite).
As a point of comparison: luajit is 600k, and unlike tcc, it actually generates good code.
Fork TCC to use DynASM, and layout code in a way that's amenable to tracing.
Extend LuaJIT to trace the TCC output, teach it to do this across the FFI, implement the hyperblock scheduler and quad-color GC.
Make a FreeBSD distribution which uses a one-sector Forth to bootstrap the TCCJIT, which compiles the source code directly into memory, and uses the GC for process allocation and cleanup. The JIT makes the happy path fast, until or unless it changes.
Binaries on disk being no longer a useful starting point, they can just be frozen process images. If they break, start again from source code, otherwise, incrementally compile changes into the image while it's in memory. Quitting is just writing it to SSD cache.
The libtcc API is minimal. For my needs that has been 100% sufficient and a pleasure to work with.
> All of it.
Hard disagree. Games, which the OP works on, cannot be written in traditional scripting languages due to performance. Thus, games use ad-hoc methods of either segmenting non-critical paths off that can be written (and reloaded) in slow languages, or putting reloadable code in DLLs, which comes with a host of exceptions and special cases you have to keep in mind.
What he's proposing is basically writing a clever .o loader that allows you to ignore all of the above, write plain C or (presumably) C++ code, and have it automagically show up in your running process when you hit the compile button.
Sounds pretty compelling to me.
o Call any function in your program in an interactive read-evaluate-print loop
o Visualize function compiled sizes
o Visualize function references
o Introspect on program data
…and more things I haven't thought of yet!"
Brilliant idea! I'd love to see this linker when it is complete! I hope you succeed wildly in this endeavor!
(Weirdly it was on by default and meant that your builds were non-distributable, as the executable depended on dynamically loaded object files on your system.)
My own attempt at solving this right now is a C variant that's amenable to SLIME-style incremental development, and I'm having fun doing it, but it'll be a while before it's useful. This linker idea is really smart. Being able to drop it in to existing C projects is a big win.
People under Unixen have been using ccache forever.
* all pointers in all new objects,
* all pointers into the old-range,
* and all live on-stack values into their new on-stack locations.
Did I forget something? using volatile-only vars would solve only the third problem.
It's not to say that this would be a reasonable thing to do, but that these sort of shenanigans may be done (e.g. to pass data across an excessively narrow API callback)
* all arrays of such structs.
Dealing with deleted fields is more or less straightforward, but what should be put into newly added fields is anyone's guess.
Edit: Or just live with the fact that certain kinds of changes require a program restart.
We've been using Edit and Continue for games at my job. It requires some care and attention to detail to keep it working, and it definitely has its faults, but it's pretty effective for constant-tweaking, IMO.
In nim's case when hot reloading is enabled functions between modules become pointers.
I'm picturing a long-running server, and people wanting to run experiments on it. Like, the next time an HTTP GET comes in, please use my new code.
In my head, I had imagined that there would be a versioned method table for each function (or at least each function that I wanted to be versionable.) I would compile my new version of the code, and then dynamically link it in, and then use some experimentation flag to say which version of the system the execution should follow.
So, it's a bit like Git. Each object is versioned, and you can request a snapshot of the entire tree, at a particular version. And because you can walk the tree of function calls, you can have a root version, etc.
It's super cool to see people experimenting with this.
[0] https://joearms.github.io/published/2013-11-21-My-favorite-e...
So basically, in your example, change the code that handles GET and then trigger the CI/CD. Your code will be compiled, containerised then a rolling update is started to replace the old version of the running web server with the new one.
I'd like the server to be able to execute a code path which mimics my Git branch, and on the next request to execute a code path which mimics your Git branch.
If you and I have two different versions of f(), I'd like them both to be available to be executed, based on some flag or condition.
As we merge our branches, f() will collapse back into the CI/CD pipeline you imagine, yes.
it aims to replace numerical python and matlab for high performance interactive numerical computing.