The Build is Always Broken
gbracha.blogspot.com
gbracha.blogspot.com
The problem is guaranteeing correctness. Correct cache invalidation is Hard(TM). Therefore, the CI system builds from scratch in order to guarantee a clean, correct and reproducible output from a given snapshot of the source.
As for 'live programming' against a running image, this write-up starts with the implicit assumption that this is always a good thing. While it can be useful on occasion, the reason for restarting the application after changing its source code is the same reason we favour functional, immutable styles: mutable state is a pain and it's best to push it out of the thing you're working with if at all possible. Modifying the running image risks ending up with internal states which no single snapshot of the source could ever have yielded.
Doing an `npm run dev` type of command in a reasonably set up JS project is very much this staged execution model the author talks about. Everything gets cached, on disk and in memory, only changes are recompiled, and reloads are fast, partial, and often hot-swapped. It's quite close to the maximum level of "liveness" I can imagine.
It's kind of amazing how JS took over older and more established languages in this sense - and we're still not content (judging by how popular it is to complain about webpack). This is great.
Plus webpack got a lot better even in that regard, and there's alternatives like Parcel pushing the curve. Web devs sometimes don't realize how good they have it :-)
The hypothetical build system can, whenever files change, run the build again (in a container), execute tests (complaining when you break one), or let you keep the debugger open while you're editing (best attempt made to fast forward).
This is probably better suited for modern languages with proper modules and incremental builds, like Rust, C++20, Java9.
Common Lisp had true liveness for decades; it's almost mandated by the standard, it's designed in such a way that you can actually recompile a function, or even redefine objects, adding or removing properties or methods, at fucking runtime without restarting the application. Objects already instantiated will remain so and will be updated to reflect the change. You don't restart. That's the true "liveness", but seriously you don't get it until you try it, and once you get it you become depressed because you realize it doesn't really exist in any mainstream language.
JavaScript is some decades behind this dream, even with stuff like Webpack.
I’m not an expert, but I think the way it works is that Webpack provides a way for JS modules to register callbacks for what should happen when Webpack detects changes in them and recompiles them. React then provides such a callback implementation for component modules that causes the individual component to re-render. Webpack’s development server also provides a way to inject new code into your app running in a browser.
Things get a little more complicated if you want to hot reload things that deal intimately with state, like Redux reducers or middleware. AFAIK people usually just have the browser do a full page reload when changes are detected in those modules. I think Webpack will actually automatically trigger a full page reload if it detects a change in a module which does not explicitly declare how hot module reloading should be handled.
I only mention React and Redux because it’s the only major stack I’m at all familiar with; I don’t know if other stacks provide similar functionality.
Erlang provides this as well, as do many dynamic languages. However, aside from certain Lisps and Smalltalk that are image based, and a few others, liveness and source persistence are typically traded off for each other. I can make a live change to a Ruby system if it is set up to provide a REPL, but unlike a live Smalltalk change I won't then have the source code for the changed but available the same way it is for the rest of the system. As a result, in many languages live coding features are mostly a tool for experimentation (usually in nonproduction environment!) on code that will eventually be incorporated into traditional source tree for a from-scratch build.
- Zope (python app server) had a feature where you could debug a web application (in production) in a private session, so all code modifications you did in your session were private until you committed the code.
This can be done in a number of other common languages today such as Java (JVM languages) and C# (probably all .NET languages). Most IDEs I've seen for those platforms support it.
Substantial enough changes to the code can require an application restart, but most changes you might make like changing the implementation of a function, adding new functions to a class, etc., will not. Hot Code Replacement (HCR) has been supported since Java 1.4: https://wiki.eclipse.org/FAQ_What_is_hot_code_replace%3F
Fully dynamic languages like Lisp and Smalltalk permit a greater degree of this than statically typed languages do, however, since they don't have types and a type system to wrangle with. When I've used it, hot code replacement supported most of the changes that conceptually make sense to support.
I do procedural music in JS/web audio, and my projects are set up so that changes to the logic and instruments all happen live, while the music continues to play.
How is that different from what you're describing? What specifically can't webpack/HMR do?
It's not just build time, it's the indirectness. I'm seriously contemplating dropping it for a next project and just writing straight JS, whichever dialect (es6?) has the widest adoption.
For most people working on larger projects, though, the advantages of real dependency management (however broken NPM is) and transpiling JSX or TS is worth the compilation, which you get used to.
Living without a bundler was alright.
At my last job we checked Google Analytics and almost 100% of our users were on greenfield browsers, so we moved to that and had no issues.
The longer you run without a restart, the more dependent you become on the current state. And the more dependent you become, the more likely it is that restarts will be catastrophic due to lost implicit state.
There was (is?) a live code editing capable browser which updates the source code. My failing memory seems to remember creating a project and mounting directories. Chrome? Great idea, but IIRC, also very brittle. And because it had it's own notions of "project", there was some impedance mismatch with the IDE.
Obviously even a compiler on a server can be caching and clever and not rebuild everything (if you use a source package manager like cargo then you may run into this, but if you have a binary package manager like nuget then you don't - each compilation unit is either required to build or it isn't, and external dependencies are always just fetched).
I believe that's why he goes on about 'live-ness' which is a very smalltalk concept of a "live-image" of the running program. I doubt what he means is directly translatable to the Javascript world, although perhaps a running browser with JS in it can considered to be the live-image (like, say, the Clojurescript/React/Figwheel way of working).
But as you note, in the JS-dev world, very often you do need to reboot and start again (say if an external resource like a CSS has changed).
[1]: https://squeak.org/
Modfiying web apps from the browser seems to be a rare case, interestingly. I think there were a forays[1] made in that direction, but the need for external tools to shoe-horn a half-shod module system onto JS was the final straw for that. Never mind the convenience of some nigh-essential language features (SCSS, ES6).
At this point, you can either copy over the difference or just replace the whole production application. I'm not sure which would be faster.
I don't think liveness is important for a production system. I prefer production to evolve in discrete chunks. Liveness makes sense for development.
I suppose one other benefit of a live system is that you could deploy an update without without restarting the application. But at some level, you would be restarting part of the application, and you are just shifting the update logic from the networking layer to the function call layer or object layer or whatever minimal layer of swappable component you can deploy.
Or is the Smalltalk idea that you never update your software?
At worst the user would have to rewrite their code from time to time.
So why should I get excited about the ability to live edit when it's held up as bad practice in one of the few areas it's possible.
As many comments note, the proposal of this post seems a disaster regarding reliablity/predictability.
I'm not sure why this blog post is written like these things don't exist.
EDIT: on a re read, I'm a bit off the mark. The author seems to want the compiler or interpreter that reads the code to automatically process the dependency graph of what changed (at the finest grain possible) and take appropriate action.
The company I work at has tools for automatically updating our dependency graph relationships, and then commits/diffs execute commands against the dependencies of modified targets as part of the code review and CD process. This is pretty close to what the author is suggesting.
I still think this blog post is written a bit too dramatically.
This is way less live than the author hints at. Let's take Emacs as an example -- it can be extended using Emacs Lisp. If I write such an extension in Emacs Lisp, and I find that one of the functions isn't right, I go to the function and hit Ctrl-Meta-x and then the function is redefined in the running Emacs, and so I can try again to see if it works now.
Or take the Fish shell as an example. It provides a method to define a function that is automatically saved. So you write a shell function "foo" and you try it out (in your interactive shell) and you see it's not right. You go "funcedit foo" which pops up an editor with the function in it. You make your change and save and exit the editor. You run the function again to see if it now does what you wanted.
A similar experience can be had, I guess, using a Java IDE when working on a web application, if your stack supports hot reload (with JRebel?). You run the application in the debugger, you find it's incorrect, you make a change to the affected method, you save and the IDE hot-reloads the new method into the existing web application and you just try again.
But for all of the above examples, it is still the case that there are two processes -- one redefines things while developing, and the other process kicks in when you turn off your computer and then turn it on again.
I think in Smalltalk, you make code changes using something like the "hot reload" thing, but the new code is automatically saved to "the image". And when you want to run the system, you open "the image".
On the other hand, I would argue that the idea of a build is an entirely valid and useful construct, it merely represents a point in state space against which test cases (and production) must run. It is impossible to get away from that state. Some tools can make managing it easier, and some make it virtually impossible to manage. Imagine that we discovered a magical halting oracle and that we could compile everything instantaneously, we would still have to figure out what combination of states were valid, and we would call that the 'build'.
I was just thinking that a running application is a little kernel of Code (the stuff that gets checked into source control) and big wrapper of State (what the user is currently doing with it).
The build system view is that the Code is sacrosanct, and what we need to focus on is checking the right stuff into source control and trying to ensure the latest commit is always correct and self-consistent. We should always be able to throw away all the state and rebuild everything from scratch.
The live coding view is that the user is more important than the Code, so we should focus on letting them get stuff done as effectively as possible. And programmers are users too! Therefore we should be able to modify the code without losing any of the user's state.
Both views are correct and what's actually needed is a good balance between the two.
Also, the more we incrementally patch a live environment the harder it becomes to specify what, exactly, is running in production.
If you don't need modularity then as others point out Smalltalk solved this long time ago.
Also as others pointed out, gradle, bazel etc solve most of the issues mentioned (incremental builds, distributed (!) caching of build resources).
I've used live reloading in HyperCard, where your application (the "stack") is always running and you can interactively edit the code for each UI element. I think this can work really well, but only under certain constraints:
- You have a very clean separation between code and data (in HyperCard, the data is the set of cards and backgrounds, and the contents of all the fields; the code is the event handlers attached to each object).
- The data format is stable.
- There are clear checkpoints where your data is stable and no code is running.
Those are sometimes the case in some applications. Arguably, all of them are nice properties and therefore good targets to aim for.
But there are plenty of normal scenarios that don't meet those conditions:
- You have an expensive long-running task. Trying to change the code half-way through is a fool's errand. You need to re-run it from the start, and/or find a way to split it up into smaller steps. Unit tests and a good incremental build system are your friends here.
- You're working on the core data structures. Again, trying to update the code while all the data is still in memory is a bad idea. Just use a good build system and try to make the compilation and startup time as fast as possible.
- You're changing the navigation system, so the user's current position in the app won't make sense any more. You'll need to reset back to the front page (or whatever).
For that last one, "apply changes" in C# or Android Studio will generally do the right thing, even though those are mostly traditional build systems.
There isn't really a hard distinction between a build system and live coding, it's a spectrum. I'd argue further that a build system (in the traditional "make" fashion) is more fundamental, because you can achieve anything with a build system (even if it's tedious and inefficient) whereas some tasks are just intractable with live coding. A fast, reliable and correct build system (like Bazel, as several people have mentioned already) is the best foundation for everything else.
Depending on how clever the build system is, I have been burned by "build hygiene" in the past. Some would go so far as even nuking the source repository and cloning it again rather than updating or running "git clean" etc, to be absolutely sure that no a single byte was left from a previous build (Git is pretty good at this, but other version control systems not so much).
- Use an incremental build during development
- If you hit weird errors, try a clean build
- Always use a clean build for releases
Make is vulnerable to weird errors, as are newer systems like Gradle. Better build systems work hard to be 100% correct. Bazel is the best example I've seen.
I'd say working with an extremely reliable build system is just as eye-opening as working with a good live coding system. If you never have to do a clean build, and the incremental build is fast, that's most of the benefits of live coding right there.
You could maybe distinguish between "incremental build while running" and "no build, just interpret the code directly". But I don't think there's a major difference between those in practice, it's all just a question of how quickly code edits are picked up.
That's what drives me mad about all these configuration heavy (declarative?) systems. Just give me a debugger and an imperative language and I'll step through it myself when there's a problem. It's what is happening anyway so it's a really leaky abstraction.
Indeed, the building blocks for fast iteration are there, but as always in the 40-years-old ecosystem of C, no coherent solution has emerged and dominated the landscape.
The system I have been using is: https://zuul-ci.org/
Not exactly what the main focus of the article. But when the article talks about "the build is broken" then using a CI system which prevents that is good. It doesn't prevent an individual developer from breaking their local build of course.
OP worked at Google as one of the creators of Dart so I assume he had a chance to encounter Bazel. Reading the post I got the feeling that these takes are partly informed by his encounters with die-hard Build System Believers.
I enjoy live coding, but the idea of checking in an "image" of the running system sounds horrible. You want to clearly distinguish between the key stuff and all the incidental state. The key stuff is what gets checked in, and that's your build system right there.
(I don't ask just for the sake of arguing, I'm interested in the answer! And I don't see it addressed anywhere in the article.)