Why is it so hard to write a scaffolding tool? (2019)
jfreeman.dev
jfreeman.dev
I primarily work in Python, so that is my reference. Django provides its own scaffolding tool, but it is too opinionated. Flask provides no scaffolding. Framework specific tools are rarely sufficient to meet the needs of an organization. Regardless of language and framework, choices must be made, and when you have a team of any size, everyone needs to agree on the same standards. We all know this.
If one limits a scaffold to the lowest common denominator, the base standards of an organization: directory structure, configuration, CI/CD, testing, library management, documentation, linting, packaging, container configuration, i.e., everything that precedes actual functionality, scaffolding is not hard, and is very helpful.
But as soon as one tries to drill down into specific functionality that is not common to all projects, then scaffolding quickly becomes difficult and unwieldy.
IMO, the key is to not think of scaffolding as a tool to speed development, but as a tool to enforce those standards of an organization that are common to all projects in a given language/framework. Resist the temptation to go further. If your scaffold needs more than a half dozen or so inputs, then you are probably expecting too much of it.
Scaffolding breaks my mental model for reading code. This often results in more mental overhead and slower/worse outcomes.
When scaffolding generates code for me, I read it as psuedo-blessed code. Rather than starting from a clean slate and working through my mental checklist, I must interpret what the scaffolder created. Figure out what I _actually_ need and figure out what needs changing.
By contrast, without a scaffolder, I feel like it's much easier for me to build a mental model around the problem I'm solving - focusing in on the more important components first. Yes, I know that I'll probably need 20 files to complete a feature. However, I don't need those files right now.
It's very silly.
"No," said the zoomers, "we want to write XML, build scaffolding tools and eat shit."
"Your language will never be a LISP," was the God's reply.
If I start a new Python project I have a short checklist I go through
1. Poetry (dependency management) 2. pyenv and pyvirtualenv (execution environment management)
Just these two are somewhat of a headache to get going. Once they're up, though, they do the job they're supposed to in an admirable fashion.
How would a language with homoiconicity and macros handle this?
But let's take Common Lisp for an example and why I don't feel the need for a scaffolding tool. Well, actually I do need _some_ very basic scaffolding, namely: create two directories called src and tests + add two very short project files with my names and project names embedded in them, along with a couple of package files. Nothing fancy and certainly nothing that's particularly problematic. Everything that's part of the project code (such as package description with imported symbols and such) maybe be done by using a macro from an external library (ex: certain functions of some library that I import in each of my projects).
But you are right anyway, not everything is a language-specific problem here, only the stuff that gets repeated in code, from one project to another. In CL that includes documentation and the testing framework (from the article's points). Well, the build/project description is just a macro too. Anyway, I didn't mean to diminish the major points of the article in any way. The author is right about the fact that generators are not aware of the user's changes and that can be problematic at times (as with his license and license metadata example).
Okay, you might say, Sindre is an exceptional case. But on any new project, you might forget to include the correct set of files or screw up how they are set up by not filling in or updating a copyright notice, project version number, repo link, etc. Sure, you can cut down on the number of files you start your projects with and the fill-in-the-blank values they might need, but that only gets you so far and it reduces the initial usability of your project.
A scaffolding tool helps maintain consistently high quality across projects. For example, my Yeoman generator automatically fetches the latest version of my favorite CLI helpers and testing framework and other dependencies, so I never accidentally start a project with an outdated, potentially insecure codebase.
https://github.com/sholladay/generator-seth
Not everyone needs something this fancy! Another commenter mentioned editor snippets, which is probably what I would use if I couldn’t have Yeoman. But when you make new projects on a regular basis, scaffolding is 100% the way to go.
Of course, you don't always get to choose or modify everything about your environment, like how git works, how your CI is configured, etc. But it's not a bad thing to admit that scaffolding is at best a papering-over of deficiencies in these tools, and not a desirable thing in itself.
I'm not familiar with Yeoman, but like other necessary but dirty things we do, we should at least try to encapsulate it and do some implementation hiding. You wouldn't generate language bindings from protobuf and then manually modify them--and the same kind of discipline should hold for these scaffoldings. You should be able to repeatedly regenerate all their artifacts (even if they get checked in) with confidence--and once you've done that, you have effectively removed the "scaffolding" and replaced it with a build tool.
---- [1] As a bit of a tangent, there's an even worse sin lurking here, which is that you have the same fact but need to capture it twice (or more) in different places: in package metadata, however you do that in your language; possibly in project metadata (if however you host your project doesn't discover it); and in the text of the license that must be distributed. If you don't rely on a single source of truth for this, you're asking not just for an ergonomic problem with updating your scaffolding, but an actual, material inconsistency, like when the COPYING file contains a different version of a license than the one that's referred to in your .gemspec. I point this out as a conceptual detail, not because I think it's an actual frequent problem.
If that's the case, then I know of zero languages/ecosystems with tools that are not deficient in this manner.
https://github.com/norskeld/arx
Curious what everyone else uses? Or you still start each project with empty folders and manually add things?
One idea I've been implementing for fun lately is a git repository where each branch represents a combination of optional features that can be generated. E.g., the branch `aad/aspnetcore/hotchocolate` is the starter code for an ASP.NET Core GraphQL server that uses Azure Active Directory for authentication. Whenever I make a change in the aad/develop branch, it automatically merges to aad/aspnetcore, aad/aspnetcore/hotchocolate, aad/aspnetcore/rest, etc. The problem is there are a lot of merge conflicts as you can imagine.
Very few of my projects existing templates closely enough that I feel like I could get away with using a scaffolding tool. When I first used create-react-app, I spent a half day ripping apart the results and adding the parts I liked to a new project, incrementally.
I feel like copying build scripts, configuration, and project structure from existing projects you have experience with is no worse than using a scaffolding tool. Copying isn't a bad thing.
[0] https://bitbucket.org/Nzen/zaftig-offal-hamisha/src/master/s... (file templates)
[0] https://bitbucket.org/Nzen/zaftig-offal-hamisha/src/master/s... (plugin config)
[1] https://maven.apache.org/archetype/maven-archetype-plugin/us...
Scaffolding tools in particular have the problem that once you union all the different requirements the resulting scaffolding is insane and impossible for anyone to understand. Some people need to have C modules in whatever non-C language you're working in. If your scaffolding doesn't include that they need to do a lot of work. If it does include that, a whole bunch of other projects now have a lot of crap they don't need, which may be difficult to extract properly from the resulting scaffold. Multiply this process by dozens of features and you end up with something that is basically guaranteed to be a monstrosity that nobody can understand, to be missing critical features, or most excitingly, both at the same time.
The natural solution many will jump to then is "Oh, we'll make it configurable!" Now the scaffolding authors are responsible for testing the exponentially-growing state space of possible configurations, and the scaffolding, whose whole point in life is to be simple, is now complicated. And most of the time, if you discover your scaffolding configuration was wrong after you've already started work it's non-trivially challenging to fix the config and re-run. This basically turns the scaffolding into a non-Turing complete domain-specific language. Those have their uses, but "operating in a fundamentally Turing-complete programming space" isn't one of them. It's also arguably an inner platform, replicating what the language already has.
But jerf, you protest too much. You seem to have disproved the utility of scaffolding in general, but I know that I've used it before successfully!
Yes, you can solve all these problems with basically the Gordian knot solution. You just provide scaffolding anyhow, and it can be helpful. But my proposition here would be that you basically just accept that any given scaffolding can't be The Solution For Everyone, you don't go crazy trying to provide endless configuration, and you just accept that it'll be missing features for some people and have unnecessary features for others. There is absolutely value in simple scaffolding even despite what I say here, especially for things like "Here's how to start getting into Ruby on Rails!" or other similar frameworks where it honestly may not be practical to expect someone to be able to bring up a new project from an empty directory. My point here is more to not jump into the Turing tarpit of spending tons of time and resources trying to 100% solve this problem everywhere and for everybody. Grab some quick wins, be aware of the risks, be aware of the infinite regress of feature requests that lies down this path with ever diminishing returns, and move on in life. There's plenty of other situations like this in the programming (and engineering in general) world.
My personal opinion is that the disadvantages of this approach still far, far outweigh the advantages... but let it not be said I don't recognize that there are some advantages to it even so. :)
What about monkeypatching makes Rails tolerable? Why do "the disadvantages of this approach still far, far outweigh the advantages"?
That's not trivial to maintain, but neither is the tool.
The combinatorics problem with "options" on generators makes it pretty impossible. I'm sure we've all search for something like "typescript react Apollo jss boilerplate" and crossed our fingers that it was updated in the last year. Imagine changing just 1 of those things and how many code decisions change.
I don't want to do that work but I guess the knowledge compounds, knowledge about intricacies of various JS tools and how they relate to one another, and no one else want to do it.
The amount of customization that can happen in JS world makes scaffolding tool really hard.
I think today's scaffolding tools (eg. cookiecutter) are just a fraction of what they could be. As the article points out (at least that's my reading), there's tension between expresivness of the tool and ease of use. Incrementality and composition are two ways of trying to handle that.
My approach is interactivity, using a GUI instead of command-line tool, and instant preview. Within a graphical UI, the user can more easily grasp much more complex configuration scenarios and explore options, and they get instant feedback on how that will look in practice. My bet is that this will allow adding a lot more options while not bogging down the user the way the command-line tools have to.
From the scaffolding perspective, this means that every time you want to change something in the scaffolding, you rebuild the entire thing. The Linux kernel follows a similar tack; if you make changes you rebuild the entire thing.
We need something like a "gradually typed version control system".
I know I'm asking too much but I wish I can just update the template and have all the files generated by it updated too.
This was a core tenant of the code generation. Additionally, you can also edit the output fields and still keep support for regeneration after template updates.
/folder/subfolder/
/file.txt
some text
/some.sh +x
#!/bin/sh
…
I could use a zip/tar file, but repacking it on every change would be cumbersome. There must be something simpler.Augeas is your guy here.