In praise of --dry-run
gresearch.co.uk
gresearch.co.uk
Here are two other things I've found:
1. Regardless of where you push your run/dry-run dispatch to, the underlying "do it" logic really needs to be factored properly for individual side effects (think functional). Otherwise, you inevitably end up with loads of "if dry-run / else" soup or worse, bugs _caused_ by your dry-run support.
2. You still need safety nets for when someone is doing dangerous operations without dry-run. Pointing and calling [0] is a great trick for this. For example, say your tool is about to delete X, Y, and Z. Instead of a simple "Yes" confirmation which the user would quickly end up doing on auto-pilot, you could have a more involved confirmation like "Enter the number of resources listed above to proceed" and then only proceed to delete if the user enters "3".
Very curious to hear about design idioms folks have come up with!
What I don't like is passing those as params as args. Currying this stuff around is very error prone and cumbersome. I prefer to set environment variables for DRY_RUN and DEBUG (with debug accepting higher integer levels for more debugging).
This works wonderfully for me since I never have to worry I didn't pass something along correctly, I'm always asking the global set at runtime.
Using environment variables for the user to configure their environment or pass information into a program seems normal, but code passing arguments to other bits of code in the same language with environment variables seems odd.
It's definitely my guilty go-to-hack when I'm not up for refactoring everything to be more functional and/or take a dry_run function parameter everywhere.
I end up writing lots of library code. Sometimes that library code is called from within a command line utility I created, sometimes it's called from a web service, sometimes it's just a small driver script because I'm not developing or testing the utility, but the library code itself.
I can make that include to set up the shared global and try to make sure it's included in all instances and all ways I want to use the code, or make all the call sites resilient to it not existing, or I can just use the included OS mechanism for doing this and since that's always available, I get it for free.
Also, dry run mode isn't necessarily something you want set in a config. It's generally something you run once or twice prior to running for real (the normal case) or while in development/debugging. It's not something you would want to set in code and accidentally forget and push live, and generally a good dry run mode will look like it succeeded without actually succeeding, mocking responses that would fail along the way, because you aren't testing one small thing you're testing a workflow of some sort generally which has a few steps.
That said, I fully admit the trade-off might go a different way for different languages. Using a compiled strongly typed language may mean there's enough bits to check that you need to write a debug/dry run helper function to make it convenient, so there's not a lot lost by requiring setup in that as well. But for something like Perl (and I assume Python and Ruby and JS, to almost the same degree) where I can do:
warn "Calling out to foo() with args: " . Dumper($args) if $ENV{DEBUG} and $ENV{DEBUG} >= 2;
foo($args) if not $ENV{DRY_RUN};
and it will be completely valid, obvious and idiomatic with zero additional work, there's a real draw to using environment vars for these two specific cases (even if not for all config).So you can do something like nuke $ENV{PATH} in your perl script and it'll apply to any subsequent child system() calls.
The contents are private to a process, and the only method to modify them from outside the process is to have write access to process memory and change the pointer to the memory block.
When you start a new process, ultimately your calls turn into execve family with full set of arguments, i.e. actual binary to replace your process with, arguments, and contents of the environment it's going to use. The wrapper functions that don't ask for new environ value simply copy the current one.
In the instance of dry run, I think it's far more important that it's correctly seen and honored than that globals should be avoided. I view dry run mode as a contract with the person running, and if I can avoid having to pass arguments along every time, and avoid having to got change the arguments of all the utility functions I call and then all their call sites, then that's a win because if that needs to be done to correctly honor that sometimes it won't be done and some worse solution will happen, if anything.
> but code passing arguments to other bits of code in the same language with environment variables seems odd.
I don't pass the environment variables around, they exist as part of the program state. if foo() calls bar(), I don't pass the environment variable to bar, it exists as a global flag, I just check for it in bar(). That's the benefit, it's a global and set at runtime.
The other benefit is that as an environment variable, it's inherited by child processes. Even if I system out to another utility script, I don't have to pass a debug flag their either, it's inherited as part of the environment the child runs in.
1. Passing around a configuration or context (similar to your gripe around params)
2. Referencing some kind of global configuration (a la env vars)
3. Referencing a local or scoped configuration (for example, a method of an object can check instance variables that dictate behavior)
I personally prefer either a passed context or a local configuration; I find both easier to test in isolation. Global contexts have their uses, but tend to become problematic when they clash with other libraries or tools that may also be present in the execution environment.
That said, I'm not married to it, if I saw something that seemed obviously better, I would switch. I also suspect that different languages may make one approach easier/better than others based on their capabilities, idioms, etc. In many scripting languages, accessing an environment variable is extremely easy. In some compiled or more strictly typed languages, the access and conversion to the expected type might be cumbersome enough to do on site that it's worse, and if you are standardizing in come parsing routine, that might tip the benefits in favor of some global context that is used instead.
CQRS (Command and Query Responsibility Segregation) is an alternative to a simple CRUD-based database interface. CQRS separates reads and writes into different models, using commands to update data, and queries to read data.
Commands should be task-based, rather than data centric. ("Book hotel room", not "set ReservationStatus to Reserved").
Commands may be placed on a queue for asynchronous processing, rather than being processed synchronously.
Queries never modify the database. A query returns a DTO that does not encapsulate any domain knowledge.
The models can then be isolated, as shown in the following diagram, although that's not an absolute requirement.[1]
It’s a really practical treatment of CQRS & ES with JS examples. I had heard about CQRS & ES on HN for years but I had never seen any system that actually used it or any practical examples so this book is amazing cause it takes you through how to build a system using them.
do a bunch of stuff
make a plan
if dryRun {
plan.Print()
exit()
}
plan.Execute()Everything on a computer has some side effects, and the purpose of --dry-run is to stop execution without certain effects that we care about. It's impossible to eliminate side effects entirely, so this is not a goal.
https://www.youtube.com/watch?v=ZLNduq2pnN0
At some point it doesn't matter if there's "the guy who does that", because the right choice is to let "the guy who does that" deal with the consequences of their own crazy setup.
https://en.wikipedia.org/wiki/Poe%27s_law
Use a winking smiley face emoticon ;-) or /s
I feel like the context for this comment thread was lost. Just to repeat it to forestall any further discussion in this subthread on non-plan dry-run, this is how I saw it:
* Igor: a good way to do dry run is to make plan
* * Me: first time i have seen 'plan' as explicit step for dry-run is with Terraform
* * * fickle: rsync is where i first encounter it [it here is not plan, it is dry-run]
* * * * me: really? rsync has 'plan'?
* * * * * john: rsync has dry-run. read manpage
i.e. Igor and I are talking about explicit plan (which is interesting model) and you guys are talking about dry-run.
--dry-run is one of the best motivations for using granular, extensible effects
when you constrain literally every bit of IO your program is allowed to do, you can now 100% know you've stubbed them all out when doing a dry run
Another important property is idempotency. Especially if the script involves network requests or moving files around, you'll want to reach the goal by re-running in case something breaks half-way.
- cancel
- undo/redo
- progress bars
Progress bars with accurate timing are notoriously difficult, or even impossible to get right. But even regardless of timing, having a progress bars that really shows progress and doesn't freeze is hard. Every slow operation has to have some sort of callback mechanism to update progress, and you have to know the number of steps in advance.
Undo requires, for each operation, to know how to roll back. You also need to have every operation go through the undo stack. Another option is to have an efficient snapshot system, which may be just as hard or even harder.
Cancel is actually the hardest to do right because it combines the difficulties of both the progress bar and undo. You have to have a way to interrupt a long operation at any time and get back to before it started. And because "cancel" is most often used when things go wrong (ex: disk full, bad connection, ...) you have to be very careful with error handling.
It's an architectural step a lot of people skip, and by doing so you often end up in a situation where the only way to make things faster is via caching, and chasing caching bugs for the rest of your tenure.
One of the less appreciated aspects of model-view-controller is that usually it demands that you plan out your action before you do it - especially when that action is read-only. Like a cooking recipe, you gather up all of the ingredients at the beginning and then use them afterward.
By fetching all of the data early, you reduce the length of the read transactions to the database, allowing MVCC to work more efficiently. You also paint a picture of the dependency tree in your application, up high where people can spot performance regressions as they show up in the architecture - where there's a better opportunity to intercede, and to create teachable moments.
To make dry-run work you are best served by book-ending your reads and your writes, because if the writes are spread across your codebase how do you disable all of the writes? And if any reads are after writes, how do you simulate that in your dry-run code? It won't be easy, that's for sure.
The problem is that we tend to write code stream-of-consciousness, grabbing things just as we need them instead of planning things out. This results in a web of data dependencies that is scattered throughout the code and difficult or impossible to reason about. This is where your caching insanity really kicks into high gear.
To my eye, Dependency Injection was sort of a compromise. You still get a good deal of the ability to statically analyze the code but you can work more stream-of-consciousness on a new feature. But it does rob you of some more powerful caching mechanisms (memoization, hand tuning of concurrency control, etc)
Certainly it's not something that can be applied for every use case, but when it can, it has so many benefits - an audit trail of what actually happened, ability to undo, and a common structure and approach for what dry runs should do (just output the log)
I'd like to one day see a global flag on operating systems to dry run nearly all commands that changed anything.
https://android.googlesource.com/platform/frameworks/base/+/...
Seems to be followed most of the time.
Same here. In addition if there's never a use case for running in production (e.g. at one point we had a command that reset a dev environment to a clean slate) then the command will actually specifically check for the production environment and refuse to run as an extra failsafe.
struct RunToken;
fn do_the_thing(_: RunToken, ...);
The general approach works for just about any statically typed languages (although null is the enemy).You can make it a bit more tedious to generate that RunToken so that a `do_the_think(RunToken, ...)` looks more natural. But the idea is simply that if the only place you create a RunToken is when parsing the --no-dry-run flag you can't accidentally do_the_thing when you move the code out of the `if is_dry_run` block without noticing.
Of course this is a simple way to get fairly reliable dry run. I do agree that having seperate plan and apply steps is approach is the best when you can do it.
sudo apt-get purge login
..
WARNING: The following essential packages will be removed.
This should NOT be done unless you know exactly what you are doing!
login
0 upgraded, 0 newly installed, 1 to remove and 303 not upgraded.
After this operation, 1,212 kB disk space will be freed.
You are about to do something potentially harmful.
To continue type in the phrase 'Yes, do as I say!'
?]Uses `--go` to confirm.
To actually carry out the obliterate action, you have to add the -y option.
https://github.com/Distrotech/hdparm/blob/4517550db29a91420f...
<div dangerouslySetInnerHTML={{ __html: "Some raw HTML" }}></div>
https://reactjs.org/docs/dom-elements.html#dangerouslysetinn... 1. do-step-one --dry-run
2. (do validation)
3. do-step-one --real-run
4. (do validation)
Many commandline parsing libraries assume you're doing booleans, and I think --dry-run and --no-dry-run is confusing. And, internally, you have a boolean flag so there's always the possibility of some code getting it backwards.Internally, I'd like an enum flag that's clearly dry_run or real_run, so the guards are using positive logic:
switch(run) {
case dry_run:
print("Would do this...");
case real_run:
do_real_thing();
}While I kinda like this idea in principle, I haven't really seen any CLI apps that require an explicit option for both modes, so it might be a bit of unexpected UX for people.
I really came back to this post because I was struggling with Gradle and if there is a more perfect illustration of how a dry run mode could help than that, I don't know what it could be. Gradle builds are a fucking mystery, every time, and there's no excuse for it.
—i-acknowledge-this-is-dangerous-and-i-gave-this-more-than-cursory-thought
A whistlestop tour of the technique "insert a lightweight API boundary down the middle of a tool, for fun and profit". It leads you towards a structure that has lots of benefits, such as meaningful `--dry-run` output and more ready librarification.
It's not material to the point, but I think there's a small typo on line 12 of the second "Finishing the example" snippet: it looks like
let instructions = gather args
should be let instructions = gather inputGlobsYou're quite right; thanks for letting me know. I'll get that fixed.
For example
rm PATTERN
Is really a shortcut for something like (pseudo code): find -name PATTERN | rm
If things were always factored for command/query seggregation you'd not need dry-run, you'd simply just run the query.Consider a conceptually very similar case: instead of finding and deleting files in a filesystem, think about finding and deleting lines matching a pattern within a file. Then the query is something like grep...but what do you put for the command on the other side of the pipe?
Of course, you can tell grep to output line numbers and use a command that operates on line numbers, or similar. The point is, in order to achieve this kind of segregation, you need some common way of naming operands on both sides of the divide. And naming things is hard, so segregating things this way is hard, and thus there's a lot of tools with mixed query/command responsibilities.
I still love PowerShell.
You would first build a pure function that read the files and returned matching lines. You could then display the count of the matched lines, and if the user wants they can see those lines and what line number it is. This looks a lot like a dry-run.
You can then pass these results into an executor. You then build a system that combines all deletes of a file and passes that to your DeleteLinesInFile fuction that does so in a single transaction. Cycle through each file and voila.
The example I was responding to was from the (Unix) shell, which as far as I am aware does not offer any such convention other than filenames. Real programming languages of course offer more options, but even there the problem is not solved: as you say, you must then build the system which understands how to take the results of the query and execute something on them. That means designing such a convention inside your program. The article is about using a type system to help you do that cleanly.
DRYRUN=1
VERBOSE=1
cmd() {
[ $DRYRUN -eq 1 -o $VERBOSE -eq 1 ] \
&& echo -e "\033[0;33;40m# $(printf "'%s' " "$@")\033[0;0m" >&2
if [ $DRYRUN -eq 0 ]; then
"$@"
fi
}
And now you can put this all over your scripts to very easily implement dry-run behavior without any quoting worries. Note that you can't invoke aliases this way. cmd echo "don't worry about quoting"
cmd do_something_dangerous "$scary_arg" "$scary_arg2" DRYRUN=echo
cmd() {
$DRYRUN "$@"
}
?run() { if dry; then echo "${(q-)@}" else "$@" fi }
It doesn't work all the time as the previous poster noted, but it is very low friction which is especially important when writing quick throw-away scripts.
For irreversible (such as deletions) actions I even go one step further with a `--no-dry=<some-non-trivial-value>` in order to force some thinking.
Edit: apparently this is way more common than I thought, by reading comments here!
I have found it a frustrating nut to crack because it seems like it needs to involve either writing a simulation layer that tracks the state of all the file calls virtually, or else copy all the files to a temp folder (so the dry run isn't dry, it's just on a separate version of the data). Both of these seem like bad solutions.
So if you need to install a package which sets up a systemd service, the subsequent dry-run of managing that systemd service would fail because the service doesn't exist yet. Or you need to make assumptions in the service dry run that it should have been set up before.
You can just assume that the service would have been setup and report that you were, say, enabling and starting the service. Or you could require some kind of hint to be setup from the package installation to the specific service configuration (for well known things this could be done by default, but for arbitrary user-supplied packages this cannot be done). Or you could list the entire package contents and try to determine if the manifest looks like it is configuring a service.
And it still may fail because you don't know the contents of that file without extracting it and the service may not parse at all.
And that's just a simple package-service interaction example. You could spend a week noodling on how to do that fairly precisely, then there's a hundred or a thousand other interactions to get correct.
You're being told not to do the thing, so there's fundamentally a black-box there, and you need to figure out how much you're going to cheat, how much you're going to try to crack open the black box into some kind of sandbox so you can figure out its internals, and how much you're going to offload onto the user. Not actually an easy problem at all.
In practice though what you'll find is that its easier to treat whole systems as black boxes, and test changes in throwaway virt systems, then test them in development environments, then roll them out to prod (or that whole immutable infrastructure thing and throw it all away and replace so you don't get bitten by dev-prod discrepancies in theory).
Maybe in your case the level you are trying to report dry-run information at is too granular.
Or maybe you have too tight coupling between your code that determines what needs to be changed and that which actually does the change and you might want to refactor the code that determines changes to the code that makes the changes.
This is problematic in multiple ways:
1) The script uses a single repository of project metadata to derive the rest of the build. This repo needs to be present and updated to figure out what to do.
2) There are sequential dependencies on a great deal of steps that can fail without much predictability (including network traffic, git commands, and the module build itself). But in a mode where you're compiling every day, even a transient module build failure does not necessarily mean you shouldn't press on to also try and build dependent modules.
I've ended up using an amalgamation of logical steps:
* If a pass/fail check can be performed non-destructively, do it (e.g. a dry run build will fail if a needed host system install is missing, or the script will decide on git-pull vs. git-clone correctly based on whether the source dir is present) * If pass/fail can't be done, just assume success during dry run (e.g. for git update or a build step) * For the metadata in particular, download it to a temp dir if we're in dry run mode and there's no older metadata version available. * Have the user pass in whether they want a build failure to stop the whole build or not.
For separate reasons, I've also needed to maintain logic to help modules with build systems that didn't support a separate build directory to act as if the build was being executed in the source directory (this was done by copying symlinks to the source directory contents into the proper build directory, file by file).
Having this separate, disposable directory made it safer to not have to be perfect on tracking files changes or alterations perfectly. And of course it helped greatly being able to offload the version control heavy lifting to CVS (back in the day, then Subversion and now Git).
if (clear_db_flag) {
if (verbosity_level > 0) print("DROP DATABASE customers;")
if (execute_flag) exec_sql("DROP DATABASE customers;")
}I.e. instead of a --dry-run flag to make the script harmless, I have an --execute flag that weaponizes it. This means that while developing the tool I'm more likely to run in dry-run mode (i.e. without the --execute flag) and so test the dry-run path more than I perhaps would have.
It also means I'm not very likely to delete the production database by accident, through running my script in what I intended to be dry-run mode but actually ending up deleting something I really didn't want to delete.
I think it's a nice methodology, to always have a print statement first, then perform an action that the print statement informed the user about. Easy to remember to add the print statements and more likely that you keep the dry-run code path up to date.
That way you don't need to worry about issues with bugs in the dry run logic and you have a record (a script) of what you actually ran.
Even better if you can extend this output script so that it's idempotent and fails on the first failed command.
I have never actually encountered such a race condition while using any tool I've written this way. But I do remember one place where I specifically didn't add a feature to the tool because it was so vulnerable to a TOC/TOU race condition. Instead, I turned that feature into a "did I just manage to do the thing successfully, and if not, why not" check at the end of the "do-it" stage.
So yes, that concern is extremely valid, and if you're doing an operation whose validity is liable to change after you've done the check, then you do just need to follow a different approach for that operation. (But then there's probably no possibility even in principle of a `--dry-run` for such an operation anyway!)
Obviously the effort of doing this becomes a trade-off but it's worth considering.
"Would have gotten all servers tagged foo (currently 1,341) and deleted them"
With server deletion, of course, the problem is less visible because it's naturally idempotent (unless you're referring to servers by some non-unique key). In that case, you can safely just pretend the leftover servers came into existence after the entire tool came into being.
Of course, if you're referring to servers by name or IP or something, then you hit exactly the problem. Concretely: I run a command that deletes all servers older than one day. The 'dry-run' section determines what servers need to be deleted and it turns out that server "foo" needs to be deleted. Elsewhere, someone spots that "foo" exists, deletes it, and creates a new server with the name "foo". Then the 'execute' flow deletes the new "foo", which is entirely not the server we wanted to delete.
Planning stage would generate `Delete(Server{CreatedBefore: 1621887357}))`.
It would output something like "Deleting servers created before $time$, currently $servers$."
The point is the planning stage would encode the operation being done, not a list of servers to delete.
That's a quibble though, the article is well-put and quite correct about the value of a dry-run mode.
Has anyone written about this idea? It seems to have some interesting interaction with laziness.
Nice aspects of this technique:
Inside the decorator you have access to both the function’s name and its arguments so you can print descriptive messages to the user e.g. “In dry run, would have run function ‘delete’ on ‘/foo/bar’”
You can use the name of the decorator as de facto documentation, e.g. “@destructive”
To implement this in your own scripts is not difficult at all. The template snippet below automatically wires up "-Confirm" and "-WhatIf" support:
[CmdletBinding(ConfirmImpact='High',SupportsShouldProcess=$true)]
PARAM(
[parameter(ValueFromPipeline=$true,Mandatory=$true)]
$TheThingToProcess
)
PROCESS {
if ( $PSCmdlet.ShouldProcess( $TheThingToProcess, 'Do scary thing' )) {
# Scary step
}
}
That even wires up the full "[Y] Yes [A] Yes to All [N] No [L] No to All [S] Suspend" logic, so you don't have to track boolean variables all over the place.It also passes through that preference to commands invoked witin the script, so things like "Format-Volume" will do nothing if you invoke the outer script with "-WhatIf".
Its been really helpful when hacking on things, adjusting logic, or validating that input data does what I think it will.
I felt it was a good design decision since the effects of bulk renaming can be substantial. Most other similar tools have it backwards with their dry-run mode relegated to a secondary action.
more examples I've collected:
- Angular CLI https://malcoded.com/posts/angular-fundamentals-cli/#--dry-r...
- AWS CLI https://docs.aws.amazon.com/cli/latest/userguide/cli-usage-h...
- Rspec https://relishapp.com/rspec/rspec-core/v/3-8/docs/command-li...
- Serverless framework https://forum.serverless.com/t/dry-run-with-serverless-frame...
at netlify we didnt do a dry run but Netlify Dev is close: https://news.ycombinator.com/item?id=19615546
I'm requesting one at Gatsby: https://github.com/gatsbyjs/gatsby/discussions/16384
pretend=${scriptname_pretend:-false}
run() { echo "$@"; $pretend || "$@"; }
Sure, this doesn't give copy/paste re-usable output, but it's really just a quick and dirty sanity check when things start getting complicated.Yesterday I wrote a cleanup tool. In fact it’s two tools: one that outputs a parseable list of things to clean up...
# snapshotX # keep
snapshotY # delete
...and another that consumes the list and does stuff. It’s the only way to be sure.The most common use is as a write/modify gate, but many tools have uses where it may need to avoid certain classes of reads as well (network, db, etc).
See this past submission: https://news.ycombinator.com/item?id=18499712