Escape the scripter mentality if reliability matters (2011)
rachelbythebay.com
rachelbythebay.com
Unless there are some serious domain experts on hand, I feel mapping the problem space is the most valuable approach - and one of the best ways to do this is to attempt a fast and basic solution. The alternative is a big design up front with minimal experience in the subjext domain.
It feels like the inverse of "I'm going to scale my foot up your ass"[1] - how useful are your reliability and metrics if your service is never used / becomes irrelevant / newspaper business dies / company folds / huge paradigm shift in tech.
I feel the dirty hack is a valuable step in the process, and its not possible to jump to the end without mapping the route first.
[1] http://widgetsandshit.com/teddziuba/2008/04/im-going-to-scal...
But some domains are very poorly understood. You won't get far by trying to iterate - because you don't even know what the problem is, never mind how to solve it, and your dirty hack is going to introduce naive assumptions that turn out to be spectacularly wrong.
The core is always the same, sure - you're basically wrapping read and write transactions to a storage subsystem in a pretty veneer. The complexity is never in the CRUD actions themselves. But transforming the things that have been R-ed so we can U them or identifying the full population of things to be D-ed is often non-trivial and deeply tied into understanding of the business domain and process at hand.
Often (analytics and reporting lenses on here), the complexity is in modeling the data appropriately and the logic of transformation. The work of presenting this stuff to a user is basically trivial arithmetic.
> Let's say I want you to get the fifth word of the fourth paragraph of the third column of the second page of the first edition of the local paper in a town to be determined. It's going to be a color, and we have a little deal with that paper to get our data plugged in every day. You can get the feed from their web site.
For the one-off case, I'd just go open up the paper and look for myself. Surely that'll be faster than trying to write some sort of a script to do it.
> Now I want you to be able to do this reliably every day for the next two years. I need this data on a regular delivery schedule and it can't rely on some human being there to constantly fine-tune things.
Ok. So now we're automating a 5 minute job that can be done by the office assistant, which will be performed 730 times. That puts the upper bound of time saved at around 120 hours of work. Due to their specialized training, software engineers probably cost the company more than 3x per hour than the assistant, so this automation task only makes sense if it can be completed in less than a week of work for the engineer, including maintenance over the next two years.
That sounds like a reasonable, but tight, time budget for the given task and reliability specifications. It's not an obvious win, which also means the ROI will be small. Also, there's opportunity cost to consider: that week of development time now can't be spent on other projects that may be more urgent.
Your analysis ignores the cost of failure to get the right word, or maybe just assumes the office assistant is infallible. That cost is almost certainly non-zero, and might be much higher than the whole development cost.
But in the general case you’re absolutely right: any analysis like this should include the costs and probability of faults occurring in each option.
Mandatory xkcd reference: https://xkcd.com/1205/
Then, you need to evolve to something else that better fits the job at hand. This is not a technical problem or even one of "mentality", I think.
It's more of an organizational challenge. It means acquiring the level of agency needed to adapt or evolve solutions to problems.
If you need to dig a trench, once, in your backyard, sure a pick axe and shovel and some hard-labor is just fine. If you need to dig a trench every day... you need a backhoe. Sadly so many organizations choose to do the equivalent of operating "chain-gangs" to dig trenches with pick-axes and shovels, at scale, every day. The people in charge of these chain-gangs just don't know any better and the people digging don't have the agency to demand the right tools and right approach.
This is a dangerous assumption. It seems at least equally likely that they are just responding to the incentives the system presents them. In other words, when you see a chain gang, you're seeing an organization that views hardware as expensive and people as cheap.
In cases like this, where both the solutions are software, the reason the company is relying on a progressively Matryoshka'ed system built on top of the original, "dumb" solution, is usually that they value developer time (to build a better solution) greatly over ops-staff time (to implement the infrastructure required to scale the dumb solution) + additional hardware costs (for all the overhead the dumb solution's method of scaling introduces.)
But even that doesn't explain why organizations refuse to switch from bailing-wire scripts to a pre-built, commonly-available infrastructure-component better solution. In such cases, both staying and switching are ops costs.
The costs of staying are known, because they’re the ones you’ve already been paying for a while. Switching to a new system is inherently risky as there’s always a significant chance you haven’t correctly identified all of the requirements from the existing system.
It depends on how critical the task is. Can you detect an error occurred that you now need to fix? Can you tolerate the lag between error and fix? etc
I saw solutions that started like this and very quickly became an unmanageable mess. I guess it’s some sort of pattern where it’s hard to see this fine line between “it works so don’t need to reengineer it” and realising a huge technical debt incurred. It requires experienced person to make a decision to drop the former and start with some more sane solution.
Where does she get this nice data, where integers are integers and fields have nice names?
Edit: The meat of the article:
> This gets into a whole thing I call "scripter mentality". It seems like some people would rather call (say) tcpdump and parse the results instead of writing their own little program which uses libpcap. Calling tcpdump means you have to do the pipe, fork, dup2, exec, parse thing. Using libpcap means you just have to deal with a stream of data arriving that you'd have to chew on anyway.
Edit 2: “Scripter” here refers to operating on text files. The Rule of Composition [1] encourages the use of text processing, for example:
> Text streams are to Unix tools as messages are to objects in an object-oriented setting. The simplicity of the text-stream interface enforces the encapsulation of the tools. More elaborate forms of inter-process communication, such as remote procedure calls, show a tendency to involve programs with each others' internals too much.
[1] The Art of Unix Programming: http://www.catb.org/~esr/writings/taoup/html/ch01s06.html#id...
The guy who handles this has set up a SQL database and pulls data in via FTP and email. Outlook sits open on a server so that Outlook message rules can be used to fire up a VBScript which in turn calls a stored procedure to ingest the data. There are a bunch of combinations of these to deal with different feeds.
He's a brilliant guy and he's used the tools he knows (not a dev) but it's brittle and only he knows how to support it. It's unlikely they'll move to our platform (invasion of his turf) but I do hope he finds another solution that's more reliable.
What’s your iPaaS product?
This is the problem with most places I have worked: "I want this and I need it today." Scripting is the obvious answer.
Then it will change to "I want it to run every day" and even if I say "it needs more work" the response will be "well, it's working now, so why do you need to fix it?" And so a shell script or several go into production and I can't stop it.
In any case, even if you use something more formal, you can still write fragile, error-prone code. Bash or Perl is not the problem - you can do great work with those tools, if you have time.
If you're going to stack the deck like that, then sure, scripting isn't the right approach here.
But in reality, for many things, automation that works 90% of the time and reliably attracts manual intervention the rest of the time is often fine. If you can build it quickly, it wins out over something more robust that takes significant investment to build.
Error handling, edge cases, unit tests, type checking, packaging can all be performed in Python.
Type checks are the biggest bug bear. Otherwise, it’s relatively easy to get things clean and tidy (packaged).
If you let your Apache Groovy scripts get too large before switching on static compilation, you often have to modify the types and logic in your programs before they'll even compile. This problem occurs because static compilation was only bolted onto Groovy for version 2.0 and there's an impedance mismatch between what's required for its dynamic and static modes.
This made me smile, because it made me remember a suite called “DSC” that I could imagine inspired this post. It did work, but seemed like a kludge. Last time I used it was when this post was written actually, and looks like it does use libpcap now.
And those people are right. Too many systems have too many unnecessary layers and "just in case" code, which instead of making them more robust than a small script, serves to slow them down, and increase the possible failure modes exponentially...
I've done lots of shell and Perl, and you can make them as reliable and bulletproof as any other option.
Do lots of scripts ignore return codes, lack retry logic, sanity checks, etc? Sure, but so do lots of applications written in compiled languages.
"Scripting" isn't really the issue.
If you really have finite, bounded requirements, then treat it like a full project with front-end design, granular types, kit & kaboodle.
So often these shiny objects are more moving targets.
In my mind, there are two ways to write software.
One way, represented by traditional product shipped to customers (think Microsoft Excel), is writing software in such a way you can show it works correctly under all possible circumstances. You need to write every piece of this software to be resilient to changes in environment and always do the right thing no matter what. Writing software like that takes care and is very expensive.
Another way, represented by typical business application written for use by the same company, is writing software to work in only very narrow possible set of circumstances. You demonstrate that it works and you ship it. Do you need it to work on all possible OS-es? Hell, no, you have an image and it works on it, it is enough. You insulate yourself from outside interference as much as possible, docker images, tightly controlled network environment, dependencies, only single supported integration architecture, etc.
The second way is the way to go when building "enterprise" software for internal consumption. You need to understand what you are doing. You need to understand why you are doing. If you are doing it correctly it will always be cheaper and better than investing in "perfect" solution.
http://www.catb.org/~esr/writings/unix-koans/ten-thousand.ht...
How do I run your node.js thing on my system without installing dozens or even hundreds of packages and modules?
It's not the same thing.
The actual way to do this would be scripting using beautiful soup and similar tools.