The 10:1 rule of writing and programming
ybrikman.com
ybrikman.com
The way VCS works, any edit would register as an insertion and deletion.
The way writing a book works, is you write some kind of draft and then mercilessly edit by rearranging things and changing words. But you don't change every single word 10 times.
Moving stuff around creates a lot of edits very quickly -- moving a paragraph creates many lines of insertions and deletions but I would never count that as rewriting all those lines.
The same goes for editing words -- changing a word will change entire line (insertion/deletion). If this also rearranges paragraph (due to line breaking) it may cause a lot of edits.
So I don't really buy into the whole premise of this article.
"On sticking with it until it’s right
“Challenge every single sentence for lucidity, accuracy, originality, and cadence. If it doesn’t meet the challenge, work on it until it does.”
While I was editing Jitterbug Perfume, Tom would read me a passage aloud to see how it sounded. Sometimes I’d comment, sometimes I wouldn’t. But each time I heard it again, it had changed. I saw how many times he would rewrite a passage and how much he relished doing it.
“Sometimes 40 times,” he told me.
He took the process of conception, research, trial and error very seriously, moving things around, changing voices and pitch. He wrote slowly and carefully, revised constantly, refining and evolving the novel over the course of about two years."
https://alanrinzler.com/2013/07/happy-birthday-tom-robbins-t...
YMMV
I cannot comprehend how he wrote 37,000 lines just to get 7,000.
Perhaps my problem was easier, but it's a domain I didn't know particularly well though and made a few erroneous assumptions about.
Most of the code I got right first time, save a few minor bugs here and there.
When I worked as a translator (I had an episode translating books from English to Polish) I would more or less write the final version of the text and there would be very few edits. The plan for the text was set by the original so my job was basically to figure out the way to express it in Polish.
On the other hand, when I started writing my own texts I would spend much more time editing. I would write passages that I wanted to include, but then I would move them around a lot so that they make more sense for the reader, for example.
The same with programming. When I know very well what I am building there would be very few edits. I would start bottom up writing modules and tools and abstractions that I know I will need later and build on it. A website would be a good example. It is not a rocket science. Once you design your UI, processes, it is fairly straightforward to translate it to code.
Again, if I work on something very tricky, there is A LOT of edits. My other project is an embedded controller for an espresso machine. I am building this controller from scratch, meaning I design the board, then I receive PCB from manufacturer, then place and solder all components. The code is for ARM Cortex-M4 microcontroller which I am getting to know. I would say most of the code was edited many, many times over because of many iterations I need to get stuff working and then to refactor it to integrate with the structure of the application.
The main exception is when you literally throw everything away and start from scratch. The odds of writing the exact same words twice are basically nil. But I'm not sure this should count as "editing" in the same sense as when I'm doing word-and-line copy editing.
I haven't done a rigorous comparison with code, but it would be interesting to take a tool like Moss and see how similar it thinks the code is after <N> revisions of the source.
1. I don't do a commit for every single line that I change. In fact, I may change a line 10 times, and commit only once.
2. This is actually even more pronounced for code. While doing a code-test cycle, I may change a few lines of code 50 times over, but only do one commit.
3. For my books, a lot of edit rounds and writing happened outside of Git (e.g., O'Reilly does copyediting in a PDF).
My guess is that these two factors roughly cancel out. It won't be exact, of course, and the actual ratio may be 8:1 or 12:1, but the order of magnitude is probably correct.
For anyone writing non-fiction, I would highly recommend the book "on writing well" by William Zinsser, who also stresses the importance of editing relentlessly your draft.
It's funny, editing has always been my favorite part of writing. It's when all of my favorite bits typically get written.
- Building the writing habit of words on a page and how you make that happen.
- Being comfortable removing all of the words that are unnecessary through the editing process.
I knew I crossed the bridge when I found myself editing my own work and saying: "this is an incredible sentence, but it's not the right one here, axe it."
TOTAL_CHANGES="$(git log --all --numstat --pretty="%H" | awk 'NF==3 {plus+=$1; minus+=$2} NF==1 {total++} END {printf("%d\n", plus+minus)}')"
TOTAL_LINES="$(cloc . | tail -n 2 | head -n 1 | awk '{print $NF}')"
echo "scale=2 ; $TOTAL_CHANGES / $TOTAL_LINES" | bc
echo -n ":1"
pwd
One snippet from https://gist.github.com/sephiroth74/d3ce654dcaefe1d2c4cd (modified)
How to run:
1. Clone a Git repo
2. cd into it
3. run bash ../script.sh
Output something like this:
polygot@server:~/tinkerpop$ bash ../script.sh
14.17:1/home/polygot/tinkerpop
Here's a sample of some ratios from random git repos:
158.25:jgthms/bulma
2.37:dthree/cash
32.49:Rochester-NRT/AlphaGo
7.04:apple/swift
5.00:robbyrussell/oh-my-zsh
108.86:github/gitignore
10.17:facebook/draft-js
12.77:matryer/bitbar
7.47:Microsoft/vscode
9.42:ptmt/react-native-desktop
6.06:p-e-w/maybe
1.00:parkjs814/AlgorithmVisualizer
9.54:hollance/swift-algorithm-club
17.63:alexjc/neural-doodle
4.55:fbsamples/f8app
8.80:nodejs/node
10.64:jiahaog/nativefier
325.12:Microsoft/CNTK
28.56:atom/atom
3.21:jcjohnson/neural-style
40.01:tensorflow/models
In other words, your work isn't done when the program works and makes the customer happy. It's done when other developers can read it. Otherwise you can't continue making the program work and make the customer happy.
I'm programming for 20 years and I still can't provide good ETA-s for non-trivial problems.
Well, that's why they are non-trivial problems. Don't be too hard on yourself ;)
You can only be 100% accurate when it's done. Here's a chart: https://cdn-images-1.medium.com/max/1600/1*hS2GObQarMYhamPeV.... Basically, the closer you are to done, the more accurate your estimation will be.
There is a really cool tool called `gource` http://gource.io/ that allows you to visualize the changes in version control systems like git and hg. I highly recommend you try it on some of your projects---it can help you see who did what. I bet you'll find some surprising things when you watch the history of the project... which parts were done first, which things are new, etc. Is there a genius colleague on your team that wrote the whole thing? Who are your colleagues with good work ethic that consistently push solid bugfix commits and refactoring all over the code base?
Here is a sample run of gource on the .tex source code for my linear algebra book https://www.youtube.com/watch?v=tlfrPRe-BeM here is a more interesting one showing 5 months of git's git history https://www.youtube.com/watch?v=GTMC3g2Xy8c
And I think it comes back to the principles of Slack. By shooting more than you need, you have room to cut. Hopefully, you'll have to wind up cutting good stuff because you have too much good stuff.
This method of practicing and improving used to be expensive with film, but with digital the process is now free.
The other thing you learn is your camera shortcuts so that when you do run into a photographic moment you can quickly take a bunch of shots with different settings.
I've gone down this road and I don't think that it's good advice any more. The secret to good photography is having a good editor, someone who can tell you which photos you have are good and bad, and why. Someone you can have a discussion with over a dozen 11x14" prints. Continuous improvement. Develop your eye and shoot neither too many nor too few photos. Every time you go through the review process you get a better sense of what you should be looking for in your photos.
It's so easy to fill up an SD card with nothing but garbage. I've done that. I've gone through bulk rolls (100' of 35mm film) and filled them with garbage I just don't want to see any more. By shooting more than you need, you can easily fall into the trap that because the photo will probably get cut during editing, there's no point to putting any special amount of effort into it. So there's no point in shooting more pictures if you're not putting enough effort into each one that it could be good. When I was at college I remember having discussions where we thought that there was not really any point to going through more than a single bulk roll per academic quarter (18 rolls in 10 weeks, or ~65 exposures per week). People who shot less than that much weren't showing as much improvement, and people who shot more than that were just ending up going through the process a bit too mechanically. That was just a rule of thumb for the classes at that particular place and time.
There are good, great, and amazing photographers who shoot such a small number of photos it would probably shock you, and others who shoot so much it makes you wonder how quickly they must go through equipment.
The downside is that I'm usually lazy/reluctant to delete the extras, especially when there's a tossup as to which is the "best", so I end up with tons of photos I don't want in my library.
As a side gig, I take photos of domestic animals (primarily dogs). Some of the animals can be posed like little dolls. Others can't sit still to save their lives. I also take action photos of them. I have to spend time manufacturing perfect, because I often can't make very many on-the-fly changes. But to capture that perfect, I have to take a lot of photos.
As a hobby, I practice general photography, and my usual method is to lock some variables, and take several photographs while floating other variables.
Also as a hobby I practice astrophotography, and here you absolutely have to manufacture perfect, because you only take a lot of pictures in order to stack them. A single photo can take hours to produce, so you do what you can to get it right in the first place.
terraform-aws-couchbase (2018) - 5:1
Terratest (2016) - 8:1
Terraform (2014) - 9:1
Express.js (2010) - 14:1
jQuery (2006) - 15:1
MySQL (1995) - 16:1
This is a small sample size but it seems easy enough to run in some popular open source perfect and see if there's a statically significant trend. It would also be cool to see the lowest and highest ratios on popular projects.
It would also be interesting to think through how many changes there are _between releases_ of a project. MySQL is 23 years old [1], but what's the effort/change between major releases at this point? That's where the rule sort of falls down for me: a book has a few editions (if it's lucky); software, on the other hand, has lots of releases if it's successful.
[1] Crap, now I just feel old, too...
So if it was normalized, we could see whether newer projects are rushed to production faster or it's just a matter of time passing.
N = number of API endpoints, then N/cloc would be something like efficiency (get more stuff done with less code)
M = number of lines of code for all code paths accesses during average day on PROD, then (cloc-M)/cloc would represent the "dead code" ratio---how much of your code base is not used
X(c1,c2) = an arbitrary function computed on the diff between commits c1 and c2
And all of the above can be run using some sort of rolling window from git init to today.
Looking at own git logs (for code, not prose), I notice git is very happy to add and delete lines even if just some characters on those lines have changed. Is it not the case that the author of the article has simply edited 90% of the lines in his book, rather than rewritten them completely? That seems much more likely to me.
Git doesn't really have good statistics (I think) for e.g. token changes (it would need a tokenizer for that). There is however --word-diff for showing a diff in words instead of whole lines; it's probably possible to count those as well, the output syntax looks parseable enough.
On the other hand, editing takes a lot of time per word changed, because you have to consider the whole sentence or even paragraph to make sure the grammar is still right, you're not repeating things, and you haven't removed any necessary context. So maybe a bit of overestimation is good.
Such numbers of whitespace changes do not happen often, but have a big impact on LOC when measuring a long period of time. I realized that when counting my changed LOC for a past internship.
Stayed at that job for 4 months before realizing the other engineer was just going to keep writing the same crap. No tests, no comments.
The real kicker was that the “database” was several different instances of mongo. So doing joins was a pain, and having transactions was almost impossible.
I would however agree that:
* those who produce & then cull generate a lot more than we see in the end product (due to those who do little editing, probably more than 10:1)
* writers/programmers get better at their craft, reducing the chaff to wheat ratio as they practice (like the original author saw in his work)
* well understood domains or themes are also more efficient (Michael Crichton didn't throw away 90% of his output; your average CRUD app doesn't either)
Finally, the written word tends to remain unchanged once committed (revisions tend to be minor in the overall scope), while code that is used changes dramatically. Long-running projects that are successful will see their application increase and evolve accordingly, so it's likely less valuable to track change over time if you're looking for total vs. finished effort metrics.
For instance, I may have to edit a project or UI component and generate “churn” just because it decides to reformat some XML for a minor edit. Or, in code, some styles are clearly more “vertical” than others, comment paragraphs may or may not have changed, etc.
Then there are tiny edits that have a huge amount of “churn” in the actual project, such as “#if 0”.
While lines of code is not an entirely useless measurement, it would definitely be good to have at least a couple other things measured. For instance: compiled binary sizes before/after edits, number of bugs logged per month, or something. Even then of course, these additional measures can obviously be affected by external factors (how good is your compiler, what kinds of bugs actually appear in reports, etc.).
My takeaway is that measuring productivity accurately is hard, and if you want good accuracy then you have to put a lot of effort into the measurement process. There are no easy 10:1 rules for these things.
"Does the amount of churn depend on the type of software? For example, Bill Scott found that at Netflix, only about 10% of the UI code lasted more than a year, and the other 90% had to be thrown away. What are the rates of churn in backend code, databases, CLI tools, and so on?"
I used to do a lot UI. This jives with my experience.
At the time, I resolved to divine concise ways to prototype and implement UIs.
20 years ago. I didn't get very far. But I am now making another run up that hill, so we'll see.
Instead of getting lines removed + lines added, there could be some sort of content aware diff algorithm that works with 'data chunks' at another granularity (lines = good for code, words ala diff --color-words = good for text)
diff = fancy_diff(c1, c2)
# diff ~ {'added':<int>,
# 'removed':<int>,
# 'changed':<int>,
# 'moved':<int> }
where `moved` represents # of `-` lines in diff that have an identical corresponding `+` line,
and `changed` represents lines where 60%+ of - line matches the text in some + line.Using such a fancy_diff, moving a 200-line chunk of code from one source file to another source file won't show up as -200s and +200s, but as {'added':0,'removed':0,'moved':200}
Looking at such "fancy diff" numbers will show more what's going on...
Not 100% sure on that though, nor am I sure if the same applies at the line level.
To be honest, I'm not sure if computing a fancy diff of the form {'added':<int>, 'removed':<int>, 'changed':<int>, 'moved':<int>} is a well defined problem because there are multiple ways one could interpret a given set of changes, e.g., is it a move with a small edit vs. and add and remove. We might need to define some "economy of diff" objective to optimize... something like Kolmogorov complexity for diffs?
I'd believe that if the guy was talking about the book text only. Let's say you use bootstrap and then decide to use xxxStrap. You just added and deleted 100k lines of code. Congrats, you have a "churn" and stuff like that. But the reality is that you neither "wrote" nor "rewrote" these lines of code.
I might be wrong but I'd like arguments to the against. This happens a lot in projects I work within.
I've seen that in some projects too, usually where developers don't trust their package managers and/or just want to inflate the size of the project to manipulate the client.
Rule of thumb: don't commit external dependencies (like node_nodules) when you can lock dependencies in place. If you really have to do it then have a downstream repo that will do that automagically.
This makes your strap situation far less problematic. You'll still touch everything that uses said strap but then it's actual programmer effort and not manual package management.
Checking node_modules into VCS is a code smell and footgun. Don’t do it. Check the lock file into source control and make sure you have a reliable cache in between build and registry (eg the yarn public registry, which caches everything, or a private registry and/or proxy).
[1]: https://arendt.io/#!posts/Pareto_for_Estimate_Correction
There were some issues, but overall it was an awesome experience. Unlike my friends, I actually enjoyed writing my thesis.
Also my friend told me how to estimate the timeline of a project - double (yours or supplier) estimate and increase the unit o measurement (so 1 week becomes 2 months, 2 months -> 4 quartes etc.). Works every time...
If that was the case, Carmack would still be working on Wolfenstein!
Git's cool and all, but I wasn't a huge fan at first. One version control system was like another. But the impact it's had? Now that version control is free and ubiquitous, using git for stuff is everywhere. That's turned out to be extremely cool.
The details may vary for each person's style.