How To Make a Quick Buck
jacquesmattheij.com
jacquesmattheij.com
One night I came in from "lunch" (2nd shift...) and saw three guys standing around a laptop trying to make something work for a customer. They had an Apache install which was supposed to log something to a given file, but they'd poke the machine with their browser and would get nothing in the log file.
I walked by on my way to my desk and one of them grabbed me. I couldn't type anything (because one of them was already in the way), but I could see what was already on the screen. It was a chunk of the config file and it looked a little like this:
<VirtualHost 72.3.x.x:80>
# (directive to log to that file was here)
...
</VirtualHost>
I asked them to run 'ifconfig'. They did. The machine had 192.168.x.x IPs because it was sitting behind a firewall doing NAT -- a fairly common config for our customers. Apache is pretty strict about matching things when you use an explicit IP+port, so it wasn't using that virtual host to service the hit, and thus the log directive wasn't used, either.I just said, okay, change the IP on that line from 72.3.whatever to 192.168.whatever and reload the config and try it again. They did, and it worked, so I continued back to my desk.
From what I heard, they had spent over an hour trying to figure it out. I didn't type a single character and never even sat down. I chalked it up to luck and went on with life.
The time I spent working with those folks was one of the most educational periods of my life.
Sure, he made 2,500 guilders in 5 minutes, but the author says of his early QnX days: "After a couple of visits there to talk over strategies on how to solve certain problems using QnX [...]". So to me this anecdote is about how guy spent an amount (several days? months?) learning QnX, getting involved in the (1-man) community around him, and became somewhat of an expert in QnX such that when the time arose, he was the guy they called.
The way the story is told is akin to patio11 saying "an easy way to make a quick buck is consult with a BigCo!" True... but that ignores the time, effort, and experience Patrick has put into building his skills.
TL;DR: "Invest time in learning something valuable" is a better takeaway from this article, IMO.
Seems like you're pointing out the obvious, at least to me.
I think the better lesson which is directly applicable to all coders is "question your assumptions". To me this was a huge benefit of learning ruby as it happens to be one of the easiest languages to pull your assumptions right out from under you.
When you’re looking for a bug it’s terribly hard to not get
bogged down in details at the expense of seeing the bigger
picture. An outsider doesn’t have those details so that
helps tremendously. The lesson I took away from that day is
that when I’m stuck I ask an outsider.
This pretty much sums it up and isn't limited to software engineering at all. Having chased some hardware bugs on my own I came to the same conclusion: one conversation with a not involved colleague ending in the question Well, have you checked (the obvious) X? was more productive then some hours or days of bug hunting.For me, the epiphany usually hits when I finally break down and decide to ask for help. I realize my cognitive blind spot while composing a message for a forum post or email.
It's always good to step back and acquire a fresh perspective.
Ah! The ancient feud between the School of the Duck and the Monastery of the Bear! I think modern times have calmed down the waters and all's right in the world of Engineer Fu.
I'd expect such an argument has been made by people - perhaps someone here might know of some examples.
- Woman: "It's You! Picasso. Could you please paint me?"
- Picasso: "Here you go Madam!"
- Woman: "Wow! great picture in only few seconds. How much do I owe you?"
- Picasso: "$5,000"
- Woman: "That's outrageous. It only took you few seconds"
- Picasso: "Madame, it took me my entire life!"
Unverified story about value-based pricingA very important thing to get comfortable with.
The root cause was a piece of paper on his desk with a diagram of the systems and their IP addresses.
Direct shell access to the machine wasn't possible. To get a shell on the machine that ran the processes he had to go via a separate gateway host, and the connection from the gateway host to the destination machine was done by hostname (as the gateway machine had access to the internal DNS).
To get a copy of any config files he was able to FTP to the server directly from his laptop, but he wasn't able to use the same DNS servers as the gateway host so he relied upon the piece of paper on his desk which had the names and IP address of the machines. Unfortunately the IP address for the system containing this config file was the one for the corresponding dev system, not the production machine.
Every time he ran the process he was using the config file that was missing a ; character and so it spat out an error.
Every time he got a copy of the config file (via FTP) he was actually getting it from the dev system which wasn't missing the ; character.
Much back and forth from support on "are you sure that's the right config file?" and he always responded with "yes".
When I got there I looked at the config file via the shell (on the right machine obviously) and spotted the problem straight away.
I didn't make any money out of it, but I did get a day and a half to look around a new city on expenses.
This strikes me as curious. Were you billing hourly?
[EDIT] I didn't quite walk out straight after fixing the problem. I stuck around to help out with any other queries they had with our software but they ran out of questions by lunch time.
[EDIT2] Customer wasn't being billed at all for my visit. They pay annual maintenance for the software though. If it's a serious (or weird) enough problem (or they shout loud enough) then someone is sent on site.
Years ago, I'd run into a somewhat similar issue, as Tomcat serializes its internal session representation on shutdown, so that when it comes back up, state can be restored. If one has truly gigantic session objects (which is a no-no anyway), it takes forever to serialize them, and as muster them back into memory.
They tweaked the config, the problem went away, their "chief architect" was fired, and I got the gig. Turns out this problem had them weeks away from going out of business, they'd been working on solving it for months.
The gig turned out to be a major disaster anyway, their problems were far deeper than technology.
Weinberg's Second Law of Consulting: "No matter how it looks at first, it's always a people problem."
In the early years of the 20th century, Charles P. Steinmetz, who stands among the electric industry's greats, was brought to GE's facilities in Schenectady, New York. GE had encountered a performance problem with one of their huge electrical generators and had been absolutely unable to correct it. Steinmetz, a genius in his understanding of electromagnetic phenomena, was brought in as a consultant -- not a very common occurrence in those days, as it would be now.
Steinmetz also found the problem difficult to diagnose, but for some days he closeted himself with the generator, its engineering drawings, paper and pencil. At the end of this period, he emerged, confident that he knew how to correct the problem.
After he departed, GE's engineers found a large "X" marked with chalk on the side of the generator casing. There also was a note instructing them to cut the casing open at that location and remove so many turns of wire from the stator. The generator would then function properly. And indeed it did.
Steinmetz was asked what his fee would be. Having no idea in the world what was appropriate, he replied with the absolutely unheard of answer that his fee was $1000.
Stunned, the GE bureaucracy then required him to submit a formally itemized invoice.
They soon received it. It included two items: 1. Marking chalk "X" on side of generator: $1. 2. Knowing where to mark chalk "X": $999.
The one I heard involved a Russian sysadmin, a giant server, the same 'X' mark, and a huge hammer.
Imagine if Jacques had charged for 5 minutes of his time at his hourly rate at the time.
The funny bit is that he probably said it in jest but out loud in front of four very wealthy customers so rather than cheapening out he paid up and ended up cementing our relationship in trust. Jan is an absolutely awesome guy, I learned more about telecommunications from him than from the CCITT books. His knowledge is literally encyclopedic, even today he's current on the state of the art.
Jan single-handedly designed and implemented a clustered system that would scale horizontally to insane volumes of messages (faxes and telexes) and he did that with early 80's technology.
Even today, with our current tech you'd be hard pressed to create a system that is as elegant as what he put together, he's both an awesome teacher and a good friend.
What was the computer?
The machines were mostly powered by 486/33, on a micronics motherboard, adaptec controllers (1542) to hook up drives.
Using the 286 version of QnX, all the joys of 'mixed model' programming. QnX took forever to release their 32 bit version (this was the main reason I wrote a clone).
Funny that I still remember those details after all these years.
Mine was a Adaptec 2940 (PCI bus, SCSI) that came out in 1994, it was so well made that you can still buy it today, almost after 20 years.
Short scope, well-defined, less risk, more fun, easier sell for both parties.
So I'm with you. I would not want to jump into a 3-month fix priced contract blind.
It was the main way I studied in university. People would come with some "difficult" problem thinking that I'd know the answer. They'd proceed to explain the problem, teaching me a ton of stuff while doing so, and there would inevitably be some small obvious thing they'd missed, I'd point it out and they would be super happy. I'm sure I learned way more from those interactions though.
It's the same in work as well, if you're "a troubleshooter" you end up learning an amazing amount of things from people coming over to explain something so that you can help them find the (usually) obvious thing they missed, or if it's not obvious you get to do some interesting research with them which teaches you some new stuff.
Bas troubleshooters have a hard time assessing just where to start, do we look at the immediate problem as reported, or do we track back through dependencies to a non-apparent possible root? Also, in a stressful situation like this it seems that it takes a certain amount of resolve to change 1 thing at a time and have a way to (quickly!) test your results so that you can actually solve the problem, instead of just pushing out the next occurrence a few weeks.
Where this gets really annoying is when you present them with direct evidence that the problem is elsewhere, and you get back an emotionally driven argument for ignoring it...
More actionable advice would be better :)
You can spend a week's worth of evenings plus the weekend and net a good few hundred dollars.
The advice I just gave, while accurate, isn't very valuable because it doesn't scale. I've done it a few times, bought myself a nice holiday last year with the proceeds. But it's not something that grows, you have to keep grinding away to create a recurring income.
I preferred it when my websites regularly produced £60/month in adsense revenues without me having to lift a finger. (That not-lifting-finger part only lasted about a year though. Now I barely make £6/month because I've neglected my sites).
Unlike the original story, "a buck" is almost literally accurate in this case.
So far as I know, nobody was in the lab over the weekend and certainly nobody would have touched my workstation. When I came back Monday the code worked. It worked the next day and the day after. The only thing that changed over the weekend: daylight savings ended. I have no idea why that would have any effect on the components involved and it is a mystery to me to this day. (Yes I restarted everything numerous times prior to that which would have explained the sudden fix much more simply). The worst thing is that I don't know whether things stopped working 6 months after when daylight savings started back up.
1. Create a 1 GB file of "nothing" (e.g. dd if=/dev/zero of=/path/meh.bin bs=1000 count=1000000) on each filesystem (exception: /boot, which is typically 256 or 512 MB).
2. Configure each filesystem to reserve 5% of space for the root user. This is typically the default, but I do it explicitly just to make sure.
When the first "low disk space" warning comes in, that 5% can be adjusted down to 1% to free up some space (on today's large filesystems, this 4% usually amounts to quite a bit of space). This buys you enough time (hopefully) to expand logical volumes, etc.
If you manage to about run out of space after freeing up that 4%, you have one more chance: delete that 1 GB file.
Another favorite of mine is: looking at the right file, but on the wrong host (for example when connected via ssh over a fast network, you simply forget that you're in a remote shell).
I've heard of people doing things like having a terminal background colour escape sequence in their login scripts as well, so it's immediately obvious if you're on a production host.
[1] http://manpages.ubuntu.com/manpages/natty/man8/molly-guard.8...
Now, shutdown and rm are the worst to add to the mix.
Hint: it was only once.
At my last job, we had a MySQL instance that kept falling over. It turns out that someone was dividing a table up with tons of partitions, and MySQL creates a new data file for each partition. The MySQL instance created several hundred file descriptors while trying to open each one of these files, and "mysteriously" crashed.
On a hunch, after seeing the pattern of errors in the debug logs, I dug up the old link list handler routines in our codebase, and lo and behold, there was an OBOE [1], a very latent one that too, dating back to early 90s. The fix was just an addition of a '=' which got things back on track. That incident made me the official schroedinbug [2] mangler of the unit, causing me to spend countless night-outs thereafter.
[1] http://en.wikipedia.org/wiki/Off-by-one_error [2] http://www.catb.org/jargon/html/S/schroedinbug.html
Clearly he just wanted to make sure that it actually was connected properly... but I like to think he was making sure the Internet was flowing into my computer. With those unidirectional Ethernet cables you have to check.
"Please reboot your computer" serves a similar purpose. Rebooting hasn't actually fixed any network connectivity troubles since approximately Windows 95, but the action serves as proof that 1) the customer actually has a computer, 2) they are at it, and 3) it's turned on. They really will have those basic facts wrong and lie about them if directly asked. ("Their computer" could be at their nephew's house, or be a TV.)
He could type for a while after I was back to my desk and before the keyboard went dead and he started to curse about a quarter of an hour of work lost. Everybody was laughing out loud, and louder because he was seeing that I had something to do with it but he can't figure out what.
The reason I think it may be apocryphal is that is bears a striking resemblance of the "knowing where to drill the hole" story told elsewhere in this thread.
More versions of the story and many others: http://www.snopes.com/business/genius/where.asp
Interesting, there's a very similar lesson in this post from James Bach's blog yesterday about programmer pairing with a tester: http://www.satisfice.com/blog/archives/852
Which, with inflation, is roughly US$2,400. http://www.usinflationcalculator.com/
Especially you're a ruby dev doing a bit of PHP haha.