Saving a Project and a Company
jacquesmattheij.com
jacquesmattheij.com
This is terrifying to me, especially as the CEO (I use the term because it's technically accurate in this case, not to be a douchebag) of a well-funded startup who is semi-technical. I can build basic apps and scripts and started out building things for the company, but now that we have real talent I would slow things down if I were in the code every day.
As such, I end up in a semi-product manager role, and consider myself responsible to make sure we're quickly shipping product in line with our goals, but at the end of the day I have to trust the people I work with, all of whom are experts in their respective areas. Luckily I can.
I have no idea how a non-technical product manager who has hired a contractor from an Eastern Bloc country would know when they're being reasonable and when they're just trying to screw him. That's just a recipe for disaster. It's a blind man in a new city taking an unofficial taxi cab. Makes me cringe.
I know of similar projects where the contractors who fucked up where from all around the world.
Incompetence knows no national borders.
But it is relevant because there is no real legal recourse, so you're pretty much crossing your fingers and hoping that they won't screw you.
I too will say that all of the Eastern Block developers I have worked with, both on site and remotely have been excellent.
"I don’t blame them for this, they were placed in a nearly impossible situation by their customer, the (dutch, not associated with the company that built the whole thing) project manager was dictating what hardware they were going to have to run their project on (I have a suspicion why but that’s another story)"
Sounds like the project manager was also outsourced to a 3rd entity and may have had some vested interest in the hardware part of the project. This just goes to show that the financial interests of those steering the "ship" better be aligned with those of the "ship" itself, or hilarity/misery ensues.
Certainly, many EB programmers are very highly-skilled.
I am sure that if we make an "Poll HN: At what distance away from you was the team/person that turned your project into nightmare?" we will find pretty even distribution in the brackets.
It is a truth of human nature, we are min-maxers to the core.
It's how you end up with a guy from marketing as the CTO of the Sony branch that got hacked. This is every single company I've been in that's not a software provider. The thing is, they are ALL software providers now, they just don't know it yet.
If we design our organizations such that one person can make or break each project and we know there are way too few people to fill those roles, then frequent failures like this are the expected outcome.
In particular, I think the role of "project manager" is mostly wrongheaded. The theory is that you can take a person who knows nothing about the details, give them total power and perverse incentives, and then expect things to turn out well. I think it's only popular because it's an artifact of our current managerialist business culture.
I much prefer cross-functional teams where the team as a whole is accountable for results. I also think we technical people need to stop thinking of ourselves as minions and instead act as professionals. If a "project manager" told a professional (like a doctor or a structural engineer) to do something unsafe or wrongheaded, they'd say no. But software developers routinely go on building something terrible after token protest. I'd love to see that change.
How in the world do you get more financial experience, short of bashing your head against it like I'm doing? Are there books you can buy that don't suck? We're going to double our development staff in Q1 (10->20) which is sort of terrifying....wish I had more knowledge!
As CEO your job (IMHO!) is not to "manage" the tasks so the project hits it's deadlines. It's to ensure that the risks in the project are recognised, seen as likely to actually occur and some action is taken Beyond crossing our fingers, to ensure that the project survives the risk becoming reality.
That might mean tests, or in house developers, or encouraging people to give bad news, it might mean hiring jacques before launch, it might mean building two teams to compete for launch whatever.
Just focus on the risks. Not the nice optimum path of perfectly executed tasks.
From having worked on timelines like this, it was probably less about the time to enable it vs. "If we turn on op code cache and that breaks part of our code, we don't have time to fix it."
Solution w/ nebulous oversight and looming deadlines? Punt it to the ops team.
I have a question about you find such work - or possibly, how such work finds you :)
For pretty much well my entire career, since 1998 or so, I've been the guy you go to when when you something's gone wrong, and no-one knows why. Unknown codebase using an unknown language on an unknown platform? I don't know how I do it, but I have a talent that lets me figure such things out when all others have failed, fast.
It feels like this should be a valuable skill; is it simply through meeting enough people that you're the go-to guy when situations like this flare up?
I've been tempted several times to form the League of Extraordinary Gentlepersons or something similar. You broke it, we fix it.
Making a corporate structure for something like this hard, I came up with a model a couple of years ago called 'The Modular Company' but the bookkeeping is a nightmare if you want it to be fair.
Still, I keep coming back to it and one of these days, who knows...
Anyway - I would be interested in any thoughts or traps to avoid - if indeed we are talking about the same structures.
Do drop me a line (contact details in profile)
Then in a way you do get paid (eventually, I hope...).
> Yeah, maybe just consulting?
That's one way to keep it simple.
But there has to be a better way, especially because the nicest flowers grow at the edge of the abyss.
I leave the process management stuff up to other folks, they're better salespeople than I am. They get paid to beat other people into submission; I get paid to beat servers into submission. :-) (Though not in Jacques' or Rachel's league.)
Perhaps some sort of league would solve some of the problems in doing that sort of work. E.g., the burstiness, the sudden need for specialized skills, the pipeline issues.
I used to be pulled into projects where the managerial side was messed up a lot. Usually the rep from one project is what would get me pulled in to the next.
An old non-technical version of this story is at http://www.amazon.com/Calumet-K-Samuel-Merwin/dp/1561141453
I think you need to work towards a system and network which will help you get that kind of work. I know of a person who has worked and works only on what he considers premium projects.
One advice I got from him was to that I should think about work in terms of projects and not companies. All companies have great projects and routine 'keep the wheels running kind of work'. Your chances of ending up working for the routine boring projects even in big successful companies are very high. Plus most companies have closed allocation policies, and tend to execute critical projects from one specific geographical location.
So to look at one's career in terms which project you wish to work on, and not which company you wish to work for helps.
Another factor is to seek out and work with smart people. Once you've been a part of a good team and proven your worth and are actively seeking out good work, you are always going to find some one in your network who will get you a good project and then you use new opportunity and the new connections to get more ones.
Nice gigs, usually no more than a few days, and working with tech I had no real experience in and otherwise wouldn't encounter.
Usually very similar problems: system slows down to a crawl or completely stops working because it hits some bottleneck (often after years of working flawlessly), original supplier no longer exists. Fun things to figure out if you have no immediate stake in it.
'Professional Troubleshooter' or something similar...
It turns out that postgres has an ‘auto vacuum’ setting that when it is enabled will cause the database to go on some introspective tour every hour which was the cause of the enormous periodical loads. Disabling auto vacuum and running it once nightly when the system is very quiet anyway solved that problem.
Often vacuum problems can also be fixed by running auto vacuum more often too - this means it has less to do per run, so should be able to keep up a little more easily. Loads of stuff on vacuum on the postgres wiki: https://wiki.postgresql.org/wiki/VacuumHeadaches#Perverse_Fe...
For what it's worth, autovacuum can be enabled/disabled on a per table basis, too. Some tables need frequent vacuuming, others less frequent, and others none at all. If you manually VACUUM a table, don't forget to also ANALYZE it!
Essentially, if (Load average / number of CPUs) > 1 (for CPU bound work) then your system is overloaded.
That seems like an insanely risky proposition on top of the risk you already assume through consulting, failing company, hiring colleagues as temp workers etc.
That said, there were quite a few details that made this project harder than it should have been.
Is it something they proposed or you? Seems like you could lose a lot and win little - and for them, the biggest risk was not your fee, but whether or not the system got up and running, so why bother?
Is it something you do a lot?
In all the years that I've been doing stuff like this professionally I've had one customer that didn't want to pay the full amount (they asked for a discount after the work was already done) and I told them to tear up the invoice but never call again. Everybody else was more than happy to pay. Maybe I've been lucky in that respect but I think that it's more of a way business is conducted here than anything else. You stand by your agreements, it's a small scene and word really does get around.
I thought I was being all nice and responsible with money: keep it on a single server, watch out for memory and CPU, minimize harm to the environment, outsource to maximize use of our limited resources. Now I'm looking elsewhere for employment and I have no "relevant" experience.
In the meanwhile, the clowns who made this mess get to claim J2EE, cloud, HA, VMWare, Redis, Angular.js, Symfony2, and a living client for their resumes, and their product didn't even work correctly.
| A single clueless person in a position of trust with non technical management, an outsourced project and a huge budget, what could possibly go wrong...
I know you might still have some degree of an NDA pinch preventing you from giving too many details, but if possible, can you give some more info on how you went about setting up the tracking instrumentation?
As always, a fun read. Thanks!
If you're running a store at any one point in time the store contains the number of people that have ever entered - the number of people that have left. So by just adding two counters (person entering, person leaving) you can validate the current state of the store by subtracting the second from the first and doing a quick count of the aisles. If you have more (or fewer) people in the store than you think you should have you have either another door somewhere that you're not aware of, people are being born or dying on the premises (that might work for a hospital ;) or they're climbing out through the roof.
If the counters match there is no guarantee that that is not the case but it certainly helps to gain confidence that you know where your entrances and exits are and that people aren't keeling over while shopping in your store.
Adding a large number of checks like that will eventually give you a very quick way to test your assumptions about how things should work and to determine the impact of a change on the system. We logged all those counters on a minute-to-minute basis (1440 records per day is peanuts), and have established a number of baselines indicating what 'normal' behavior is, what 'perfect' behavior should be and this in turn (over time) gives you a goal to shoot for.
If after a change you're below normal you've probably messed something up and should roll back, if after a change you're doing better than before than good, don't change, establish a new 'normal' in a couple of days time and strive for 'perfect'.
This trick has made it fairly easy to steer the project in the right direction and saved us from making stupid mistakes a number of times (most notably: at some point we realized the sessions weren't cleaned up at all, but cleaning them up too fanatically caused some of the relationships between the counters to indicate that we had a problem, it didn't take too long before we realized that the session cleanup routine was the culprit, without having that system in place this would have taken much longer and would have done a lot more damage).
I have done this so many times now that I have my own little mental model for what to look at when I get airdropped in: - Project managment. Do they have one dedicated project manager who is reporting status correctly and frequently to the stakeholders. Are plans available and follow-up, etc. - Product management. Are requirements from the business gaterhered and negotiated down to clear and concise things that can be built - Technical leadership. Do they use suitable technology, proper infrastructure setup (in your case not), is technical design simple to understand and not overly complex. - Change management. Is the team communicating the comes changes effectively to the end users. Is training done and being planned correctly. - Work process. Is there a good process that with good flow from requirements to created and tested feature.
My theory goes that if one of them fails, the project usually survives anyway, the others compensate. If two or more fail, the project fails.
Finally a question. You write that it is usually not a good engagement for you financially. Would be curious to know your business model here. I end up doing these projects on a hourly rate for the most part.
Not 'not a good engagement financially', rather the opposite. Just more risky. Typically I'll do these for daily rates depending on the perceived risk but I'm pretty flexible.
If it is just static content then you should be able to saturate your uplink from one single machine.
The 10K number applied to the whole setup, and that's now comfortably served from one box (as it should be). It could probably handle 10 times that number now without too much in terms of additional tuning (if any), above that it might require more work.
I have seen caches having more inserts v/s reads.
I have seen them replacing MySQL by NoSQL as it does not scale for them.
This is common for early stage funded startups founded by non tech founders.
"Let's make our change password system handle 100k requests per minute but the front page starts to get wonky at 1000 req/min"
Only time this table is touched to check user entered correct password or not.
Are they just logging in twice a day to punch in/out, like a online timeclock? Certainly too many resources for that amount of traffic.
But are each of the 10,000 users logged in all day, each working with large files or data sets, and doing intensive tasks?
How much space does the client need for this trading app? DBA figures 1k transactions per day at 1kb/record for 1 MB storage, round that up to 1 GB per year, so tell them 10 GB should do for 10 years.
CTO hears 10 GB from the DBA, adds his own safety margin factor of 10x, tells the client 100 GB.
Client hears 100 GB, adds his own safety margin factor of 10x, tells their operations 1 TB.
So now we have a client building a giant SCSI terabyte array (this was when 72 GB SCSI disks were the high end server standard) to hold a database that's a year away from reaching even one gigabyte.
Why all the virtualization? Lack of experience? I haven't done a large amount of work with virtualization, but stopping and thinking about it would seem to have indicated a problem with the design. Did no one look at this and say, "That's a bad idea..."?
So, yes: someone did say 'that's a bad idea' and got sidelined for his effort.
"The job ended up being team work, it was way too much for a single person and I’m very fortunate to have found at least one kindred spirit at the company as well as a network of friends who jumped to my aid at first call. Quite an amazing experience to see a team of such quality materialize out of thin air and go to work as if they had been working together for years."
I have done similar sorts of things for my current employer and at previous jobs. One thing we discuss here, since the opportunity keeps popping up, is forming a specific team of people to parachute into a flailing project and get it back on track. That frequently seems to involve taking it away from the then-current developers, paring it down to "the good parts", and then rewriting the rest, but it does yield results.
Then management wonder why it failed.
While freelancing, especially for people in another country, it isn't uncommon for customers to just disappear without any trace (or pay) half-way through a project. As such, when confronted with some problem mid-project, finishing the job in a half-assed manner can be a way of ensuring you get paid. Of course, a better method is making sure you have a solid plan before agreeing to the job, and trying to avoid unreliable customers - but accomplishing this can be difficult in any setting, even non-freelancing.
On the customer side, making sure you have a number of reasonably sized milestones and pay for them immediately on delivery can help keep freelancers confident, and thus encourage better quality work.
What about the jobs is financially risky? Do you have a downside beyond "might not get paid if the company fails"?
Another risk is that when I call my friends in to assist I assume their risk of not getting paid, in other words, if the company would not be able to meet its obligations I would make sure my friends and colleagues would be made whole (those relationships are worth more to me than any job ever would be).
There's likely an alternative scenario where a consultant runs into a different but similar set of problems with a company that has mis-configured their Rails app across multiple AWS EC2 machines, in the wrong security groups, with their EBS settings tuned improperly for their MySQL instances. All resulting in extremely poor performance of their flagship application which is costing them a lot of business.
Not a job for the faint of heart, especially when it's your own history you are now fixing, and I appreciate seeing the experience of someone else.
Though typical story of any company/person that assumes a framework are great for their problem, product, not realising what and what not happens in the background. One has to perfectly understand which cogs, axis and wheels turn when an operation is done. Know which wheels always do the exact same thing (apply caches) etcetc.
But more important, best wishes for 2015 from nearby your office,
RB
This kind of thing irritates me. User numbers are important, financially, because "10,000 users daily" can tell an investor or manager how much money is involved. But technically? That number doesn't mean anything to me. Are the visitors making one request or a hundred? Are they clustered into the five minutes before and after a horserace or are they spread out?
As far as interaction goes I would qualify this particular product as halfway between twitter and a social bookmarking site. More interaction than HN but signficantly less complex than twitter. Both twitter and HN are deceptively simple on the outside but remarkably complex underneath, so maybe I'm overstating the complexity level but it's not too far off the mark. By my estimate and using my own websites as a benchmark they should be able to run their current product on a single machine up to or over 100K users daily (using their current set of technologies), session times and concurrency of course play into that heavily.
As for the clueless PM, I have met far too many of these in my travels. If you can't write software, what makes you think you can 'manage' a software development project?
It's a really useful trick to stop yourself believing that the systems works the way you think it does just because you think it.
Deleted comment
This is all stuff that is so basic. I gotta laugh.
And I have to wonder how bad the actual code was...
It sounds like they didn't have a good systems person though nor good (and general) software leadership, often good software leadership is also your early-stage systems person. Jacques here acted as their systems integrator to save the day along with what sounds like mild programming support to cleanup some of the unfinished software product that got pushed too early.