Performance Culture
joeduffyblog.com
joeduffyblog.com
- The cultural transformation must start from the top
- X often regresses and team members either don’t know, don’t care, or find out too late to act.
- Blame is one of the most common responses to X problems
- Performance tests swing wildly, cannot be trusted, and are generally ignored by most of the team.
- X is something one, or a few, individuals are meant to keep an eye on, instead of the whole team.
- X issues in production are common, and require ugly scrambles to address (and/or cannot be reproduced).
Unexpected performance problems typically are not that hard a problem to deal with once you can measure and find the bottlenecks. This tends to be local, rather than global. The other kind of performance problem is scaling - traffic is up, things are slow, and the architecture needs to be scaled. That can be hard.
I've seen features that were regression/unit tested and code-reviewed to death fail in production because all that testing was done at < 10 transactions/sec with a dataset of about 100, when typical production values were in the neighborhood of 1000/tx/sec with a dataset of perhaps 10,000.
At the very beginning of the design process, someone needs to ask, "How frequently will this feature get invoked? What kind of latency can the user tolerate?"
Obviously you need good infrastructure to do so.. But you can't fake a prod exchange. Too much bandwidth and too many users each on their own agenda.
cost of testing w/ production-like data > cost of potential production issues
and decide whether it's prudent to omit or scale down the tests.For an HFT system I'd have the dev version doing overnight testing with the day's data or chunk thereof, in delayed-realtime mode.
I'm starting to tell people "go beyond TDD with PTSD (performance testing stipulated development)". It will give your machine trauma, but it helps gamify development in an area that is otherwise dry and boring. Even worse, our company, GUN, does JavaScript development (we're an Open Source Firebase), and JS engines are sadly slow, but it is critical to our engineering team. In the process, I've collected a humorous sample of some very basic JS operations that you can run yourself, starting with nothing all the way up to loops and closures - http://db.marknadal.com/ptsd/ptsd.html . Hope you enjoy it and don't get too depressed.
Particularly as you mention trying to create positive buzz for your team.
I love a weekend hack as much as anyone, but I'm assuming this is unpaid and uncompensated... for FOSS that might be fine but for a multi-million dollar project? No way.
Manager B would most likely be politicked out of his/her job.
The trouble is when illegitimate (generally political) business concerns are allowed to be trumpeted as more important than legitimate business concerns, or when illegitimate views about how to discount the future value of quality are permitted to dominate the comparison with extreme short-term opportunities.
This assumes that you are building something you will be using in that future which for most startups is not true. For most it will be a better trade off to iterate faster now and solidify the system once you know for a fact that it matches actual needs.
I tried it a little while ago and I had to put Windows into test mode or something like that, and write lots of code On linux it seems to be much easier (for example just run 'perf').
Edit: Joe - If you read this, please ask the Microsoft kernel team to expose the CPU hardware performance counters to userspace. Thanks!
Also, when is M# / Midori going to ship?
Also some of the M# features are coming in C# 7 and later versions.