Distributed Systems Testing: The Lost World
tagide.com
tagide.com
I don't understand the stigma attached to logs and debug-by-print. When is a tool that uses x to provide y 'glorifying' x?
I've built and regularly use two tools that rely on logs:
a) I write white-box tests that allow me to check not just what a function returns but specific events in the course of the function running: http://akkartik/name/post/tracing-tests. This allows me to check, for example, that searching in a sorted array doesn't require double the lookups when the size of the array doubles. It also allows me to test subcomponents without calling them directly, leaving their interfaces underspecified and flexible to change later without needing to modify a bunch of tests.
b) I often debug my programs by dumping a trace and running it through a 'trace browser' which starts out showing just the coarsest level and allows me to drill down in specific places as I want. This 'zoomable' google-maps-like UI gives me all the benefits of time-travel debugging at a fraction of the system engineering it usually takes. (I still can't get the feature to work in gdb..) Try it out:
$ git clone https://github.com/akkartik/mu
$ cd mu
$ ./mu browse-trace .traces/factorial-test
It'll take ~30s to compile the first time you run it (C compiler required on Linux/OS X). Once it compiles, you'll see an ncurses program. Use 'q' to quit, hit enter to zoom into lines (you can see how many lines are collapsed in parens), backspace to zoom out on a line.These are both things nobody can do on a conventional toolchain, and super easy to build. Maybe print seems trivial because nobody's building tools around it?
And if you have an outage where you didn't have sufficient monitoring, then you add it afterwards. But at least by breaking it intentionally you can at least be watching.
This is the entire idea behind Netflix's Failure Injection framework (http://techblog.netflix.com/2014/10/fit-failure-injection-te...) which was briefly alluded to in OP's link.
1. Making things that actual work is industry, which writes very few papers comparing to academia. Academia cares more about design, less about working large scale systems.
2. Testing is more about well-oiled machinery, less about some super clever algorithm. Testing is more of a craft than art.
When the focus is pushing out Grade-D goods what's the point of rigorous testing? If you get it wrong, it barely matters. If it matters at all.
Perhaps the social view of software testing is just a corollary to the state of software engineering?
Because as more feature-laiden software is written to facilitate growth in consumer and industrial markets that is deployed on something more akin to an embedded device than a PC or server there's going to have to be more rigor applied to the correctness of the software to compensate for how tightly coupled it is to the largely immutable use-specific hardware.
However, I think there's a growing surface area where our consumer expectations for iteration and features are going to bleed into places that were previously left largely untouched by such demands. "Smart" cars are probably a good example. There's growing and intensifying competition for the ownership of the experience in the cabin/cockpit of the car, and in order to draw lines of competition there's a lot of focus on how to integrate that space with the rest of your normal and your digital lifestyle. That ends up being a place where, to avoid ending up the subject of very public screw ups, some amount of rigor that's applied to the lower-level stack will also eventually bubble up to the feature/application level at the same time that enablement for features/applications (network connectivity, new kinds of input devices, etc.) will end up being forced into the lower levels of the stack. Or at the very least a useful set of stable abstractions and tools will be built to prevent today's software developers from having to care very much about applying that rigor and the abstractions/tools will just take care of it for them.
Also, I would argue that academia is about 7 years behind industry in this area, and so papers don't match up to reality well.
On the other side, industry testing is usually very specific to the internal product created, and so there is an incentive not to publish anything.
There is some nice academic research outside of CS, notably mechanical engineering, on systems testing. It is pretty applicable. Contracts, sub-unit testing, etc are all anaolgous.
My two cents.
(We work on an open source distributed data sync engine and needed a way to test distributed aspects of that engine - https://github.com/amark/gun/)
I think this gets closest to the truth. That first slide regarding "testing a microservice architecture" with its complicated block diagram hints at why the level of abstraction is too high to be useful. In the end a distributed system consists of components that have interfaces and which interact with each other through those interfaces. You test the component modules in unit testing, the components themselves at the interface level, and the whole system in end-to-end request and response flow. That latter is assisted by centralized event collection and storage tools like logstash + elasticsearch, but it could just as easily be dumped log files. Whatever works. That's the closest I've personally come to "distributed systems testing."
Once that's in place, write tests with selenium or unit test framework with parallel extensions (depending if cases are frontend or backend based) and analyze using home made tools.