Scientist: Measure Twice, Cut Over Once
githubengineering.com
githubengineering.com
https://blog.twitter.com/2015/diffy-testing-services-without...
A few reactions from reading through the release:
Scientist appears to be largely limited to situations where the code has no "side effects". I think this is a pretty big caveat, and it would have been helpful in the introduction/summary to see this mentioned. Similarly, I think it would be nice to point out that Scientist is a Ruby-only framework :)
You don't mention "regression test" at any point in the article, which is the language I'm most familiar with for referring to this sort of testing. How does a scientist "experiment" compare to a regression test over that block of code?
Anyway, thanks again for writing this up, I'll be thinking more about the Experiment testing pattern for my own projects.
That's one of the things I was initially thinking too, but then as I thought about where I could have used it in the past, I can think of only a few cases where it wouldn't be possible to keep it isolated.
For example, have your new code running against a non-live data store. Example: When a user changes permissions, the old code changes the live DB, while the new code changes the temporary DB. Later (or continuously), you could also compare the databases to ensure they're identical (easier if the datastore remains constant, a bit harder if you are changing to a different storage schema or product).
Where it would be in the difficult-to-impossible range is when touching on external services that don't have the ability to copy state and set up a secondary test service, but even in that case, you could record the requests made (or that would have been made) and ensure they both would have done the same thing.
Definitely an interesting concept overall.
You can only test hypotheses of the form "A is exactly like B" - no bug fixes are allowed, because they will show up as differences.
So a more accurate (but less cool) name might be "Refactoring" - you assert that all your tests still pass, where your tests are your production data.
Of course they will. That's a feature. The system doesn't intrinsically know whether a change in behaviour is a bug added or removed, it can only report that there's a difference in behaviour, it's your job to investigate whether the control or the experiment is correct.
f(x, y) = f(y, x)
without a reference implementation.However, I don't get the restriction on code with side effects.
Would it not be possible to introduce another abstraction layer around those side effects to allow comparison between the old code's side effects and the refactor's code side effects?
For example, suppose you have a db object and two versions of the code new_code and old_code. You call something that looks like:
experiment.run(new_code, old_code, mutables=[db])
Then the infrastructure runs new_code normally, but records the arguments and return value of every call to db (and any other object defined as mutable). Then, the infrastructure runs old_code, but whenever a method of db is called, it tries to match it with a call made by new_code and just directly returns the return value that call returned. If it can't match the call, it signals an error, but it never actually tries to call db, thus negating the risk of side effects.
Obviously this would fail when the two versions perform different operations in the database, even semantically equivalent, but non-identical operations (say one retrieves a value and increments it inside a transaction, while the other uses a stored procedure in db to increment without fetching). But it still relaxes the constraint, now you can do this for code that has no side effects and code that has the exact same side effects as represented by exact call/return pairs to mutable objects.
I'd think the latter would be the common expectation by virtue of a different implementation.
* An STM implementation (ruby doesn't have one afaik)
* All network services you connect to (eg databases, APIs etc) would need to support transactions for write operations
Looks like a great tool! We'll give it a spin.