Testing Firefox more efficiently with machine learning
hacks.mozilla.org
hacks.mozilla.org
I assume the overhead of the project (and subsequent tweaks to model, re-training and validation) is sufficiently negligible compared to the measured benefits even if those weren't as clear-cut as 70%. I'm unaware of how much compute is required for the task, but likely less than many compute-years per day :)
One thing I did not notice in the approach to the modelization of the problem is any link/tag regarding the platform for which the code changes are made, and the programming languages used. There seems to be some evidence that certain languages could lead to more defect fixing commits[1], and I don't know if there's evidence that some platforms are more prone to bugs (I'm sure wars of words have been fought over this). But would it make sense to have that sort of information inform the model in a way? I fully understand that I might be out of my depth here.
[0] https://firefox-source-docs.mozilla.org/setup/index.html
[1] https://cacm.acm.org/magazines/2017/10/221326-a-large-scale-...
I wouldn't be surprised if language is explicitly included as, for example, a flag in the "data" object. But the model should be able to figure it out by itself otherwise by identifying keywords that only some languages (often) use.
The cost of a bug slipping through because a test being skipped will be higher than running an irrelevant test to a commit.
Yes a regression slipping through would far outweigh the benefits of reduced tests. The thing the post didn't make very clear is that thanks to our integration branch, the chance of a missed regression is still nearly zero. If the scheduling algorithm misses something, the failure will show up on a "backstop" push. These are pushes where we run everything, and then a human code sheriff will inspect any failures, and if something was missed figure out what caused it and back it out.
So the costs of missed regressions are: 1) More strain on the sheriffs (too much strain means the need to hire more) 2) More backouts which is annoying to developers and can mess up annotation (though we have ideas to fix the latter).
For the record, the algorithm with the 70% reduction in tests has a regression rate almost on par with the baseline (it's ~3-4% lower). This hasn't seemed to result in much additional strain on the pipeline.
10 core-years per day sounds like a lot but it's only about a 10kW load, and they've saved 70% of that, or about $20 of opex per day.
* Many tasks run on expensive instances (hardware acceleration, Windows)
* We have OSX/Android pools that run on physical devices in a data centre (these are an order of magnitude more expensive than Linux)
* There are ancillary costs. For example each task generates artifacts which incur storage costs. These artifacts are downloaded which incur transfer costs.
* There are also overhead costs (idle time, rebooting, etc) that aren't counted in the 10 years / day stat.
All these things see a corresponding decrease in costs with fewer tasks.
I get about $1000/day based on some EC2 prices for typical machines I've used, though I'm sure Mozilla's requirements are different and they can negotiate better prices than I can.
It really depends on the type of bug, and perhaps this could be factored into the model by also correlating change sets with outage severity or complexity of a fix.
It's for mobile apps mostly though.
> In the past, how often did this test fail when the same files were touched?
> How far in the directory tree are the source files from the test files?
> How often in the VCS history were the source files modified together with the test files?
But for prediction all they input is a tuple (TEST, PATCH), and XGboost works fine without the additional features?
The documentation is tantalizing, but hilariously short: https://devcenter.heroku.com/articles/python-rq
Very "And then draw the rest of the owl." Oh really, you can just do `from utils import count_words_at_url; q.enqueue(count_words_at_url, 'http://heroku.com')` and presto, your blocking function -- whose source code exists locally -- is run successfully at the other end?
I'll have to set aside some time to try this out. Python does have introspection facilities that could make that possible. I could imagine that since the code is executed on the same box, it's relatively simple to send a request like "here's which module the function was loaded from; here's the order all modules were loaded in; load those modules and call this function." But it leaves so many questions: serialization, performance, scaling, and all the tiny bugs that inevitably come up.
I guess I was hoping someone could give me a quick gut check of positive/negative reactions. The full RQ documentation is slightly better: https://python-rq.org/docs/ but has some worrying signs:
Make sure that the function call does not depend on its context. In particular, global variables are evil (as always), but also any state that the function depends on (for example a “current” user or “current” web request) is not there when the worker will process it. If you want work done for the “current” user, you should resolve that user to a concrete instance and pass a reference to that user object to the job as an argument.
Yes, sure, global variables are the root of satan, but they're also a fact of life in many scenarios.
Interesting approach... I wonder how much of a nightmare it makes devops...
* serialization: the input was passed a json string argument. the output was a file uploaded to S3, so just the URL was returned, again in a JSON string * global variables: the program was quite self-contained: there was an initial state setup that was not mutated afterwards. So RQ's fork-exec model(the default) worked well enough
Sorry I don't have much to say about performance and scaling. It was quite fine for our needs, and we could scale horizontally upto a certain point by just starting extra processes, and beyond that with more VMs. Since they all listened on the same RQ, it worked fine. (the number of items in our queue never really hit any of Redis' limitations either)
RQ lets you customize the worker model: so you could for instance use threads instead of processes.
Regarding monitoring: there's RQ dashboard[1] which gives a nice web interface to view jobs, failures, and restart them.
I guess what I'd recommend to Mozilla, although fantastic work, is that's it worth trying to compute certain features (data points) upfront to simplify the main model. These calculated features can be used as inputs to the main model possibly reducing the response time.
If the consensus is its stable enough to give it a go though, I’ll give it a go.