The Saturation Effect in Fuzzing
blog.regehr.org
blog.regehr.org
- the number of possible bad packets is literally billions of times larger than good ones - pretty soon the percentage of those that trigger bad behaviour gets close to 0
E.g. last time I fuzzed a network element with AFL, it took seconds from it to go from a starting corpus of a single ethernet IPv4 SYN packet to some double-encapsulated IP-in-NSH-IP-in-NSH-ethernet monstrosity that triggered a misparse. And seconds more for it to generate a IPv6 packet with a fragment extension header that triggered some other problem. A random walk would have no chance of finding that.
Btw, and maybe I'm misunderstanding what you wrote, if you were generating random packets without fixing up the checksum, you were already wasting basically all of your testing capacity. All that it ends up doing is checking that the negative case of checksum verification works.
I have seen it consistantly produce "magic" strings; presumably by walking the strcmp function calls, where each individual character comparison is another oppurtunity for execution to take a different path.
all modern fuzzing is coverage guided for this exact reason
Is anyone aware of fuzzers that take issues like this into account by explicitly trying problematic values, like INT_MIN for an integer variable?
Even without a dictionary if your input is reasonably small it should discover these special values given enough iterations.
I'd always run end to end tests on random real data records.
I don't see many people doing this, not sure?
There seems to be kickback because tests might fail one run and pass another. People don't seem to like this.
If you can't repro a crash a month or a week later, it's not worth it IMHO.
Usually you can get most of the benefits with data generation, based on a deterministic seed.
Mainly though, you have to ask about the purpose of the tests. Most developers want a test suite that tells them with high confidence that what they just worked on didn't create a new bug, or a regression. That lets them stay focused on their work rather than chasing through the codebase for an unrelated latent issue which might have been introduced by someone else, years ago.
There is often value in separating "exploratory" or "stochastic" tests which might uncover new bugs (previously unknown) from regular tests. To make it really work as part of default workflow you need a culture which understands a feature might take additional time because the team stopped to fix a latent issue to get back to green.
To put it bluntly, letting random-ish testing break your pipeline is making a statement of business priority (we care so much about random bugs that we will stop all other work until they are fixed) which might not align with reality.
"Is this the most important thing for me to be working on for the success of the business?"
I think the "right" way to add fuzzing / random testing to a pipeline is sell the value to the business and have an initiative to rigorously fuzz the snot out of the software. Crucially, have resources dedicated to triage and fix the identified issues. It should be a non-breaking pipeline stage right up to the point where everyone feels confident that any test failures are the result of new code.
The worst case is that an opinionated developer adds stochastic tests, with bad reproducibility, without consulting the team and randomly breaks builds in an organisation that only rewards or understands feature delivery and ticket punching. That is basically going to make their co-workers life hell and is IMHO not the right way to go about it.
I have also encountered this attitude and haven't understood it so far. What stops you from recording the test cases to reproduce failed ones?
Definitely using this idea. Thanks
We call them:
- market data feeds
- orders feeds
- users
- developers
YMMV
/sarcasm ;-)