Go: Fuzzing Is Beta Ready
blog.golang.org
blog.golang.org
My experience (real world, for actually existing code) is that property tests often require a lot of fiddling with the data generation, in order to actually stress your system in interesting ways. If you just throw "totally random" data into your system, you won't be testing very interesting properties. Amazing assurances and payoff from doing it, of course! Just like... "just generate random arbitraries" works a lot less when you're working with 100 field structs and only 10 or so "matter" in senses you care about.
To my knowledge I have not seen a property testing toolkit that leverages code coverage in the way fuzzing does.
They turned it off (not sure if they turned it on again): the problem is that property based tests of the kind implemented with hypothesis are mainly supposed to be run as part of your CI/CD test suit. So they should be fast.
Coverage guided fuzzing like AFL does usually takes longer than people are prepared to wait for in their build/test server.
You run the tests briefly in CI to check for regressions, and then leave them running permanently on a fuzz server to search for new bugs. Nelson Elhage has a good writeup of this approach at https://blog.nelhage.com/post/two-kinds-of-testing/
That said, for me: they are distinct but related, and that distinction is useful.
For example, Hypothesis[1] is a popular property testing framework. The authors have more recently created HypoFuzz[2], which includes this sentence in the introduction:
“HypoFuzz runs your property-based test suite, using cutting-edge fuzzing techniques and coverage instrumentation to find even the rarest inputs which trigger an error.”
Being able to talk about fuzzing and property testing as distinct things seems useful — saying something like “We added fuzzing techniques to our property testing framework” is more meaningful than “We added property testing techniques to our property testing framework” ;-)
My personal hope is there will be more convergence, and work to add convenient first-class fuzzing support in a popular language like Go will hopefully help move the primary use case for fuzzing to be about correctness, with security moving to an important but secondary use case.
In Rust the driver for property and fuzz testing can be shared, which is nice[0].
https://docs.rs/arbitrary/1.0.1/arbitrary/trait.Arbitrary.ht...
By describing my data with arbitrary I've written programs that had traditional property testing and fuzz-testing as well, without any additional effort.
It's only really post-AFL where fuzzing has suddenly become synonymous with instrumentation guided data generation, which is a totally fine distinction to make, but they're still fundamentally equivalent. The first fuzzers were virtually identical to prop test frameworks today.
I'd be fine with dropping the "fuzzing" name entirely and instead just having us use the term property testing, with "data generation" being the thing we start differentiating ie: "random prop testing" or "instrumentation guided prop testing" or "type based prop testing" etc.
Property testing generates these bits using a specified distribution, and that's about it. Fuzz testing generates these bits by looking at how the program is executed, and uses a black box to try to explore all paths in the program.
Most libraries for property testing comes with very convenient ways to craft the "input to data" part. Fuzz tools come with an almost magically effective way to craft interesting inputs. The two combines very well (and have been combined in several libraries).
This is why you can use one of the approaches to help the other side of the approach.
The 3rd solution is concolic testing: use an SMT/SAT solver to flip branches. The path down to the branch imposes a formula. By inverting the formula, we can pick a certain branch path. Now you ask the SMT solver to check there there's no way to go down that branch. If it finds satisfiability, you have a counterexample and can explore that path.
Coverage guided fuzzing is eating symbolic execution and concolic testing for breakfast. It isn't even close. As much as I love these more principled approaches, the strategy of throwing bits at programs and applying careful coverage metrics is just way more effective, even for cases that seem to be hand picked for SMT-based approaches to win.
I think most property testing frameworks also come with the concept of "shrinkage", which is a way to walk back a failed condition to try to find the "minimum requirement of failure". Though I am sure there are PT frameworks that haven't implemented this.
Of all the property based testing libraries, hypothesis has some great ideas on shrinking. (By comparison, Haskell's QuickCheck is actually pretty bad at shrinking.)
This here is closer to what I see as property testing than fuzzing, although it looks like they plan on coverage feedback (so guided generation).
That property can be "does not crash".
And I'd say this is the structured / language-awareness part: with fuzzing you can't generally build an oracle.
And if you can it's of course trivial: just have a wrapper script check the result against the oracle, and trigger whatever the fuzzer looks for indicating "failure" whether it's a return code or a segfault or…
It's been a multi-year effort, so congrats to those who've made it happen.
What I found over time is that is erodes the trust of the team in the test suite. When you work on feature A, the last thing you want is for the tests of unrelated features B and C to randomly fail and prevent merging your work.
So what becomes the norm is re-running the tests until is passes rather than understanding it and fixing it.
Over time, such issues accumulate and it becomes harder and harder to be lucky enough for it to pass.
While the right thing to in theory do would be to always fix it, it's not always possible depending on the organization of the team and tasks, and being bothered by unrelated stuff during a specific task is just annoying.
That said, this package could be useful if pseudo-random (maybe it is, I didn't look so much in details).
Also, this would not work with most (all?) CI runs (which are inherently stateless), and working around this would be complex and introduce a lot of negative side effects.
(Saving the seeds to the corpus still leaves your tests independent and stateless. You are just sampling from your space of all possible tests a bit more intelligently in future.)
Instead of making it part of the base runner? Maybe so the beta bits can work and it’ll be folded in more directly afterwards?
Also if it includes coverage guiding being able to know that a run is fuzzing or non-fuzzing would avoid having to include the fuzzing instrumentation for non-fuzzing run, mayhaps?
> Is there a way to fuzz non-string inputs?
From the design doc it looks like you can have any number of parameters and
> Fuzzing of built-in types (e.g. simple types, maps, arrays) and types which implement the BinaryMarshaler and TextMarshaler interfaces are supported.
It's good to hear that it's not limited to strings.
The fuzzer input is random data. You don't need a special library to generate random data. In fact that's what the post tells you with respect to type compatibility:
> types which implement the BinaryMarshaler and TextMarshaler interfaces are supported
these are just tools to convert from/to binary (unstructured or utf8).
Yes, and this one does, but you requested a library to generate inputs didn't you? A library can't get coverage feedback from the SUT, which as I wrote in my previous comment would be why you'd be building a dedicated test framework.
Fuzzing usually revolves around strings because of escape characters, escape sequences. There is a much larger set of string characters than there are for the 10 or so numeric digits. Numbers don’t have the same problems that strings do, because numbers are usually interpreted only as data, whereas strings can be interpreted as data or computation.
Not always. AFL has been used to detect issues around processing plain old binary data (eg https://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2014-8637)
I would argue that anything involving a parser of some description (either binary or text-based) is a good candidate for fuzzing.
I'm guessing that you won't always want to execute these alongside other tests. Go has also taken the same approach with benchmarks.
And you gotta start somewhere. String inputs is a good start, and you can use those to test other inputs by factoring through conversion functions.
Edit: and here's the mechanism that guides mutations towards increased coverage: https://github.com/golang/go/blob/5542c10fbf19cb199d1659c189...
I'm glad this is being made, but like many other things that have been added to Go, it shows the limitations of the language that you can't just build this inside the language. (I might be wrong, as I haven't had a chance to look at the design docs/implementation yet, but the installation instructions imply that's the case).
Not sure why you'd make that assumption. https://github.com/dvyukov/go-fuzz
When writing parsers and compilers it has proven eerily good at identify corner-cases (panics, and infinite loops).
I'm looking forward to trying the new approach out. Anything that makes fuzz-testing easier to configure/maintain and spreads awareness is a good thing in my book.
Fuzzing is similar but typically involves starting from a known-good input then randomising it at the byte level (irrespective of validity). This project allows for property testing-like unit tests, but tools like american fuzzy lop focus on detecting whole application crashes.