I would think that they would attempt sending all possible characters, especially because they've had issues with this in the past.
I would think that they would attempt sending all possible characters, especially because they've had issues with this in the past.
To be honest, though, if someone on my team suggested we implement an automated test that tries sending every Unicode character (and it would be applied to every interface of every app and API, right?) I would have objected that this was an over-complicated, over-engineered solution that will probably be too slow.
I'd argue that a set of test data selected to cover a range of patterns, especially ones that are considered risky (either is know to have caused problems in the past or appear complicated or tricky) would have almost as good coverage and be an order of magnitude more useful.
The problem with "run-every-case" tests is that they start off slow and get exponentially slower if you try to go deep you very quickly end up with tests that take too long to run to be useful. (e.g., {every Unicode character} is one thing... {Every Unicode character} X {every interface and app} is a LOT more. {Every Unicode character} X {every interface and app} X {every build} X {every device model} X etc. is impossible). So you end up with very broad but very shallow tests. And that means less coverage ultimately, not more.
Generally, I think you'll get more efficient, effective, useful tests if you tailor the test data sets to the problems you see rather than going for blanket coverage.
Sure, but fuzzing and "throw everything and the kitchen sink at it" type testing will generally expose logic bugs you weren't aware existed. Aka the unknown unknowns. Its not like this is an either/or proposition, you could run the fuzzing type gauntlet tests every week or so.
But I agree it would be unreasonable to do that for every app that uses the library.
I don't think it's a no-brainer though.
And it is unreasonable to expect a testing technique to be applied in every case where it is reasonable to do so. There are a lot of techniques that are reasonable to use in any give case but it would be no sense at all to apply every one of them. To focus on this one, now, amounts to Monday morning quarterbacking. Once you know the bug it's easy to see how it could have been caught. Of course if you somehow knew of a bug ahead of time you wouldn't need any test at all!
IMO fast probabilistic total coverage test is better than no test or partial coverage.
Yes. This seems like a somewhat embarrassing fail in terms of whatever automated testing Apple does. It's a pretty minimal and restricted case of fuzzing, in fact I'm not sure it'd qualify as "fuzzing" at all because here the crasher is just an actual Unicode character, part of a fixed set that can be completely run through in deterministic time. Unicode is certainly large in human terms but in terms of automated test data sets it's not.
Given the notorious difficulties and edge cases Unicode parsing has long created, testing every single individual character in Unicode as part of standard unit/regression testing seems like something that should just be done as a bare minimum for an operating system or parsing framework/library release. That's not to say there wouldn't be more complex inputs that would cause problems that sufficient fuzzing could discover, but this kind of a error from a single real language character shouldn't have slipped through, it's not random.
Unicode and image/document handling frameworks are both areas that should have extensive fuzzing as well as deterministic input sets for testing. They shouldn't need to be messed with much either once they're "done", so in the case of an org like Apple this is the sort of place where it might even make sense to really devote resources to getting at least some parts of it formally verified.
To be specific, here are the codepoints in the character from Pastebin:
U+0C1C [Lo] TELUGU LETTER JA
U+0C4D [Mn] TELUGU SIGN VIRAMA
U+0C1E [Lo] TELUGU LETTER NYA
U+200C [Cf] ZERO WIDTH NON-JOINER
U+0C3E [Mn] TELUGU VOWEL SIGN AA
I don't know if the ZWNJ is in there to make the bug happen or to neutralize it. But the entire group of five codepoints renders and selects as one character for me (in Chrome on Ubuntu).Likely the next two programs are disagreeing on character, and therefore widget size.
I've used AFL [1] to generate semantically valid, crashing multi-statement SQL queries when given ~50 CPU days and a simplistic SQL-like database. All from thin air, with an empty text file as the initial testcase. It can also come up with valid JPEGs [2] given enough time.
[1] - http://lcamtuf.coredump.cx/afl/
[2] - https://lcamtuf.blogspot.ie/2014/11/pulling-jpegs-out-of-thi...