I have also used Haskell's `QuickCheck` and Clojure's `spec` / `test.check` and have had a great experience with these. In my experience they "just work".
Conversely, if you're trying to generate non-trivial datasets, you will likely run into situations where your specification is correct but Hypothesis' implementation fails to generate data, or takes an unreasonable amount of time to generate data.
Example: Generate a 100x25 array of numeric values, where the only condition is that they must not all be zero simultaneously. [1]
[1] https://github.com/HypothesisWorks/hypothesis/issues/3493
Silly idea for your generator would to generate an array, and if it's zero... draw a random index and a random non-zero number and add it into the array. Leads to some weird non-convexity properties but is a workable hack.
In your own example you turned off the "data too slow" issue, probably because building up a dataframe (all to just do a column sum!) is actually kind of costly at large numbers! Your complaint is probably actually meant for the pandas extras (or pandas itself) rather than the concept of hypothesis.
But! Even though it doesn't even get that much slower at a certain number of rows it just starts hanging! Like at 49 rows everything is still fine and at 50 it no longer wants to work. It's very bizarre and I'll see if I can debug it. But I think your test case isn't indicative of some fundamental issue with Hypothesis rather than some sort of bug.
Let me know if you have a reproducer, I'd be curious to take a look.
Yes, this brings back memories. I've definitely seen this kind of behaviour as well, in many different, not-particularly-exotic, situations.
I am absolutely convinced the issue I raised on the github project was a bug or a defect, despite the maintainers not taking it seriously.
I find QuickCheck and Clojure spec/test.check much more straightforward to use. I just never ran into this sort of thing with these other tools.
(a) Filtering is a last resort and is best avoided. As an example, the Gen type in Haskell's falsify package can't be filtered, since it's a bad idea. As another example, ScalaCheck's Gen type can be filtered, but they also allow "retries" (by default, up to 10,000 times), because filtering is very wasteful.
(b) If you're going to filter, scope it to be as small as possible (e.g. one comment points out that you're discarding and regenerating entire dataframes, when the filter only depends on one particular column)
(c) Have some vague awareness of how your generators will shrink, to avoid infinite loops. In your case, shrinking will make it more likely to fail your filter; and the "smallest" dataframe (all zeros) will definitely fail.
My initial impulse is to pick a random cell which must not be zero, generate a random number for each other cell and a random non-zerp number for that one. I'm not immediately decided on whether it's uniformly distributed.
Any algorithm that cares about the number of non-zeros could have non-trivial interactions with their arrangement and count, so picking something that generates non-trivial sparsity (and doesn't just make the array look like white noise) is going to have the best chance of exposing interesting behavior. The tricky part is thinking through how to generate "interesting" patterns, which admittedly I haven't put enough thought into.
Right! And even if it were, in the sense that that's what we should expect as real world input, it wouldn't generally be the best distribution for finding bugs.
Haskell's "falsify" package takes a similar approach, but uses a tree of random values. This has the advantage that composite generators can run each of their parts against a different sub-tree, and hence they can be shrunk independently without interfering.