As far as I can tell from looking at the code that implements it, and ignoring the training for the moment, the model itself takes a binary vector, and takes the sum of (the characteristic function of) AND's of elements of that vector, and returns whether the sum exceeds a threshold.
So an example model with an input vector `x` might be
x[1] & ~x[2] & x[4] + x[2] & x[4] + x[1] & x[3] > 1
EDIT: okay, so that's the effective evaluation model, but for training purposes, it is stored internally as a series of real numbers for each potential coefficient, and if that number exceeds the "number of states", then it is considered active. So the model above in binary form would look like
+: [1,0,0,1],[0,1,0,1],[1,0,1,0]
-: [0,1,0,0],[0,0,0,0],[0,0,0,0]
But internally this could be represented (with number_of_states == 6) as
+: [7.1,5.0,4.0,8.0 ],[3.7,6.1,5.9,7.8],[8.0,3.0,6.5,5.2]
EDIT 2: Each term apparently is assigned a sign flag, and the threshold is 0, so not strictly an increasing sum.