But then you'd need to answer 100 or more test questions instead of 13. This is necessarily simplistic given how much involvement they can ask of participants.
Also, I don't think they're implying anything nefarious about the resulting biases shown in the final results. For sure a lot of it has to do with the small sample size of the questions. I didn't run the tests twice, but I imagine there is some randomness involved in the way they are generated.
If there are any heuristics at play, the results will indeed show them (in my case there were enough tests to recover the fact that I preferred saving passengers, preferred non-intervention, and preferred saving humans over pets). But it will also come up with some gibberish/noise due to the small sample size.