...only? In a sample size of 96, what's the value in analysis beyond that level?
No, this seems pretty solid to me. The math to do this is routine freshman statistics stuff, something every practicing scientist knows. The assumptions just require that wrong answers be roughly evenly distributed across the choices (e.g. you could construct a test where everyone who was wrong would be led to choice C and rarely B or D, but that's a little pathological; and regardless it's something that would be evident in the data set and seems not to have been).
I mean, standards vary but I'll bet in most jurisdictions 100:1 odds count as "beyond a reasonable doubt" for jury instructions in criminal trials. At the very least you'd bring the trial, which is what happened here. If I know there's only a 1% chance I'm wrong, it's absolutely valid for me to accuse you of cheating.