That it can make mistakes? That’s expected, it deal with plausibility like other models no? You would need to look at the actual response and the assigned probabilities to evaluate. Generally demos are not a great way to evaluate a technology, it’s a way to get hype but the next step is to actually look at the details