What is a possible use case for a legal reasoning tool that is wrong 46% of the time?
It’s barely any better than a coin flip
waiting for the actual Prodigy Corp. to show up and offer to buy your username out?
<my own consternation, stated>
54.0% isn't particularly high.
What I'm trying to say which doesnt seem obvious is that all models are x% correct at benchmarks until they get saturated, and then new benchmarks get made.