Elician here!
Accuracy and supportedness of the claims made in Elicit are two of the most central things we focus on—it's a shame it didn't work as well as we'd like in this case.
I'd appreciate knowing more about the specifics so we can understand and improve