I find it interesting how this is premised on the "one big claim per paper" model, which for better or worse seems increasingly common. In my experience a lot of 60's-90's research is of the form "the results of our first study were consistent with [some specific paper], and our second experiment was not." This nuance is pretty hard for people to keep in mind, and often seems to get lost over time (as a postdoc in cognitive science at MIT, I've noticed a big difference between what papers actually say, vs. what people say those papers say 30 years on). I'm excited to see how well Scite actually works as I test it over the next week!