The proof is in the literature I'm referencing rather than me dumping all of it in the comment box. Anyone studying PL, QA, and so on literature should have seen research on effectiveness of basic techniques. Here's a few examples of methods of evaluation I claimed were scientific vs people guessing that their stuff seems better in vague ways.
1. Experiments conducted between similarly skilled groups doing same things with only difference being language or methodology. Those with new technique do significantly better in ways that match the expected properties of technique. Hypothesis based on real-world data, experimental confirmation, fix/reject if necessary, rinse, repeat.
2. Study of problems in software from a specific language identifies specific features as root cause, proposes a fix, the fix is implemented in significant test cases, and the problem goes away entirely or is significantly reduces. This is replicated across groups and codebases. Easiest example here is safe (Ada) or automatic (GC) techniques to manage memory to reduce or eliminate its many, common issues. Supported by theory, sometimes proofs, & countless applications in industry.
3. A certain verification approach espouses specific modeling, analysis, refinement, etc approach to achieve certain goals. The work on formal side translates into finding defects, enhancing performance, etc on the code side. Implies the model or method were correct, esp with diverse replication of results. This happened a lot in Orange Book days for catching security violations but most applicable to hardware with formal, equivalence checking. Theory, some proofs, and confirmation in production.
4. In some cases, properties are mathematically formulated, proven in theorem prover, checked by humans + machines, and tested on real-world examples to further validate them. Two words: type systems. Again, theories, proofs, and real-world validation. Tons of testing, too.
Most of the best evidence in the literature works something like that. You're saying none of the above constitute scientific evidence for a claim about a software tool or method? That's a big claim itself that contradicts what most of research community thinks. The people successfully finding bugs with Astree, avoiding out-of-bounds with languages that don't allow it, catching safety violations with protocol checkers, and so on will all be surprised about your claim. They and I thought the tools would work based on evidence presented and field results that confirmed the theories. Turns out we were all drinking spiked Kool-Aid, leading to hallucinations that followed. Or wrong on a technicality.
Got a link that briefly explains why all of that is unscientific and what methods count as scientific evidence in software tools? (without 1hr videos)
Note: Looking up more on Quorum is on my agenda. Already saved What Works Clearinghouse document. I'm not rejecting it outright in our discussion. Yet, you've indicated proven models, worked case studies, analysis techniques effective on 1+ million plus samples, mathematical proofs... that none of this is science. So, I'm starting with wanting an easy resource on what you think counts as scientific evaluation, argument, and method. Then, I can go from there in future discussions about specific topics and what science says about them.