Why isn't there a replication crisis in math?
jaydaigle.net
jaydaigle.net
Math is an arbitrary framework, built from arbitrary axioms. It is deductive. Thus, all proofs are simply deduction. The knowledge here is positive knowledge, we can show things are true, false, or undecidable. There may be errors, but those are errors in execution.
Psychology is not built on a framework. It is inductive, we are trying to find axioms that map to the data we collect. Thus, all papers are trying to add/build it's arbitrary framework. The only knowledge here is negative knowledge, falsification, we know only what is a failed hypothesis. There will be errors, both in execution, and there will also be statistical errors in experimental results.
The entire point of the replication crisis is that we don't publish or pay attention to results that are boring, so the framework we build is built on skewed data. We don't reject previously popular papers that are now unfalsifiable (the idea that the now unfalsifiable Milgram experiment is still taught in every university psychology dept should be outrageous). The boring results need to be weight statistically against the interesting results, but aren't, etc. Nobody out there is arguing whether or not the axiom of choice is a true axiom? It sort of doesn't matter, it can't matter, because it's arbitrary by definition.
You can't have a replication crisis inside of a deductive framework without changing the framework. This doesn't happen too often, but we did see this during the shift from neutonian to einsteinian physics. The study of philosophy of science is fairly obscure, but is the center of this discussion.
"The results can't be replicated" is different from "the logic here is wrong". So different that this article starts from an entirely invalid premise.
> More seriously, it’s reasonably well-known among mathematicians that published math papers are full of errors.[1]
I think the author's point is that actually the distinction is less clear: many math papers can't be (easily) replicated, yet people in math aren't too worried.
There isn't even so much a notion of "results", just "logically, Y follows from X". If that's wrong, you don't need to run the experiment again (there is no experiment to run), the logic is just flawed.
While it's possible a mistake won't be found, it's a much different problem than running empirical experiments and getting different results.
It’s obviously quantitatively different in very important and meaningful ways, but the author I think is drawing an apt analogy.
I remember reading an article in the intelligencer by a mathematician who essentially said there are certain conjectures where if they were to be proved false rather than be thrown into a sea of uncertainty, mathematicians i would quickly move to investigate a readjustment of basic axioms rather than accept that those conjectures are incorrect.
Then there are fields of mathematics around selecting different axioms. Investigating the ramifications of whether you take the undecidable "continuum hypothesis" as true or false. And then theres model theory and such. Presumably they study models of interest and not arbitrary ones.
You're mostly correct that the methodology is mostly deductive but the point is that what we choose to use isn't arbitrary because there are things in math which are more important than axioms as they are things believed to be "real".
edit: I'll try to expand on this. Math is weird because it's a sort-of in-framework study. We use math when we do physics, and we use physics when we do engineering.
If the physical world ever started disagreeing with our physics, then we need new physics. If the math ever started showing inconsistent results in our physics, then we would need new math.
None of that is to say that Neutonian Physics, say, is "wrong in the sense that it's internally inconsistent", only that "it's wrong in the sense that it is not an accurate mapping of a framework to the world we find ourselves in." The first type of wrongness is the type of wrongness we concern ourselves with in deductive studies, the second type, the mapping-error, is inductive, and requires meta-analysis, and is prone to (Hume's) problem of induction, which means it's ultimately unknowable (i.e. Popper).
> You can't have a replication crisis inside of a deductive framework without changing the framework.
You may change the framework, but the replication crisis would be that people don't notice it's changed.
A proof can get so huge and complex that it would take a lifetime of study to understand it. One of the article's good points was that "you can replicate a math paper by reading it", but if you cannot read a proof, you cannot replicate it. If nobody can understand it, nobody can replicate it. There would be a big crisis if people started trusting huge papers blindly and using their results without stating them as assumptions. Mathematicians are largely not doing that, so there isn't really a crisis. But they might! And you would not notice that the framework had become inconsistent for a while. Therein would lie the crisis.
It is nice to think that as soon as the framework changed, people would notice, because surely if it the ground moved, that would be obvious! But it is not true. Nothing about the deductive-ness or in-framework-ness of mathematics changes this. It is the same for psychology, the replication crisis is a crisis because people do not notice these results are wrong, and then rely on them in clinical guidelines and affect people's lives for generations.
It is possible to stave off replication crises by using a machine, because mathematics is self-contained and you can use a computer to show that you proved something using only these specific assumptions and they can believe you without reading and understanding it fully. But this doesn't mean mathematics cannot ever have a replication crisis.
I think your arguments here are best summarised as confusing "replication crisis" with "any published results being inconsistent with previous results". Your idea of "inconsistent" is also strange, the article considers this to mean mistakes in proofs, whereas you mean new mathematics that changes the field by discovery. You're just talking about something completely foreign from the concept of a replication crisis.
They even found a flaw in it, but Wiles was able to fix it after a year.
Deductive frameworks can’t be wrong in the way that experimental results can be statistical anomalies. The author conflates errors with perfectly reasonable experimental results later shown to be anomalous upon replication.
You can't define your way out of the fact that mathematicians will make mistakes and write papers that draw conclusions that are not valid. They are just pretty good at discovering the mistakes, and finding them is culturally essential to the practice of mathematics; arguably if it were not, it would be a big game of Numberwang. Psychologists are, on the other hand, not nearly as good at discovering mistakes. We didn't have to use any properties of the self-containedness of the disciplines to get to that result. It was not necessary to prove mathematically that mathematics does not or cannot have a replication problem. You can just observe it.
(Edit, all of that in brief: if it were truly impossible to be wrong in maths, then why do mathematicians spend so much time trying to verify each others' proofs? Are they wasting their time? Or is it the entire reason there isn't a replication problem?)
As the author stated, the math can be reproduced by reading it: ergo, thee is no barrier to entry except knowledge. The other sciences do not have this luxury - they require access to specialized instruments, participants, funding, and most importantly - time. Math can be checked nearly instantaneously, so the "might" becomes inconsequential my small.
Can it though? Pickovers 2013 edge response comes to mind, can hundreds of pages of a proof based on axioms that require particular expertise to understand be checked instantaneously? I think it would easier to replicate the typical psychology experiment than to check the math in that case.
That’s fundamentally it, there’s a barrier to entry, and it’s knowledge. Other sciences may require specialized instruments and time, but math is also subject to the problems of knowledge not being instantaneously transferable. And I imagine, although I’m not certain how prevalent it is, that many branches of mathematics are also dependent on instrumentation these days, in the form of software.
And we all know how software can be…
People have a hard time understanding things on extreme ends of the scale because they just assume everything is within some reasonable bounds of effort, this is a failure of imagination.
Mathematics is not a reasonable amount of effort.
What's more it is depended on by literally everything else we do (as a civilization).
Theorem proving software (so far) makes the fatal mistake of just trusting; the machine it is running on, the operating system it is running in, and several other things that can subvert assumptions. The correctness of a proof is justified by the trust we place in the heavily audited kernel which will check the proof script outputted by compilers for the human (textual) interface but that kernel is just another program on your (mostly) proprietary platform.
Mathematics is a religion that currently dominates the planet (and as with most religions, it is totally ambient to people in that they don't realize to what extent they base their behaviour on it) but it is not without competition and the thing we must realize is that this competition is as motivated, ruthless and mission-driven as anything else on the planet.
Just as we have people sitting around learning to recite the qur'an or the bible from memory we also have universities full of mathematics students reciting proofs and theorems by memory, trying to keep the knowledge that brought us this far alive.
Computer memory is inherently volatile and we do not have the same conventions and procedures in place as we do for safeguarding the legitimacy of written texts. These take time to develop and are usually motivated by tragedies (such as the loss of our "prehistory").
Try to imagine the amount of effort it would take to re-assess every piece of information that we have (at current time) because we accidentally let the entropy into our libraries.
The reproducibility of computer-derived knowledge is what makes it something we can trust. Especially if it's being reproduced on disparate hardware and disparate software.
Mathematics is possibly the only thing in the entire world that can _not_ be described as a religion.
Math works, whether you believe in it or not.
You can add the Stanford Prison Experiment (Zimbardo) to the list of, now unfalsifiable, experiments that inexplicably are regularly taught in university settings when they cannot possibly provide any useful data to the sciences: https://en.wikipedia.org/wiki/Stanford_prison_experiment
The idea that many if not most universities create educational programs based on these very obvious problematic studies, leads me to believe that the scientific community can be as guilty of info-tainment bias as the general public. I only wish that I'd been able to continue in academia so that i could have at least a small voice in changing this problematic dynamic we live with.
They literally made a movie about the Zimbardo experiment in 2015. It couldn't be more obvious that people care more that it be true than whether it might be false: https://en.wikipedia.org/wiki/The_Stanford_Prison_Experiment...
"[D]espite the fact that error-correction is really hard, publishing actually false results was quite rare because “people’s intuition about what’s true is mysteriously really good.” Because we mostly only try to prove true things, our conclusions are right even when our proofs are wrong."
Deduction can't explain this surprising reliability of mathematical intuition, only induction can. The question the author tries to answer is why the intuitions of mathematicians seem to be more reliable than those of social scientists.
I think the fact that our brains are computing/pattern-recognition machines is the obvious reason computing (deducing) is 'easier'. However, human minds are not actually mysteriously really good at math. Our brains are very clearly broken at some types of mathematical problems. It's just something acedemics tell themselves. When it comes to gambling, large numbers, power laws, etc.
The number of times I've had to explain to my friends and family how exponential growth "will appear" with the spread of covid is testament to the fact that our brains don't actually have that math built in, probably because it wasn't particularly important for evolutionary survival and reproduction. The entire field of behavior economics studies this delta specifically.
I'd say the deductiveness of a problem, itself, is what makes people have the intuitions. It's vastly easier to use deduction that it is to do induction... because induction is by ultimately unknowable and effectively solipsistic.
I'm saying we are taught the results are correct, when, due to the impossibility of reproduction we can't know whether the results are correct or incorrect.
What should be discussed in the Milgram experiment is "we can't know the results of the experiment because one of the subjects killed themselves due to their participation. Try to consider the ethical implications, don't design experiments that might lead to psychological trauma, etc." The idea that the results being true is why we shouldn't do the experiment is wrong to do, is nonsense, because we can't know whether the results are true because can't replicate the experiment.
Yet, it's results are published as a pop-science book: https://en.wikipedia.org/wiki/Obedience_to_Authority:_An_Exp...
Instead, these two studies if their findings, not the ethics, are discussed at all, are a perfect illustration of the problems that Karl Popper discusses regarding falsifiability.
EDIT: Hell, pretend play is for weaklings, I'll be your confederate jumbo-dumbo test taker faker, and let's go a step beyond the original experiment, let's have me actually suffer serious shocks in our Milgram Repro, for your n >= 30 guineau pigs to see. Then we'll know even better and that can be the new standard. And I'll be your confederate convict prime, set aside in especially horrible solitary confinement conditions for the Stanford Prison Repro. And when you watch me and others actually suffer horribly for days on end due to your decisions, I'm sure you'll simply have zero suicidal ideation, and you'll repeat your "we can't know!!!11!!1" drivel.
https://www.theatlantic.com/health/archive/2015/01/rethinkin...
https://www.vox.com/2018/6/13/17449118/stanford-prison-exper...
We also should not do it again.
We can know all kinds of things with a reasonable degree of certainty even if we can't or won't verify them experimentally.
Isn't it trivial to drop something out of an airplane along with the same thing but with a parachute and measure the impact velocity?
We don't have enough data on people jumping out of planes without parachutes to conclude that the parachutes are actually working. They could just be superstition.
That this remains true is due to a) no one being allowed to repeat Milgram's experiment exactly and b) more general scientific funding bodies not allocating money for pure replications.
Is it obvious that humans will generally shock people until they are told not to? Is it obvious that randomly assigned students will take on the roles that they are given?
I don't think it's obvious. I think that is why the studies were done in the first place, and I am generally quite skeptical of the results. There is no way to verify them, thus teaching them is genuinely bad science.
It's certainly not obvious, so verification is informative, and that likely was the motivation for doing the experiments. However, despite their flaws - they were limited and biased in various ways - it certainly would be far, far worse science to base our teaching solely on one's assumptions/scepticism/opinion about how things should be, instead of taking into account whatever limited data these experiments provided. Sure, it would be better to have more and better data, but since we won't get it, this does provide relevant information.
Makes a lot of sense.
This would only be true if the hypothesis being tested in these experiments was "psychology as a field needs a code of ethics". That wasn't either hypothesis.
"We need a code of ethics" does not imply "Milgram and Stanford prison proved ___ about authority figures".
Why? The experiments would still be unethical, even if they led to the opposite result.
If the opposite result is "some people are prisoners, some are guards, and they all sit around and have a jolly old time" then what's unethical about running that experiment?
But why? The standord experiment was performative, it wasn't even an actual experiment, as shown from the experiment notes. How did you come to believe that it is mandatory to believe that study is valid?
Ironically, the "correct results" are often cited as the reason for the tragic suicide of one of the participants.
Any psychologist worth their salt knows better than to claim any interpretation of experimental results as unassailable
Pop science is another story (literally) but the profs I've had the privilege of studying under spent a lot of time warning us against the shoddy reasoning you describe
Psychology has a lot of bullshit in its orbit but the discipline as taught in quality departments is just that, disciplined
I hope you are right and I am wrong, I was taught these experiments this way at a good school, i can only hope that is more uncommon than i think it is.
After a brief look around, I also couldn’t find anything about harm to the participants.
If we could clone the participants of a psychology experiment prior to the experiment, with their memories and personalities intact, we could possibly narrow the gap. Or maybe not! But that would be a fascinating finding! ;)
I'd consider it appropriate to say that physics is indeed immune to replication crisis unless it turns out that a significant proportion (e.g. more than 5%, one order of magnitude less than for psychology) physics papers fail replication.
This is a big difference to humanities or even biology.
Psychology will always have ethic, sample size and cultural/time/space constraints. Your "psychology experiment" done with Mechanical Turk is more an exercise on sampling bias than anything else
And they keep running into "replication failures" because they fail to account for that and try to chase the illusion that it is possible to have a perfect controlled environment in psych. They can't, and they're only fooling themselves they can.
The problem with Psych is that you can't really experiment on humans in a meaningful way. Animal studies for example don't suffer the same replication problems as psych.
While absolutely valid this isn't the only problem: One other complication is that conclusions from experimental studies often don't apply to the real world, i.e. people behave differently in experimental settings. Thus in any case you can either have good control over confounding varables but very limited scope (experimental studies) or bad control over the setting but the ability to observe real live behavior (field studies).
But it gets worse: Two people with identical behaviour can have quite different motivations and thus identical observable behaviour can result in quite different consequences. I.E. 5 hours a day playing games can be both healthy behavour or an addiction/a means of suppressing thoughts or memories, depending on the person. So you'll always have that massive confounder called the mind, in every setting. Which means you have to inspect the mind itself, which only works indirectly, i.e. by interacting with the person, asking questions, trying to find out the motivation behind a particular action. And the person often doesn't really know, because a particular decision probably wasn't chosen consciously.
[0] Guess when I was experimented on? Of course it was during university and of course I'm a white english speaker that lives in North America.
The entire point of the replication crisis is that we should expect statistical anomalies to occur, and the problem is that anomalous results are wildly more likely to be published. Which is why you probably hear about evidence of life on Venus but not about calibration issues.
It’s just generally easier and more problematic to do experiments testing framework boundaries in sciences with more variables like medicine, or anything to do with behavior.
Testing a new drug, for instance, is very difficult because there are about a billion chemical processes going on in the human body (or whatever body is being tested), and trying to distill the effects of adding one more, and making sure that the effects are correctly attributable to the right causes is something we can almost certainly never know 100%.
This kind of thing certainly happens in physics, and we can see results like neutrinos appearing to travel faster than light until they are corrected, but it's usually much easier to isolate the processes of interest in a physics experiment than in a biological system, or at least easier to eliminate most extraneous effects.
However, math papers can still be wrong! Especially the actual proofs can be wrong. The question is, how many wrong proofs are out there, and how many of them have false conclusions.
If we accept that most published proofs have errors, why isn't that a horrifying failure of Mathematics that shakes fundamental trust in its systems. That is what the article is about.
This hypothetical shaking of fundamental trust would be highly analogous to the replication crisis. And there are more parallel arguments. So to me the premise makes sense, and I see little problem with some semantic wiggle room for sake of analogy.
> In experimental sciences, the experiment is the “real work” and the paper is just a description of it. But in math, the paper, itself, is the “real work”.
In other words, there is no replication crisis in math because there is no replication to be done. There is no experiment to be replicated, just work to be checked for correctness.
Total nonsense. Experiments are never falsifiable, what does that even mean? Hypotheses are falsifiable. Did any hypothesis of Milgram's suddenly become unfalsifiable one day in the last 50 years because the psychology research community found a new moral compass? Did the foundations of what is knowable by science shift during the night? "Now unfalsifiable" implies change. So what do you think changed, exactly?
You seem really hung up on a strict Popperian falsificationism that you have misunderstood.
In any case, the article mentions that you replicate a proof by reading it and checking the steps for yourself, so I feel your criticism about deductive reasoning is misplaced. It's implied that soft science papers require experiment to reproduce because they are inductive, just because the author doesn't spell this out for you doesn't mean they missed the distinction...
Yes! The moment the approval committee decides that certain experiments are no longer allowed, that makes some results no longer reachable by science, or at least not in a reasonable manner.
I agree with this point in general, but the Milgram studies have been replicated many, many times so potentially not the best example.
The Stanford Prison Study, on the other hand...
It's been an important part of modern scientific advancement that falsification is a good, but not perfect way of learning about how the universe works.
For example, macro-evolution is basically unfalsifiable, and yet it is generally regarded as good science.
This and other crises led to grounding modern mathematics with set theory, Zermelo–Fraenkel axioms, etc and understanding what's possible (e.g. Godel's theorem).
Psychology and other social sciences are barely a century old.
That replication crisis has led to efforts in formal verification such as HoTT, Lean, etc.
https://homotopytypetheory.org/
https://xenaproject.wordpress.com/2021/06/05/half-a-year-of-...
> This site serves to collect and disseminate research, resources, and tools for the investigation of homotopy type theory, and hosts a blog for those involved in its study.
> Exactly half a year ago I wrote the Liquid Tensor Experiment blog post, challenging the formalization of a difficult foundational theorem from my Analytic Geometry lecture notes on joint work with Dustin Clausen.
???
> Question: Was the proof in [Analytic] found to be correct?
> Answer: Yes, up to some usual slight imprecisions.
This has been the case for almost all math formalization efforts. Even when (very rarely) proofs were revealed to be incorrect, the result was salvageable.
When you enter truly new grounds mathematicians don't even agree if the distinctions being made have a meaning, let alone if they are true.
What examples are you thinking of? I can think of only one case where this has been the case (Mochizuki and the ABC Conjecture), but that turned out to have so much fanfare precisely because mathematicians did not agree on what distinctions were being made and this was the only time in living memory that this had occurred (the general consensus is that Mochizuki's proof is simply too obfuscated to make heads or tails of). As such the ABC Conjecture is not considered solved.
However, that is almost always not the case, even on the cutting edge of mathematics and even when making other breakthrough discoveries (e.g. Fermat's Last Theorem).
> E.g. it would be like a psychologist asking how an results of a well proven result would be different if all participants wore red shoes.
And to be clear the major reason why fields like psychology are termed to have a "replication crisis" is because well-known results are being overturned, not just cutting-edge ones.
EDIT: I see now that OP also refers to the ABC Conjecture.
Back in grad school I was very much into types which let me do things sideways to how most mathematicians did them.
The worst one was the derivative of a derivative - NOT the second order derivative. Using R for the real numbers, the type signature of first (and any) order derivative is (R -> R) -> (R -> R). The type signature for what I was talking about was ((R -> R) -> (R -> R)) -> ((R -> R) -> (R -> R)). It didn't have very many interesting properties I could find but trying to explain it to anyone else in the department was like pulling teeth. They'd start thinking about second order derivatives every time and there is no common notation for talking about third order functions like there is for first order and (some) second order functions (derivative, integral, transforms, etc).
Mathematics only works as well as it does because mathematicians are all working on pretty much the same thing, not because there is some innate quality in mathematics that pushes mathematicians towards truth.
Usually notation is defined for higher order functionals when the need arises - when a new result is found to have interesting properties. Topological proofs on function spaces can use third order functions.
> Mathematics only works as well as it does because mathematicians are all working on pretty much the same thing
I don't think any person alive can understand all the major branches of mathematics well. By that measure mathematics is very broad.
What you are talking about already exists, in many different versions at that. I think you really underestimate how varied objects mathematicians works with, they work with so many things that it is hard to come up with original ideas, this idea was already explored close to 200 years ago.
https://en.wikipedia.org/wiki/Fractional_calculus#Fractional...
I also have a PhD in maths, thanks.
I got tired of having units mismatch in physics equations so I picked up a ton of type theory and used it for everyday work. Using it, rather than talking about it, has given me a very different perspective on it than anyone else I've talked to.
Developing a practical rather than a theoretical understanding of type theory rapidly makes you express intuitions which other people can't lex/parse.
Because you begin to think in functors/compositions e.g constructively; and as you've already pointed out many of those functors don't have corresponding English nomenclature in a classical setting.
Find some Category Theorists to talk to instead.
Most mathematicians are also category theorists today. They were the ones who invented category theory in the first place and today it is a basic topic most takes in grad school and then used just about everywhere.
You could very well mean something different than what that article is talking about, but it isn't like I am just saying that a second order derivative is the same thing as what you are talking about. But if you mean something different than "the derivative of the derivative operator" then you aren't very good at explaining what you mean, you have to be more precise as math is a wide field and if you describe an object imprecisely then people will misunderstand.
> I also have a PhD in maths, thanks.
You yourself argued that people with PhD in math gets this wrong, so I don't see why you would bring this up. But the most likely scenario is that you just failed to communicate your thoughts properly since you used imprecise language. Maybe mathematicians should get taught more how to be precise with their statements, but usually they can rely on the crutch of old notation. I did invent new notation and solved some old unsolved problems that way in grad school, and the other mathematicians had no problems understanding what I wrote, so mathematicians have no problems understanding new things from my experience.
Except like, instead of the 2nd order function returning a number, where the functional derivative would be ((R -> R) -> R) -> ((R -> R) -> R), uh, instead, the thing you got.
Ok, but, differentiation is linear, so, (\frac{d}{dx}(u(x) + \varepsilon \eta(x)) - \frac{d}{dx}(u(x)) )/\varepsilon is just \frac{d}{dx}(\eta(x)) (and so there's not even much need to take the limit as \varepsilon goes to 0, as it doesn't depend on \varepsilon anyway) ..
Uh, that seems a little odd, that the derivative of differentiation in the direction of a function would be the derivative of that function? Like, the derivative of something linear should be a constant, ah, but, the directional derivative of 5 x_1 + 3 x_2 also depends on the direction in which it is taken, though it doesn't depend at all on where it is taken.
Ok, yeah, same situation here then.
Is that the sort of thing you are talking about?
Though, you said "of a derivative" not "of the derivative", so I guess maybe you mean like, more generally things that satisfy Liebniz's law?
If everyone commits to using approximately the same definitions then (at the social scale) you get normative semantics as an unintentional by-product. Everybody in your tribe uses the same terminology/notation and means the same thing by "==".
It's a desirable property because it minimises miscommunication and it enable effective asynchronous communication. Good luck trying to read a paper by somebody using different semantics.
And then computer scientists come along and point out that all definitions (and by proxy - all Mathematics) are arbitrary.
Because the chosen axioms are arbitrary.
One of the more important results discussed was not salvageable.
I think you're underestimating the amount of theorems now that have been verified without issue. Those several examples are an absolute drop in the bucket.
> One of the more important results discussed was not salvageable.
I don't see any important result in that talk that was not salvageable. I see a lemma that had to be changed in support of a larger theorem, but again nothing that "broke downstream papers" so to speak. All the results ended up being fine for their purposes.
The trials and tribulations of abstract math where these are little errors that turn out to (mostly) be salvaged translates into billion dollar mistakes in mathematical models in industry — errors that have caused a hiring freeze at a major tech company.
I respect other people feel differently — I personally think there is a crisis in mathematics, where the mistakes/errors of Voevodsky et al is the tip of the iceberg, but the most visible since it happens in academia versus industry.
I hope everyone can agree that:
a) it’s currently hard to verify mathematical models and proofs; and,
b) we could make that better — and currently are working on it.
What you lament as “business trying to use math that is not made for it” is precisely the problem:
Mathematics can’t be reliably applied, even if you hire a dozen world class PhDs, hundreds of software engineers, and let them spend years trying to make it work.
That’s a crisis.
Elizabeth Bik has put a lot of work into finding faked images in published papers. Her Twitter feed is pretty entertaining; she presents images and challenges reader to find the duplications. https://twitter.com/MicrobiomDigest
For more, see Nature: https://www.nature.com/articles/d41586-020-01363-z
- The "Italian school of algebraic geometry" of the 19th century used intuitive methods that, while groundbreaking, ultimately proved to be unreliable and generated many false results. https://en.m.wikipedia.org/wiki/Italian_school_of_algebraic_...
- Mochizuki's abc "proof", which seems to be believed a proof by many of his Japanese colleagues, but fatally flawed by most everyone else. https://www.math.columbia.edu/~woit/wordpress/?p=12220
Mathematicians have definitely gotten out ahead of their skis in the past, but I have the impression that the community today is incredibly good at finding flaws in flawed work and making solid work fully explicit and rigorous. It can take years or decades though.
> The replication crisis is partly the discovery that many major social science results do not replicate. But it’s also the discovery that we hadn’t been trying to replicate them, and we really should have been. In the social sciences we fooled ourselves into thinking our foundation was stronger than it was, by never testing it. But in math we couldn’t avoid testing it.
But the post doesn't stop there! The second part of the post (effect sizes etc), with the examples of power posing and the impossibly hungry judges, is even more illuminating. Thanks!
Some of the early results on pseudorandom generators and hash functions aren't holding up well, but I think that's just progress. We understand the problem a whole lot better than we did back then.
Perhaps more interesting is the literature on memory models. The original publications of the Java and C11 memory models had lots of flaws, which took many years to fix (and that process might not be totally done). I worry that there are a bunch of published results that are similarly flawed but just haven't gotten as much scrutiny.
there's a couple levels there:
rote translating pseudocode into your target language isn't likely to pan out well.
so instead you run the pseudocode in your mind, develop an intuition on how it works, and that's the "replication" bit this post talks about with reviewing math papers.
but both the pseudocode and your code will likely have edge cases you didn't handle. this isn't a problem for math - that's the category of common trivial/easily fixable proof errors that don't really affect the paper. but they're a problem for machines that run them literally.
maybe a good compromise strategy for formal verification is to declare the insight of the algorithm - recurrence relation or whatever - as an axiom, and then use the prover to whack the tricky edge cases.
The thing is, the general sense I get is that people in CS already have so little confidence in these results that it's not even considered worth the time to try and refute them. Which doesn't exactly speak well of the field!
Or there is, but then you're doing statistics not just ML.
But does it mean there's a crisis? maybe that's just a way to foster an environment that will let great ideas emerge.
And yet, the concepts of binary search and merge sort are fine.
I think that's quite similar to the situation in math papers? Because math isn't executable, a math paper being "significantly" wrong would be like discovering that a program uses a fatally flawed algorithm and is trying to do the impossible. It can't be fixed.
Programs that can't be fixed seem rare?
[1] https://ai.googleblog.com/2006/06/extra-extra-read-all-about...
I'd say that's an example of the difference between science (the authors don't need to show every detail of a practical implementation and can assume infinite bits in int) and engineering where you do need to make such consideration.
Since proofs are programs one can basically say that mathematical theorems are incredibly detailed software that is completely open source and invites people to identify programs that don’t work and or fix issues.
A famous one is Fermats last theorem which needed a fix but was largely right.
Others have said that it takes 6 months to a year to get published. The other thing with math is the fact that you can get completely scooped and your work is worthless.
Edit: I am using "proofs are programs" very loosely and yes Theorems are much more than programs as other commenters have pointed out.
It's a fun rabbit hole to go down :)
With programs-as-proof it really wouldn't matter. It's either "computer says yes" or "compu'er says noooo".
EDIT: Whoop, sibling post mentioned, it's Mochizuki.
He's considerably less esoteric than e.g. Grothendieck was even during his more "public" years.
It might not lend more understanding to people not invested in "field X" (or even people who are invested in field X!), but it would be proof.
Proof in the current world of math is quite intangible.
A proof is whatever convinces sufficient mathematicians that the theorem is consequent! Classical logical systems are a very good way to do that so they get used a lot. But they're not the only way, and involving a computer program makes most proofs less convincing rather than more.
Sadly not sometimes it is just:
computer says
Or "computer says x but you don't agree the mathematical concept was correctly formalized into the proof engine language".
I would say its more like pseudocode. There can be quite a large gap between a normal proof, and a machine checkable proof, which is the computer program version.
Why math specifically? One would think this applies in virtually all fields.
I don't think it's that straightforward; proofs in papers are a mix between explanation in natural language and mechanical steps. Not every step of deduction can feasibly be written out. That's part of why computer aided proofs are not that popular in math.
Like, the more accidental edge cases people produce, the less they understand the program.
While it's harder to produce the math proof, it's probably harder to grasp what's going on in a program in a mathematical sense.
(And if you have an independent paper, that can _also_ get published; your paper is distinct even if the result isn't. I think the PT HOMFLY polynomial was independently proven in like four different papers published within two years (and it's named so that all eight authors get credit).
But also, publication lags shouldn't lead to more scooping, because you can put it up on the arXiv at the beginning of the publication process, not the end. In my experience the paper is treated as "real" once it hits the arXiv; the acceptance is mostly a formality that lets us put it on our promotion packet.
But also, publication times don't lead to scooping generally because you
In mathematics and computer science, there are many errors in published papers. However, once you point out an error, it's usually pretty straightforward to resolve whether it's really an error or not. Often there is a small error which can be fixed in a small way. Exceptions like the abc conjecture are rare.
The side effect is that math papers have an insane long time to publication. Perhaps 6 months, or 1 year or more if you are unlucky.
In physics, the publication time is like 3 month. Something like 1 month for the first review and then two months for making small changes suggested by the referee and discussing with the editor.
As a side^2 effect, some citation index of the journals count only the citations during the first year. But the papers that have the citation are sleeping over the reviewer desk during that year, so the number is lower than the real number.
Wiles's proof of Fermat's Last Theorem is like 120 pages long and he first delivered it disguised as a class to a bunch of grad students who barely understood any of it and hence gave no feedback. Because this is Fermat's Last Theorem which is famous, eventually people in the math community that understood Wiles's work reviewed it and found an error. Had it been a 120 page proof of some not famous problem like random chessboard thought experiments, it probably could go years without anyone seriously looking at it.
The level of solitary study necessary to make progress is another sign of this phenomenon. Andrew Wiles famously spent six years alone in his attic to come up with his proofs for Fermat’s Last Theorem.
[0] https://sites.math.rutgers.edu/~zeilberg/Opinion104.html
However, the social science results, although unified in each country, were almost always different and in conflict.
Social sciences are politics and nothing more. It’s an opinion of how the world ought to be.
It's not like you make a hypothesis in math and then need to go away and interview a sample of 1,000 circles and report back that, controlling for ellipses that may be misreporting as circles, the ratio of the circumference to the diameter is 3.2 +/- 0.1 (p<0.05).
> But one of the distinctive things about math is that our papers aren’t just records of experiments we did elsewhere. In experimental sciences, the experiment is the “real work” and the paper is just a description of it. But in math, the paper, itself, is the “real work”.
And
> And that means that you can replicate a math paper by reading it.
I'd say most times it's "modeling", not "experimenting"
Science attempts to describe reality. Math attempts to create rules/axioms.
They're not the same pursuit, although they can often be useful together.
But it's worth considering the possibility that mathematics could have fallen, and could still fall, into a state where false results are frequently published, and isn't 'protected' by anything special in the nature of the field or its practitioners.
Just as you might find yourself asking "why did city A fall into the grip of organised crime, and not city B?". You might look for answers in the methods of police recruitment or a strong history of respect for the rule of law or anything like that, but it might turn out that the answer is really just "city A got unlucky".
But if the science agrees with a decision they've already made then they're happy to use if for justification, even if the science is junk (e.g. the crazy fines on taking children out of school in the UK).
Why is this even a question? Analytical and empirical fields face fundamentally dissimilar challenges.
I'm sure the excuse is that errors in math papers these days don't have the same kind of impact as mistakes in, let's say, medicine.
Later on my guess was that this will change when we will have a editor software that is easy to use, wysiwyg, renders math texts as well as latex, and actually the edited document is integrated with the formal proof, beneath the fancy elegant phrases and formulas.
Then I quit a job and enrolled for phd studies but that's another story.
https://philosophy.stackexchange.com/questions/14818/is-math...
Mathematics cannot have a replication crisis because there is nothing to replicate. In math correctness is tested not by redoing something but by double-checking the original work for mistakes.
The inventor of category theory's wiring diagrams, for example, has claimed that he could get middle schoolers to understand them. I suspect that success has not been replicated.
Maybe I should try. I like category theory and I am a middle school teacher and like teaching electives on improbable things :D
But because journals publish human proofs, one of the following cases can arise:
- A proof can be wrong and the underlying conjecture can still be true (i.e. a theorem).
- A proof can be wrong and the conjecture can be false - these kinds of errors need to be corrected with utmost urgency, because they lead to follow-up downstream errors.
- A proof can be right and recognized as such by peers, in which case it gets published and everybody is happy.
- A proof can also be right and be contested by peers. This is when proofs have gaps that are "obvious" to some, but not believed by others. Because a proof is a sequence of steps that are either previously derived or self-evident, and people simply differ what they consider self-evident (as an aside, in school I was criticized in maths for "skipping steps" and in English for "jumping thoughts"; to me it seemed obvious where things were going but the teachers obviously needed "comments to the code", so I learned to insert baby steps, and everyone was happy).
Because of the last case, it is important to get down to the most formal, fine-grained, nitpicky, atomic level, so everyone can agree that each micro-step is self-evident, and thankfully we can dedicate this to machines nowadays (provers like the "Isabelle" system - https://www.cl.cam.ac.uk/research/hvg/Isabelle/), at least to an extent. Ironically, that's not how real proofs in mathematics journals look like at all. They're written in prose, and often written in a surprisingly "meta" style (we had algebra lectures where the professor talked about "colored roosters" and how they behave, which is the name of some structure in the algebraic sub-field of Ramsey theory).
The author makes a point that both social science people and mathematicians are trying to prove subtle things which are actually probably true. In other words the social science replication crisis is because the experiments are often impossible to perform consistently when the effect is subtle, leading to the use of inconsistent lucky draws to demonstrate things.
Errors in science, disagreement and lack of reproducibility I think are common and prevalent but it doesn't necessarily imply that a discipline as a whole doesn't make progress. The obsession with statistical accuracy and 'science as bookkeeping' mentality seems fairly new to begin with, and science did just fine before we even had the means to verify every single thing ever published.
It kind if ignores the dynamic nature of science. Most of what is published probably has close to zero impact regardless of whether its right or wrong, but paradigm changing research generally asserts itself. Science is evolutionary in that sense, it's full of mistakes but stumbles towards correct solutions at uneven tempo. In a sense you can just look at it like VC investment. Nine times out of ten individual things don't work, but the sector overall works, the market economy is full of grifters and failed businesses, but it doesn't matter that much.
So, maybe half of math is bullshit but so is everything else but in math people just say "whatever" until they find something good, whereas in psychology people use it as an opportunity to hack away at it.
The problem is that psychology, unlike physics, doesn't really have a paradigm. There's the old saying, "Extraordinary claims require extraordinary evidence." But for that heuristic to be effective, you need a standard for what counts as extraordinary. Extraordinary compared to what? In phsyics, there are two well-established paradigms (relativity and quantum mechanics), which establish what counts as ordinary, and what counts as extraordinary. So, for example, if you're making a claim that the distribution of dark matter in the cosmos more clumpy than predicted by existing models, of that the energy level of a particular field is 12 MeV rather than 10, those are ordinary claims, which can be accomodated by tweaks to the existing paradigm. But if you're saying that the speed of light has varied over the history of the universe, or that all subatomic particles are actually tiny vibrating string-like structures, well, that's going to require a lot more evidence.
In psychology, it's much more difficult to have that kind of intuition. Take the concept of priming, for example. Is claiming that people walk more slowly when they're encouraged to think of things that make them feel old extraordinary? It makes a certain sort of intuitive sense, but, on the other hand, there's absolutely no causal mechanism suggested. So when a number of priming studies fail spectacularly under replication [1], I don't know what to think. I don't have a good sense for how much of psychology is overturned by the replication failure, in the same sense that I'd have for physics if it turned out that e.g. the speed of light is a variable rather than a constant.
[1]: https://mindhacks.com/2017/02/16/how-replicable-are-the-soci...
Actually, I would say that these well-established paradigms establish what counts as extraordinary. In other words, relativity and QM are examples of extraordinary claims that we believe because we have extraordinary evidence for them. Both of these theories say all kinds of extraordinary things, and most people who first encounter the theories start out thinking they can't possibly be true. We believe them not because they are just ordinary, but because we have taken the time and effort to accumulate extraordinary evidence for them.
In that light, the replication crisis in other areas of science is easily explained: they allow extraordinary claims to be published without the extraordinary evidence that those claims would require. So of course many of those claims turn out to be wrong.
I still don't know the answer but I suspect there is something to the simple/complicated distinction.
I do recall that there was this one professor that couldn't replicate his seminal work that basically got him tenure and it was all very dramatic. the kind of story that we would share with each other as a byword.
Aslo, I wonder if the errors get caught more quickly because in chemistry and physics to a degree, most progress is built directly on previous results. so you can't have bad results just floating unnoticed for very long.
The issues I have with the "replication crisis" are two: malicious publishing of false results for personal or organisational profit, and malicious interpretation of prestigious publications as the one and only immutable truth. Both of which are social and political problems.
Scientific publishing was functioning fine so far for the same reason the old internet was full of true information while it didn't have to: there was little to be gained by lying. Now that gov policies are build on phd papers, things are changing.
Apart from the replication crisis, the other crisis that is not really talked about is funding. Academic funding basically comes from one of 3 (connected) sources - government, corporations and the military. Somehow or other - these 3 sources have pretty much the same or non-conflicting aims. These aims relate to power and control. This is actually the largest crisis, IMO.
Given that information, we can re-assess the quoted statement the author makes.
Perhaps its not that they are proving things that are "basically true". Its that right or wrong do not matter. What does matter is that the answers provided meet the agenda of those funding the study. The answer is not that important as long as it is supportive of whatever agenda is in play. I believe this is the case for the replication crisis in science also.
A replication "crisis" is only a crisis if you are attempting to achieve truth and greater understanding. But truth and understanding are only ostensible reasons, not the actual ones. What these studies are actually doing is creating a parallel construction - the aim is actually for studies to appear 'truthey', without actually being so. What studies should actual do is increase the funder's power, wealth extraction abilities, etc.
If you doubt this and think that truth matters, consider this. Surely we should have cracked the best diet for people by now? But there is no common understanding of what is good or bad to eat - if anything there is more confusion. The reason of course, is that there's no money in recommending whole foods or whatever. However, there is money in drugs to make people 'better'. And money in making diet so confusing that people eat themselves into trouble.
Anyway, if you are in the business of governing or monetising the masses, truth and understanding is the last thing you want. Far better to have a story that gives you control, or extracts money. Such is life under fascist governance (where fascist = corporation + governance working together).
It comes from trying to make inference from noisy and often confounded experiments from limited data.
> Many papers have errors, yes—but our major results generally hold up, even when the intermediate steps are wrong!
is it possible that e.g.psychology research might become more mathematically rigorous if there were more incentives for serious mathematics people to study psychology?
I appreciate at first glance that might seem a bit vacuous: but there’s a ton of money and way better policy at stake if we understood behavioral economics better, and that’s at least an adjacent field to psychology right?
Verification and review is harder in experimental works and it might take few decades before someone finds an error, or even collect resources to verify the claim. So maybe it's harder to fight 'publish or perish' deamons in experimental sciences?
4.10 The “Lacy Macbeth Effect”
Did you know? That in Science, and therefore reality as we know it... NOTHING can be proven to be true. Proof is the domain of logic and mathematics, NOT science.
There is a huge difference between science and mathematics.
Math is an imaginary game involving logic. You first recursively assume logic is true then you assume some additional things are true called "axioms." You assume the entire universe only encapsulates Logic and the axioms as part of your theory and that is everything that exists in your theoretical universe. You derive or prove statements that you know must consequently be true in your made up universe based off of logic and the "axioms." These statements are called "theorems." That is math...
For example the "pythagorean theorem" is a statement made about geometry assuming that the axioms of geometry are true and logic is true.
So of course if it's a logical game, there's no real replication crisis. All theorems in mathematics are proven true based off of logic. You can't replicate conflicting results if a result is PROVEN to be true via pure logic.
Science is different. Nothing can ever be proven to be true. There are basically two axioms in science. We assume logic is true, just like in math. And we assume that the mathematical theory of probability is a true model for random events happening in the real world. Then we assume everything else about the universe is completely unknown. This is key.
Because the domain of the universe is unknown... at any point in time we can observe something that contradicts our initial hypothesis. ANY point in time. That means even the theory of gravity is forever open to disproof. This is the reason why NOTHING can be proven. To quote Einstein:
"No amount of experimentation can ever prove me right; a single experiment can prove me wrong."
So sure mathematicians can make mistakes... and I know the author is talking about higher level details.... but at its most fundamental level, assuming all other attributes are ideal... it is fundamentally impossible for math to have a replication crisis while science, on the other hand, is fundamentally forever open to these crisis so long as the universe is unknown.
The most interesting thing to me in all of this however is that within science, probability is a random axiom. We have no idea why probability works... it just does. It's the driver behind another fundamental concept: Entropy and the arrow of time. For some strange reason Logic in our universe exists side by side with probability as a fundamental property.
You can publish a paper in mathematics that claims to prove something, but is mistaken. A paper claiming that a theorem is proven, is not the same thing as the theorem being proven. However, that's not often the case—why that is, is an interesting & meaningful question.
It took a very long time to even notice the axiom of choice, and longer to ultimately prove its proof futile.
Please read about abc conjecture and whole saga wrt proof, Shinichi Mochizuki, Peter Scholze, etc, etc, etc
Here there is piece of the very important math research which community failed to replicated in either positive way (confirm it) or negative way (reject it)
But it does! Look, in social sciences there are a lot of lousy papers written by people who doesn't know statistics, acceptance criteria etc. There are problems with "average" papers.
In math there are no problems with "average" papers, afaik. People are following those papers and rather quickly find flows or confirm the result.
In math there are problems with "exceptional" papers. Holes in Fermat theorem proof or abc paper were taken years to find, patch and settle.
The "Replication Crisis" is about confirming long-standing results that are widely accepted, and yet when re-examined for the purpose of replication cannot be confirmed.
Mat his different from social studies, but (and here I disagree with OP) it has it's own share of replication problem(s)
As the article pretty much says.
The article also suggests an "exceptionalism" of maths in that their results hold up even when their methods don't. In other words math papers are full of "unimportant flaws and errors" but even though that's the case when corrected or verified, somehow the conclusions of the papers mostly hold up, in the author's opinion.
But I tend to think this is more because, math is groups of people (and formal verification systems) convincing themselves they know what someone else is talking about, so confirmation bias is probably in effect.
The limitations of a self-consistent set of axioms after all...
You may say that formal verification systems are perfect oracles... But if papers have "unimportant flaws and errors" it's possible that those systems have them, too, right? It's also possible that as those systems were written by people, they reflect and somehow seek to confirm the biases of those people or their understandings (or misunderstandings) about mathematical knowledge.
So I don't think maths is quite as exceptional as this author hopes it might. Tho Maybe it is