Please Commit More Blatant Academic Fraud (2021)
jacobbuckman.com
jacobbuckman.com
A few days later I got an email from the author (some professor) who wanted to discuss this with me, claiming that the paper was written by some of his students who were not credited as authors. They were unexperienced, made a mistake, yaddah yaddah yaddah. I forwarded the mail to the editors and never heard from this case again. I don't expect that anything happened, the corrective actions for a level-1 violation are pretty harsh and would have been hard to miss.
The fact that this person was able to obtain my name and contact info shattered any trust I had in the "blind" part of the double-blind review process.
The other two reviewers had recommended to accept the paper without revisions, by the way.
The organisers then made the argument that double blind was working because 50% of papers were not identified correctly! I was amazed that even with strong evidence that double blind was not working, the organisers were still able to convince themselves to continue with business as usual.
That experiment showed that even when asked to put effort into identifying the source of an anonymized paper—something that most reviewers probably don't put any conscious effort into normally—the anonymization was having a substantial effect compared to not anonymizing the papers.
Am I missing some obvious reason why double-blind reviews should only be attempted if the blinding can be achieved with a near-perfect success rate, or are you just setting the bar unreasonably high?
> Am I missing some obvious reason why double-blind reviews should only be attempted if the blinding can be achieved with a near-perfect success rate, or are you just setting the bar unreasonably high?
OP thinks you are looking at either signal or noise, instead of determining where the signal begins for yourself.
No, really: we have the same problem in software. Software developers under high pressure to move tickets will often resort to the minor fraud of converting unfinished features into bugs by marking them complete when they are not in fact complete. This is very similar to the minor fraud of an academic publishing an overstated / incorrect result to stay competitive with others doing the same. Often it's more efficient in both cases to just ignore the problem, which will generally self-correct with time. If not, we have to think about intervention -- but in software this story has played out a thousand times in a thousand organizations, so we know what intervention looks like.
Acceptance testing. That's the solution. Nobody likes it. Companies don't like to pay for the extra workers and developers don't like the added bureaucracy. But it works. Maybe it's time for some fraction of grant money to go to replication, and for replication to play a bigger role in gating the prestige indicators.
I completely disagree.
For one, academic standards of publishing are not at all the same as the standards for in-house software development. In academia, a published result is typically regarded as a finished product, even if the result is not exhaustive. You cannot push a fix to the paper later; an entirely new paper has to be written and accepted. And this is for good reason: the paper represents a time-stamp of progress in the field that others can build off of. In the sciences, projects can range from 6 months to years, so a literature polluted with half-baked results is a big impediment to planning and resource allocation.
A better comparison for academic publishing would be a major collaborative open source project like the Linux kernel. Any change has to be thoroughly justified and vetted before it is merged because mistakes cause other people problems and wasted time/effort. Do whatever you like with your own hobbyist project, but if you plan for it to be adopted and integrated into the wider software ecosystem, your code quality needs to be higher and you need to have your interfaces speced out. That's the analogy for academic publishing.
The problems in modern academic publishing are almost entirely caused by the perverse incentives of measuring academic status by publication record (number of publications and impact factor). Lowering publishing standards so academics can play this game better is solving the wrong problem. Standards should be even higher.
The alternative to not enforcing existing rules against plagiarism is to enforce them.
The alternative to ignoring integrity issues i.e."minor fraud" in the workplace is to apply ordinary workplace discipline on them.
Well... spending a few weeks reproducing a shiny conference paper that simply doesn't work and is easily beaten by any classical baseline will do that to you in the first few months of your PhD imo. I've become so skeptic over the years that I assume almost all papers to be lies until proven otherwise.
"This surfaces the fundamental tension between good science and career progression buried deep at the heart of academia. Most researchers are to some extent “career researchers”, motivated by the power and prestige that rewards those who excel in the academic system, rather than idealistic pursuit of scientific truth."
For the first years of my PhD I simply refused to parttake in the subtle kinds of frauud listed in the second paragraph of the post. As a result, I barely had any publications worth mentioning. Mostly papers shared with others, where I couldn't stop the paper from happening by the time I realized that there is too little substance for me to be comfortable with it.
As a result, my publication history looks sad and my carreer looks nothing like I wished it would.
Now, a few years later, I've become much better at research and can now get my papers to the point where I'm comfortable submitting them with a straight face. I've also came to terms with overselling something that does have substance, just not as much as I wish it had.
I couldn't agree more. I have read a lot of psychology papers during my PhD and I think there is very little signal in the papers. Many empirical papers for example use basically the same "gold standard" analysis, which is fundamentally flawed in many ways. One problem for example is that if you would use another statistical model, then the conclusions would often be wildly different. Another being that the signal is often so weak that you can't use it to predict much (to be useful). If you try to select individuals for example, the only thing you can tell is that the group on average is less neurotic. But for individuals there is no better chance of picking the right one than average. The point of a good paper is to take these sketchy analyses and write a beautiful story around it with convincing speculation. It sounds absurd but take a random quantitative psychology paper and check which percentage of the claims made in the discussion are actually based on the actual data from the paper.
But the worst part about this is that these problems exist for literally decades. Nobody cares. The funding agencies grade people not on correctness but on the number of citations. As a result, you see that many subcultures exist who's sole existence is about promoting the importance of their subculture. It is quite common in academia to cite someone in the introduction just to "prove" that some idea is worth pursuing. But does it work? Doesn't matter. Just keep writing papers.
So I'm not saying that all research is bad. I'm saying that indeed most papers are not very useful or correct. Many researchers try, but the incentives are extremely crooked.
How could the Federal government ensure that public monies only fund high quality research? Could policy re-shape the incentives and unlock a healthy scientific sector?
Ignorance as a defense needs to go too. Ignorance as a defense is too powerful and we should balance it more towards hurting the supposedly ignorant rather than everyone else. Basically, a redefining of wilful ignorance so it's balanced as stated.
I feel ashamed that my name is on it. I wish I could retract it.
So, yes please: make it hard to impossible for paper mills and kill the whole publish or perish approach.
I contributed nothing other than a statistical framework which was discarded when it broke their predefined conclusion.
I think as children if we are taught what earning a living means, people who only want to make ends meet would try to do it using other less damaging methods. For e.g., sales and marketing are not bad places for such people. When it comes to research people should know that perhaps money will not be great.
It is because we aren't aware of the full picture as children, we follow our passions (or we follow cool passions) and then realise that money is also important and then resort to unethical means to get that money. Let's be transparent about hard fields with children so that when they enter such fields they know what they are getting into.
In the past this wasn't an issue because university was seen as optional, now in most places it's ostensibly required to obtain a sufficiently well paying job, and so much more of society ends up on a treadmill that they may not really want to hop off of.
I think in some fields you walk into them with some kind of noble ideology, possibly driven by marketing but then you find out it's all bullshit and you're n-years into your educational investment then. Your options are to shrug and join in or write everything off and walk away.
I don't blame people for taking advantage of it but in some areas, particularly health related, there are consequences to society past financial concerns.
I don't think this would help. IMO, it's a money vs. effort thing. Yes, real research is hard, but if someone learns early on that the system can be easily gamed, then the required effort is relatively low.
Plus, there's the friction factor. Moving from undergrad to grad to post-grad to professor keeps you within an institution you know.
The game is this: get hired at a research university and pump out phony papers which look legit enough to not raise any suspicions until you get tenure. Wrap the phoniness of each paper in a shroud of plausible deniability. If anything comes out after you're tenured, then just deny and/or deflect any wrongdoing.
A few computer science friends of mine worked at a social science department during university. Their tasks included maintaining the computers, but also support the researchers with experiment design (if computers were involved) and statistical analysis. They got into trouble because they didn't want to use unsound or incorrect methods.
The general train of thought was not "does the data confirm my hypothesis?" but "how can I make my data confirm my hypothesis?" instead. Often experiments were biased to achieve the desired results.
As a result, these scientific misconduct was business as usual and the guys eventually quit.
Research fraud is common pretty much everywhere in academia, especially where there's money, i.e. adjacent to industry.
Observations: Firstly inventing a conclusion is a big problem. I'm not even talking about a hypothesis that needs to be tested but a conclusion. A vague ambiguous hypothesis which was likely true was invented to support the conclusion and the relationship inverted. Then data was selected and fitted until there was a level of confidence where it was worth publishing it. Secondly they were using very subjective data collection methods by extremely biased people then mangling and interpolating it to make it look like there was more observation data than there was. Thirdly when you do some honest research and not publish because it looks bad saying that the entire field is compromised for the conference coming up which everyone is really looking forward to and has booked flights and hotels already.
If you want to read some of the hellish bullshit, look up critique of the Q methodology.
At least in the social sciences there is an expectation of having some data!
There's huge amounts of data available (geography, lots and lots of maps; history, huge amount of historical documentation; economics, vast amounts of public datasets produced every month by most governments; political science, censuses, voting records, driver registrations, political contest results all over the Earth - often for decades if not centuries).
Most is relatively well verified, and often tells you how it was verified [2]. Often it's obtainable in publicly available datasets that numerous other researchers can verify was obtained from a legitimate source. [3][4][5][6][7][8][9][10][11][12]
There's lots of data available. Much is also verifiable in a very personal way simply by walking somewhere and looking. In many ways, social sciences should be one of the most rigorous disciplines in most of academia.
[1] Using Wikipedia's grouping on "social sciences" (anthropology, archaeology, economics, geography, history, linguistics, management, communication studies, psychology, culturology and political science): https://en.wikipedia.org/wiki/Social_science
[2] Census 2020, Data Quality: https://www.census.gov/programs-surveys/decennial-census/dec...
[3] Economic Indicators by Country: https://tradingeconomics.com/indicators
[4] Our World in Data (with Demographics, Health, Poverty, Education, Innovation, Community Wellbeing, Democracy): https://ourworldindata.org/
[5] Observatory of Economic Complexity: https://oec.world/en
[6] iNaturalist (at least from a biological history perspective): https://www.inaturalist.org/taxa/43577-Pan-troglodytes
[7] Coalition for Archaeological Synthesis, Data Sources: https://www.archsynth.org/resources/data-sources/
[8] Language Goldmine (linguistics datasets): http://languagegoldmine.com/
[9] Pew Research (regular surveys on economics, political science, religion, communication, psychology - usually 10,000 respondents United States, 1000 respondents international): https://www.pewresearch.org/
[10] Marinetraffic (worldwide cargo shipping): https://www.marinetraffic.com/en/ais/home/centerx:-12.0/cent...
[11] Flightradar Aviation Data (people movement): https://www.flightradar24.com/data
[12] Windy Worldwide Web Cameras: https://www.windy.com/?42.892,-104.326,5,p:cams
The complaint is that their data often doesn't strongly support the hypothesis, and dubious statistical techniques are performed to make it appear otherwise. And just poor statistics abilities (not malicious intent).
Physicists get away with it because they often just don't do any statistics. Often the data aligns so well with the hypothesis that you don't need any sophisticated techniques, or their work doesn't involve any data (like my example in my prior comment).
Most US trained physicists have never taken a course in statistics. It's not in the curriculum in most universities. When I was in school and would point it out, the response was always "Why do we need a whole course in statistics? We learn it in quantum mechanics."
No. That's probability you learn. Not statistics.
In social sciences (and medicine) people take a lot more statistics courses because the systems are much more complex than typical physics systems. A lot more confounding variables, etc. They simply need more statistics.
(Yes, yes. I know. There's probably some experimental branch in physics where people actually do use statistics. But most don't).
I wouldn't say I hate social science, that's much too strong. The rampant fraud and poor method in several of the fields just means that I put less value in peoples' academic achievements than they deserve - which I don't like, because many surely sincerely tried to do good science and spent years on getting there, but I can't filter a priori in which camp a person belongs. They should not be defunded or stuff like that, but they need to get their act together. Somehow. I suspect a lot of this is driven by extrinsics (publish or perish; need an advanced degree to get a job, but the advanced degree is actually pointless for the job; probably more things I don't think of now), and those need to change to allow for good science.
Take for example the department I mentioned above, that's essentially commiting fraud. Word is, the professor running it is actually pretty damn good at what they do. They have an accepted grant application framed on the wall: "I need 2000 bucks. Signed Professor Foobar" (like 5000$ in today's money); times surely changed for the worse for them. And I pity that, since we're often (but not always of course) talking peanuts in many of those fields. Especially for Masters level research, or for a single paper.
But I judge people in my life by their competence and character anyway, not by their degree. So politely ignoring their degree has little to no adverse effect on how I interact with them.
I’ll reduce it to a part of psychology.
The proof of concept worked, but it wasn't doing anything new. We were just doing what we used to do, but now this terrible component was involved in it, making everything slower and more complicated.
Somehow that became a paper, and somehow this paper passed review without a single comment (my feeling is it's because of the professor's name recognition). I'm ashamed to have my name on that paper.
In practice people see that $SYSTEM is rotten and most likely to doom everyone on the long span, with increasingly absurd actions accepted silently on the road. But they also have the firm conviction that not bending the knee, be brave and say out loud what’s in everyone mind, will only put them on the fast track to play the scapegoat and change nothing else on the overall.
Think about it: over-reporting of grain production was a major factor of the great Chinese Famine.
The cover ups in the article were also interesting- a deliberate staging to Mao to prevent uncovering the truth. I'm not sure how this compares directly (is there a centralized authority with power to fix the issue that is being lied to, compared to the decentralized "rotten" system, where the status quo is understood and 'accepted').
Science produces discrete units which can be used in different ways, if not in their exact form from a preceding research. I am not sure it’s reasonable to say that existing ideas, even if not cited, are not inspirational (to the researchers themselves). Peer-review isn’t perfect, but I think that all accepted papers have something academically or scientifically relevant, even if there’s no guarantee that the paper will generate hundreds of subsequent citations. I think improving your subsequent work is more important, which includes mentioning why you think some previous work may not be as relevant anymore. This last step is often missing from many research papers.
I think the author is right that it doesn’t quite make sense to publish anything you know isn’t quite correct. But I can think of several papers in different fields which someone may think are “not quite correct”, but the goal of such papers, I think, is to demonstrate the power of low probability scenarios, or edge cases. Edge cases are important because they break expected behavior, and are often the root cause of system fragility, system evolution, or poor generalization in other systems.
Some differences:
- The first one was in a space with more low hanging fruit
- The first one was after large effect sizes, not the kind where you can massage the statistical model
- The second one was a topic with far higher public interest
- The second one was primarily an analytic project, whereas the first one was primarily experimental
I feel like bad science lives in the middle of a spectrum - you have young fields/subfields with boring but impressive experimental breakthroughs, and on the other end you have highly political questions that have been argued to death without resolution. Bad science is about borrowing some of the strategies used in politics because all the important experiments have already been done.
One of the pernicious things in this area is that, even as we teach young researchers how to avoid making mistakes and engage sceptically with the work of others and that scientific fraud is a nontrivial issue, we also tell them how to commit fraud themselves and that their competition is doing it.
"Watch out for P-hacking, that's where the researcher uses a form of analysis that has a small chance of a false positive, and analyses loads of subsets of your dataset until a false positive arises and just publishes that one"
"Watch out for over-fitting to benchmarks, like a car taking the speed crown by sacrificing the ability to corner"
"Watch out for incomplete descriptions of test setups, like testing on a 'continent-scale map' but not mentioning how detailed a map it was"
"Watch out for citations where the cited paper doesn't say what is claimed, some people will copy-and-paste citations without reading the source paper"
"Watch out for papers using complicated notation, fancy equations and jargon to make you feel this looks like a 'proper' paper"
"Watch out for deceptive choice of accurate numbers, like a study with a 25% completion rate including the drop-outs in the number of participants"
"Watch out for simulations with inaccurate noise models, if the noise is gaussian in the simulation but a random walk in reality, great simulated results won't transfer to reality"
I've made no suggestion at all that you should modify your science or commit fraud - but I've also just trained you in how to do it.
I'm just saying: I don't believe anyone actually tells budding researchers that they should commit fraud. Instead I think the process is probably more like this:
Year 1: Statistics/research training. Here are a load of subtle mistakes to watch out for and avoid. Scientific fraud happens sometimes. Don't do it, it's very dishonest.
Year 2: Starting research. Gee a lot of these papers I'm reading are hard to reproduce, or unclear. Maybe fraud is widespread - or maybe they're just smarter or better equipped than me.
Year 3: "You really ought to have published some papers by now, the average student in your position has 3 papers. If you don't want to flunk out you really need to start showing some progress"
It's not that you need to be tought how to cheat, it's that you need to be tought how to avoid unintentionally cheating.
New student is shown how to read a paper, how to spot egregious errors and all the things listed above.
Student, i guess feels forced to publish. And maybe uses murkier tactics to get the paper published.
As Dr Frank Etscorn said, "I can show anything correlates to anything else." We were discussing vitamin D papers anf I was testing that paper funding AI mentioned on HN a few times last year.
I don't see a problem with this? If papers are the vehicle for conference entries why shouldn't authors submit it just because it's wrong? Conferences are for discussion. So go there and discuss it... "My paper says XYZ, but since I wrote it I realised ABC" - sounds like a good talk to me?
(Naivety check: I am not an academic)
(Experience check: I is one)
What you're describing are workshops with what we would call non-archival proceedings. Places where you write whatever you want and then talk about it.
Publications, conference or journal, are supposed to be what are called archival. They are a record of what we've discovered and want to share with the world. They are supposed to be sent into the world after we carefully complete a line of work.
Publications are not supposed to spam the system with half-baked junk. Sadly, that's what a lot of people are doing these days.
Now, my pet theory is that they knew glyphosate wasn't that great, but talked it up in papers as a sacrificial anode sort of thing "gee shucks it looks like glyphosate based pesticides are harmful to humans (or bees, or fish, or) so we'll stop manufacturing that formulation."
But, Possibly due to academia, they have fanboys and cheerleaders and I think that's why it's still around and in heavy use even though we're not sure it's a good idea.
[0] Bayer Monsanto funds studies at agricultural universities.
P. S. Just watch.
This is the funny part. There is little to no power and prestige to be had in the academic system. To a first approximation no one outside academia cares.
I was just working as a staff programmer and taking grad courses with my tuition benefit, and found myself getting caught up in the mentality of needing a PhD to really be successful and valuable. Then got a job in industry making far more money and realized how academia is a small self contained world with status hierarchies irrelevant outside that small world.
They have power over their students and relative power over other Professors. That's plenty enough incentive for most. There can also be fame and fortune for the most famous among them. See Francesa Gino, Dan Ariely, etc.
Even if they don't work for the administration, there are plenty of other bodies that will value them and pay large sums of money (or let them have large influence).
Very common amongst economists, and more and more common amongst disciplines like psychology.
Even in technical fields, if you can manage to become a big name, you can do consulting work and get paid quite well.
> Then got a job in industry making far more money
This is not a healthy way to look at it.
The average mechanical engineer isn't making tons of money in industry. A ME professor at a top university likely makes more. A biology major with just a BS degree will make less than the average biology associate professor.
But more importantly, there's a simpler reason why money is a poor metric to measure: You can always get more money in finance or medicine than as a mechanical engineer. Does it make sense to denigrate a whole profession just because one can make more money elsewhere?
> * The colluders share, amongst themselves, the titles of each other's papers, violating the tenet of blind reviewing and creating a significant undisclosed conflict of interest.
> * The colluders hide conflicts of interest, then bid to review these papers, sometimes from duplicate accounts, in an attempt to be assigned to these papers as reviewers.
Is it that common that conference reviewers also submit papers to the conference? Wouldn't that alone already be a conflict of interest? (After all, you then have an interest in dismissing as many papers as possible to increase the likelihood of your own paper being accepted). And how do you create "duplicate accounts"? The conferences I have submitted to, and reviewed for, all had an invitation-like process for potential reviewers.
And finding reviewers who know their stuff, who'll work for free, and who'll review thoroughly in a short timescale isn't easy.
Come to think of it, is there a "Journal of Academic Fraud"?
Almost all the benchmarking results I see is just a percentage difference between two algebraic means, no statistical analysis whatsoever.
Very common interaction: QA folks say "your change degraded some of our metrics and improved some others". I know they are full of shit because it's impossible that my change improved any perf metrics. I ask for statistical details, they don't have any, this meeting was a waste of time, it will be next time too.
The fact that I get these reactions suggests that everyone else just lets each other get away with it.
https://github.com/denoland/pm-benchmark
Check the run bench shell script (there's not much else in the repo anyways)
I do this sort of thing to see what tools are faster all the time. ripgrep, ag(silver searcher), grep, MongoDB was one we were arguing about for a while recently.
(I'm the author of ripgrep.)
I have a lot of subtitles. I'm partially hard of hearing and partially i can't stand the way everything is mastered, so i use volume normalization (sometimes called "night mode", vizio calls it this) and subtitles to make up for the fact that the audio tracks in most things is bad.
Well a side effect of subtitles is now i have context for every video that i can search. grep was grep.
you didn't think i'd leave you hanging https://i.imgur.com/Vs5AAT7.png
some other non-statistics from that day: 15GB sorted password list, newline delimited, UTF-8 from spinningrust drive 64 seconds (~234MB/s) to make a copy of the file. ag and rg took 3.2 seconds to search the copy. I'm actually hesitant to state that grep took 52 seconds...
Thanks for replying, thanks for making me remember the great conversations we had around those topics a couple months ago, and thanks for creating ripgrep, it's my go-to for anything non-trivial!
I've occasionally wanted to put the subtitles from all of my Simpson episodes into an easily searchable format. What do you use to extract subtitles?
also when i get stuff from a website with yt-dlp for archival i use
```pwsh
$userInput = Read-Host -Prompt '480 video download script enter URL'
Write-Output "URL:`t`t$userInput"
yt-dlp.exe `
-f 'bestvideo[height<=480]+bestaudio/best[height<=480]' `
--write-auto-subs --write-subs `
--fragment-retries infinite `
$userInput
```
I burn in the subtitles because "streaming media players" nearly universally are awful at handing anything except 100% perfect subtitles - and that's if they bother handling them at all, over the last 20 years. Also my best friend was dating a deaf person, so the impetus for burning in for streaming was because we'd watch movies together in my living room on a rear projection TV via wifi streaming from a WHS "plex-like" server in my room. The device was a western digital something TV.
I also own the first 20 seasons of the Simpsons on DVD. I'd own more but they stopped making them. Despite starting it decades ago, they took so long that the DVD (and disk format in general) has become obsolescent for the most part. But I've ripped all 20 seasons on to my network, including the commentary tracks, which I absolutely love.
I don't burn subtitles in though. I've never had a problem using them with vlc (since the early aughts) and then mpv (as of maybe 10 years ago). I don't use anything fancy though for viewing. I just have a handful of Intel NUCs running Archlinux hooked up to each TV in the house. I skipped all the Plex bullshit.
That being said, the insane emphasis on venue is what's pushing me out of academia. I can't compete with people like this.
For me it sounds counter-productive. I have a feeling that lately (tens of years) many people try to focus on the negatives, rather than the positives. Should we focus on the 3 amazing papers this year, cited by hundreds, that resulted in clear progress or should we complain that 100 papers are useless? Let's focus on 100 because "someone is wrong on the internet".
I did a PhD (so might have more experience) but papers are meant for dissemination. For me everybody that wants to have them "perfect"/"useful" papers imagines a system that does lots of work for them. The system could be improved, but if anything (just throwing an idea) maybe researchers should try to do research in the industry to prove themselves. Then come back after 10 year in academia (maybe with savings) so that they are more independent of "career progression". A lot of research was done (historically, >100 years ago) by rich people, not constrained by a career.
Agree, and it seems that this is how fields naturally evolve anyway.
Together, we can force the community to reckon with its own shortcomings, and develop stronger, better, and more scientific norms. It is a harsh treatment, to be sure – a chemotherapy regimen that risks destroying us entirely. But this is our best shot at destroying the cancer that has infected our community.
And now, nearly 4 years later, research funding gets DOGE'd...
In this case, how could the federal government ensure that academic funding flows to researchers and institutions that are doing genuinely high quality research? How can we remove funding from low-quality and fraudulent actors?
Trendiness trumps all notions of academic rigor, and as long as a field "feels like" it's on the cutting edge it can go pretty far before collapsing in on itself.
So much so that there's a wikipedia page for how bad it is: https://en.m.wikipedia.org/wiki/Criticism_of_evolutionary_ps...
Secondly, there are many ingredients required to successfully publish, communicate science, foster collaboration, etc., beyond technical brilliance. I'm sure we all know many technically brilliant people whose career never advanced because they lacked in some necessary area. People shouldn't be discouraged from improving in all areas because OP's delicate genius is offended by their technical ability.
Speaking of discouragement, it's a shame and a disgrace that you publicly called your colleague's work bullshit, including a first author that isn't yourself.
This might be true in hard sciences where a "head of steam" can only build based on real, replicatable results
But it's very common that public policy is proposed and adopted based on findings from soft sciences like psychology and sociology
If policy is adopted based on a research paper, I would count that as a "head of steam" being built.
And if that paper is fraudulent, then we are adopting well-intentioned policy on false pretenses
Conversely, computer science/AI doesn't have an equivalent of the rigor that public policy research tends to go through. CS has e.g., benchmark datasets, typical evaluation metrics, but these are more like norms rather than requirements, whereas in public policy, instruments for validations are far more rigorously tested and enforced. Depending on the area.
I agree that outright fraud would be detrimental, but I think OP overblows this issue completely and should apologise to his co-authors.
I think in any field it's natural to start out naive with an idea that may just be a few steps away from solving something important. Somewhere in the middle only to realize you're not, and then scramble with the ethical dilemma around your work, is it "good enough" or not. I was there anyway.
Horseshit. This might be true for AI research (and even there that's an awfully broad brush you're using, mate), but it's certainly not true for other areas of computer science.
As a society, we have far too much trust in science, however any time this argument is brought up, we focus on conspiracy theorists who struggle with 100+ years old theories as if the visage of the public trust in science will change their mind, ignoring that any member of the public, who accidentally discovers the Jenga tower the science is built on (but hidden) will become much more likely to believe those charlatans in the future.
As a society, there is laughably little support for science, instead the majority of policy and business decisions are based on fairy tales and snake oil. We need more trust in science.
So you don't know anyone, but I can counter with those three anecdotes.
3>1
Qed I am right, statistically.
Is there even more stuff which really shouldn't be published, and has experiments which are abused to show off how great new technique A is, while hiding that was attempt 72 at making an experiment that showed A was great? Also of course.
Fraud here basically means faking reputation. There are many ways to do this. And it's common because doing scientific work goes hand in hand with very generous funding. And money corrupts things. So attempts to fake reputation, plagiarize work, artificially boost relevance through low reputable referencing, bribes, etc. are as old as scientific publishing is.
There are a few interesting dynamics here that counter this: high quality publications will want to defend their reputation. E.g. Nature retracting an article tends to be scandalous. They do it to preserve their reputation. And it tends to be bad for the reputation of affected authors. Their reputation is based on them having very high standards and a long history of important people publishing important things that changed the world. Every time they publish something, that's the reputation that is at stake. So, they are strict. And they should be.
The problem is all the second and third rate researchers that make up most of the scientific community. We don't all get to have Einstein level reputations. And things are fiercely competitive at the bottom. And if you have no reputation, sacrificing it is a small price to pay. Also the prestigious publications are guarded by an elitist, highly political, in-crowd. We're talking big money here. And money corrupts. So, this works both ways.
With AI thrown in the mix, the academic world has to up its game. And the tools it is going to have to use are essentially the same used in other social networks. Bluesky, Twitter, etc. have exactly the same problem as scientific publishers; but at a much larger scale. They want to surface reputable stuff and filter out all the AI generated misinformation and it's an arms race.
One solution is using more AI or trying other clever tricks. A simpler solution is tying reputations to digital signatures. Scientific work is not anonymous. You literally stake your reputation by tying your (good) name to a publication and going on the record by "publishing" something. Digital signatures add some strength to that that AIs can't fake or forge. Either you said it and signed it; or you didn't. Easy to verify. And either you are reputable, by having your signature associated with a lot of reputable publications, or you are not. Also easy to verify.
If disreputable stuff gets flagged, you simply scrutinize all the authors and publications involved and let them sort out their reputations by taking appropriate actions (firing people, withdrawing articles, publicly apologizing, etc.). They'll all be eager to restore their reputations so that should be uncontroversial. Or they don't and lose their reputation.
Digital signatures are a severely underused tool currently. We've had access to those for half a century or so.
The challenge isn't technical but institutional. Lots of disreputable people and institutions are currently making a lot of money by operating in the shadows. The tools are there to fix this. But people don't seem to necessarily want to.