Tortured phrases: A dubious writing style emerging in science
nature.com
nature.com
(We merged this thread and https://news.ycombinator.com/item?id=28108111)
That said, the authors did not fabricate their research (as far as I can tell). They just did not know English well, so it was easier to just copy things that you know are phrased well than to learn to write English well. As the saying goes, do not attribute to malice what can be explained by ignorance or laziness. That does not excuse it but it makes it more understandable.
I agree with the article that this is probably just the tip of the iceberg. There are likely many more lesser evils being committed with similar tools that are just much more difficult to spot. I would not have noticed my particular example if I were not a reviewer for the paper, for example. It makes me wonder how big the problem really is.
This seems to confirm my suspicion than these cases are not so much about AI-generated content but rather a result of machine translation.
It's also a common technique for people who don't speak English to translate it... In fact, quite a bit more common.
And sure, there are reasons to translate a phrase from English, to another language, and back to English. This will be familiar to most people who've studied abroad, or done technical conferences on foreign soil, things of that nature. Let's say you're from Bolivia, attending a lecture at an English university, and are planning on referencing some of the content in a paper you're writing, in English. You speak passable English. The lecturer gets into the meat of the topic, and you realize you don't quite understand the context of what they're saying. Some of the conjugations are unfamiliar, so you just write it down as best you can and move on. Later, when writing the paper, you need a way to untangle the phrasing. A simple way is to put it into a translation application, translate to Bolivian, then try to parse it in native tongue. However, you know you have to explain and discuss this section in English; by translating it back, you'll get the English words, but some of the context and grammar structure will be from familiar Bolivian.
So yeah, never say never.
> This will be familiar to most people who've studied abroad
It wasn't to me until now but this explains a lot!
You've never had to do that before? Maybe I run into it a lot due to the nature of the conferences I attend. I know just enough of the language for functional conversation, but as soon as a complex idea is put forward, I need to be able to contextualize the familiar scientific portions of it quickly, and the round trip translation usually helps enough that I can parse it correctly.
Not really, but I came across some teammates in college who didn't seem to be able to follow a conversation yet they seemed to have quite good writing skills. This might be the reason why ;)
Also, they speak Spanish in Bolivia ;)
Good to know; I've always wondered at the peculiarities of different translation engines, but never really dug into them, as most modern ones seem like neural network black boxes to me. I was pointing out that there are some realistic use cases for doing round trip translations. I've used this technique at a few conferences to help straighten out my hazy understanding of a complex idea in a language I spoke quite poorly. And I do agree it is bad form to use this directly in an academic paper.
> Also, they speak Spanish in Bolivia ;)
To be fair they speak Bolivian Spanish, along with many other native languages! I chose Bolivia as a random target without doing any research, so thanks for the pedantic push to go learn something new; things like this are why I do love HN!
Once I re-read the submission I wanted to reject it immediately, but I realized that I should get a second opinion first. So I contacted the editors, who agreed that it was blatant plagiarism. Hence, they rejected the paper once I recommended rejection in my second review. So this wasn't just a conversation where I made some suggestions and the authors used them. Even the editors thought it was plagiarism once they looked at it.
An acknowledgment would be impossible because the review was single-blind. The reviewers knew the identities of the authors but not the other way around. What the authors should have done was just re-phrase where they used the term in the paper. They didn't even need to copy my explanation, to be frank. The paper would worked fine without the paragraph they copied. If they just re-phrased the relevant parts no other changes would have been needed and this whole thing could have been avoided.
In the absence of an explicit directive or request from you, given that the authors are from a different culture, how do you expect them to know what was required by them?
I don't mean to be snarky or accusative. Your comment was thoughtful, articulate and detailed, which tells me you are a sophisticated communicator.
But that said, I've often times wondered if this requirement of having to "rewrite in your own words" may do a lot of harm too. It obfuscates that things that people are talking about are actually exactly the same, or make it fuzzy what the exact differences are.
In a particular academic CS area I've witnessed people reproduce again and again the essentially identical description of setting and assumptions, but in being afraid of plagiarism accusations, they over and over re-formulate things which made it nonobvious that things are the same as from other authors or even from their own earlier work.
1. authors submit a paper with expository sections about (eg) some materials being flammable and others inflammable
2. reviewer tries to explain that they have incorrectly understood the meaning of the terms, explains the meaning carefully and maybe suggests the terms they might mean.
3. Authors copy in the explanation and maybe replace incorrect usages with weird tortured phrases
4. Rejection
Obviously this description reads a little bit silly and things were probably more nuanced in practice. I think I’m probably also being uncharitable towards the authors in the example.
4. Rejection
5. Authors replace terms in verbatim copy with weird tortured phrases
6. Authors submit to other journals, get published
This is a very subtle distinction and it sounds like being more explicit about it would have been a good idea.
Is there any way the authors could have kept your definition, and somehow credited you, even anonymously? Because rephrasing definitions is the pinnacle of wasted effort, and leads to confusion - you are asking them to say what you said, but without using your words.
I find it hard to think of any technical terms which have a fixed, well-specified phrase for the definition, much less ones which, if re-used, don't require attribution.
I mean, there are definitional terms like "one meter is the length of the path traveled by light in a vacuum in 1/299 792 458 of a second" or "the discriminant of the quadratic equation is b^2-4ac". Re-use those quoted definitions and no one will blink.
But, what's "evolution", or "electron spin", or "aromaticity"?
Even something as well-defined and concrete as "cosine similarity" has many different variations:
Wikipedia: a measure of similarity between two non-zero vectors of an inner product space. It is defined to equal the cosine of the angle between them, which is also the same as the inner product of the same vectors normalized to both have length 1
SciKit-learn: the normalized dot product of X and Y: K(X, Y) = <X, Y> / (||X||*||Y||)
towardsdatascience.com: the cosine of the angle between the two non-zero vectors
statology.org: For two vectors, A and B, the Cosine Similarity is calculated as: Cosine Similarity = ΣAiBi / (√ΣAi2√ΣBi2)
uchicago.edu: For vectors, it is the cosine of the angle between those vectors.
datadriveninvestor.com: Cosine similarity of two vectors is just the cosine of the angle between two vectors
While certainly equivalent, these definitions show some creative choice in how they are worded, and thus if copied, should be cited.
(I agree that the creativity level is quite low for some of these, and I believe several people might come up with the same description, but that's a different issue than re-using someone else's definition without attribution.)
It's funny because I actually find your exemples to be supporting my point more than yours.
All of these sentences are translation in plain English of the absolutly perfectly defined and commonly accepted definition of cosine similarity. SkiKit-Learn is even just writing the formula.
uchicago.edu and datadriveninvestor.com even use exactly the same words. I mean, if you came to see me complaining someone was plagirising for writing "For vectors, cosine similarity is the cosine of the angle between those vectors.", I would find that laughable.
That is, even given a technical term with a precise agreed upon definition, the description of that term (eg, in English) does not have a precise agreed-upon form.
Incorrect use of the latter may imply plagiarism, and this thread appears to concern that aspect.
Most definitions are not as simple as "cosine similarity".
Your statement, if true, would mean that most dictionaries would use exactly the same words to describe a given, well-specified scientific concept, yes?
What's "Frame dragging" in general relativity?
Wikipedia: the effect on spacetime caused by a rotating mass ... Frame-dragging is an effect on spacetime, predicted by Albert Einstein's general theory of relativity, that is due to non-static stationary distributions of mass–energy.
doi:10.1126/science.aax7007 : the mass-energy current of a rotating body induces a gravitomagnetic field, so-called because it has formal similarities with the magnetic field generated by an electric current (1). This gravitomagnetic interaction drags inertial frames in the vicinity of a rotating mass. (quoting from the preprint at https://arxiv.org/abs/2001.11405 ).
einstein-online.info: a mass’s rotation influences the motion of objects in its neighbourhood
doi:10.3390/universe7020027 : The term "frame-dragging" usually refers to the influence of a rotating massive body on a gyroscope by producing vorticity in the congruence of world-lines of observers outside the rotating object.
doi:10.1142/9789812564818_0002 : A major consequence of General Relativity and related theories of gravity is that all inertial frames are local. These local frames are accelerated, warped and stretched, and rotated with respect to each other due to the surrounding mass-energy distributions. While only their relative rotations are typically called frame dragging effects, this phrase describes a broader range of gravitational influences on inertia.
Very different definitions, because it's hard to express that concept in English. And I think re-using a few of the more extensive definitions, without attribution, is a minor form of plagiarism. (Reusing any follow-up explanation is, as you've agreed, definitely plagiarism.)
You might recall that atrettel wrote "the authors were misusing a particular technical term".
If you dig in to the papers on frame dragging, you'll note similar complaints, like https://arxiv.org/abs/gr-qc/0509025 : "Many accounts of these experiments have been in terms of frame-dragging. We point out that this terminology has given rise to much confusion and that a better description is in terms of spin-orbit and spin-spin effects."
> I would find that laughable
So would I, which is why I commented 'that's a different issue than re-using someone else's definition without attribution'.
No, it definitely doesn't unless you significantly extend what I said in a very uncharitable way to reach that point.
> Most definitions are not as simple as "cosine similarity".
But plenty are. As I said previously, you are going to be hard pressed to constate plagiarism on pure definitions unless you go towards extensive paragraph long ones which are more akin to explanations than what I would refer to as a definition as you did in the comment I am replying to. Actually, if you reread what you just wrote, you are yourself using the world explanation and definitely agree that that can be plagiarised as could very unorthodox and original forms of definition.
But when I read definition, what comes to my mind is akin to the cosine similarity example where even publications reuse mostly the same sentences while applying minor modifications to the subjects or adding an adverb. Thus me pondering the close proximity of the words definition and plagiarism in the original comment until I realized the whole thing was actually about an explanation triggering my reply to someone sharing my initial puzzlement.
I showed several examples of definitions for frame dragging, including ones where re-use would, IMO, constitute "a minor form of plagiarism".
Is your view that re-use of those definitions cannot be plagiarism? If so, why not?
Could you stop pretending you are not understanding my point considering your first example nicely underline it and you yourself admitted it would be laughable to call that plagiarism?
I never was arguing there that the copying of everything you might defined even tenuously as a definition never ever constitute plagiarism. That's a complete strawman. I'm going to stop wasting my time here.
That's a universal statement.
I've been trying to argue that definitions can be plagiarized, with examples which are not "tenuous" but ones which are drawn directly from publications.
But about three sentences into the introduction (where you explain all the background) you start going into "there is also Y's that do Z backwards". Which Y's you compare and connect with is important and says alot about how you think about your X. It might even be a new way of looking at it. So telling other people how you think of it can be important.
And another 5 sentences in you start referencing previous work on the topic. At this point you are crediting other and you get to chose whom to credit how much, with the benefit of hindsight. You refer to papers that are useful to people new in the study of capital letters. What you write here helps them much more than a mere list of papers or a google (well google scholar or ADS or pubmed or what ever) result list, because you can provide a good order to read them or which aspect of X's are best explained where. You also name papers that might be useful to practitioners in the field because they have a particular technique or a good explanation of it.
So it is very much worth while of providing the background that others expect at the beginning of your paper. Even if it requires rewriting that first paragraph several times.
I would have hated that level of navel gazing.
I guess the fear is that they don’t actually understand the definition, and anyone reading the paper will incorrectly believe that they do, improving their reputation in a way they don’t deserve?
I agree that this is not a school assignment, so the nature of plagiarism is a bit different. But you hit the nail right on the head. We should be worried that they are pretending to know something that they really do not. They copied my explanation nearly word-for-word. Doing that does not prove that they actually understand the concept. It only proves that they have copy and paste. Now maybe they did spend some time learning it and looking into, but there is no way to know for sure. The only way to prove that they really understand the concept fundamentally is to make them explain it themselves in their own words. And that is precisely what they should have done.
Why are you so concerned on this ? Isn't the goal of a paper is to communicate information?
>It only proves that they have copy and paste.
Yes they copy paste but it doesn't prove that they don't really understand it.
They may very well found your definition is the best way to describe it.
As long as the reader of the paper understand what being communicated, imo it should be fine?
I can understand if its a school assignment where the goal is specifically to prove that the author know their stuff.
For a scientific paper, requiring a concept to be described in different way for the sake of it seem to be inefficient and wasting time.
Unless this very phrase does the tripping off. Having to rephrase it would be pretty ridiculous though (and illustrates your point).
It's not uncommon for the acknowledgements section of a paper to thank an anonymous reviewer for, e.g., suggesting the authors further investigate a detail that turned out to be important. But in this case, where the authors couldn't write their own explanation of a technical term, maybe it's a bit premature for them to be writing technical papers.
It's not a, its b.
Ok b.
Hey, you copied my answer?
???
To preface: I am speaking in ignorance of the mechanisms of credit and advancement in your field, so I'm undoubtedly overly harsh. I'm not even trying to be fair, because I admittedly don't have the knowledge to do so.
I know academia is a different world from private sector industry, but I would think if the point of a confidential review is for you to do anonymous work to improve someone else's credited work, and their work was improved by incorporating your feedback, you would be happy with the outcome, or else why are you participating in a confidential review process in the first place? There are different mechanisms for publishing words that you want credit for.
When someone incorporates my feedback on code or documentation word-for-word, I might in the worst case be suspicious that they are trying to get my approval without engaging with my criticism, but in most cases I'm flattered that they respect my idea enough to put their name on it. Although, in my world, putting your name on something is more about responsibility than credit. The command is called "git blame" and not "git credit", after all.
Wanting them to incorporate your feedback, but also wanting them to make some change to the wording to avoid plagiarism, smacks of how freshman essays are graded, not how real work gets done.
Like I said, I'm not trying to be fair and don't have the right background to be fair to you. I only speak up because I grew up in an academic family and know that there is a presumption that academic work is more idealistic, more altruistic, and less mercenary than private sector work, and I think it's worth pointing out when the reverse is true.
I can't speak for GP as to whether this was true, or if there are different norms for that sector, or the author was slightly aggressive in asserting their rights, or the assumptions we're making about the content of the criticism are off, but given the article brought a new dimension to it, I thought that worth mentioning.
This is one outcome of peer review, and is an explicit goal of (good) peer reviewers. But it's definitely not the main goal of peer review. The main goal of peer review is assessment. Improvement is something that you can strive for as a secondary outcome, but it's not the main point.
This is a significant and important contrast with code review. Peer review is NOT analogous to code review!
NB: The code review style of improvement-focused review also happens in academia! But it happens within groups rather than between groups. I.e., advisors or post-docs in single research group or university critiquing one another's work will behave more like a code review. Peer review is different.
The same is true in industry, btw. Think of peer review as the thing that happens when a regulator reviews the code from a medical device. They're not interested in making pull requests and doing your work for you. Although they do in principle want you to succeed in your goals, and might provide some feedback along those lines, they're making a yes/no decision. Different function.
> their work was improved by incorporating your feedback, you would be happy with the outcome
There are really two concerns here:
1. confidentiality, and
2. plagiarism.
The first item is probably more important than the second. The typical expectation of reviews is that they are anonymous and private unless stated otherwise (eg openreview). It's incredibly bad form to publish a private correspondence without first asking for permission.
Copying a reviewer verbatim without permission is plagiarism, but it's not really something that anyone actually cares about, per se. I've sometimes asked for permission to incorporate components of reviews verbatim into my papers, and never received anything except enthusiastic "of course". I assume the same would be true in the case above. But even in those cases I say something like: "as helpfully observed by an anonymous reviewer of this paper, [insert quote]".
The key point is that I don't go around publishing explicitly confidential correspondences without permission.
> or else why are you participating in a confidential review process in the first place?
Exactly. Confidential.
Again, surely you can see how publishing someone's confidential words is quite rude, even if the person wouldn't mind those words being published if simply asked.
> Like I said, I'm not trying to be fair and don't have the right background to be fair to you. I only speak up because I grew up in an academic family and know that there is a presumption that academic work is more idealistic, more altruistic, and less mercenary than private sector work, and I think it's worth pointing out when the reverse is true.
Academic work is all of those things in terms of goals, not necessarily in terms of process. I don't know anyone who has a passing familiarity with Academia and doesn't realize that it is incredibly competitive.
Altruism and competition are different axes.
I guess I don't understand what kind of "confidentiality" was violated. To me a violation of confidentiality would be revealing your name and your role in the process, or publishing sensitive information they got via the process, maybe publishing some idea or data you shared with them that you intended to publish yourself later. But they didn't use your name, and you didn't share any sensitive information with them, so I don't get it.
Plagiarism I guess I can see, though the cynic in me says if it was so obvious that their native language wasn't English, it was in the best interest of readers for them not to do the dance of paraphrasing away the plagiarism.
> I don't know anyone who has a passing familiarity with Academia and doesn't realize that it is incredibly competitive.
Yeah, my family who are in academia talk all the time about how petty and cutthroat it is, and I'm sure that played a role in scaring me away from trying it myself, but I can tell they think business must somehow be worse. I get the feeling that deep down they believe they’re seeing a version of human behavior somewhat elevated by the ideals of academia, and however bad it is inside academia, outside, in environments ruled by cruder values, it must be worse.
https://news.ycombinator.com/item?id=28112208
https://news.ycombinator.com/item?id=28112358
To sum up, the authors did not just use my wording for a short portion. They copied an entire paragraph of my review nearly verbatim. Both the journal and I thought that the authors acted unethically. They violated one of the ethical standards of the journal regarding plagiarism, and these ethical standards were made available to them when they submitted the paper. Those were the rules that I had to evaluate them with for the review, so my hands were tied in some sense. I would also quibble with saying that they mastered the concept, because I really have no way to gauge their understanding if they just copy my own words.
Isn't that what journals are supposed to be for? To help you reach a wider audience?
"flag to clamor" for signal to noise
"individual computerized collaborator" for PDA (personal digital assistant)
"haze figuring" for cloud computing
"information stockroom" for data warehouse
"focal preparing unit" for CPU
"discourse acknowledgement" for voice recognition
"mean square blunder" for MSE (mean square error)
"arbitrary right of passage" for random access
"arbitrary timberland" for random forest
"irregular esteem" for random value
ETA:
"notoriety examination" for sentiment analysis
It sounds like something you'd find in 30s, 40s, 50s sci-fi for sure! Like “visiplate” (E.E. “Doc” Smith, Heinlein) for a computer display screen. (Along with ticker tape printouts and tape reels in the far future of course.)
E.g.,
Signal -> flag
To -> to
Noise -> clamor
…and… Data -> information
Warehouse -> stockroom
This would be a lot easier than running through multiple translation steps (as proposed elsewhere here).From top to bottom: 信噪比, 個人數字助理, 雲計算, 數據倉庫, 中央處理單元, 語音識別, MSE(均方誤差), 隨機訪問, 隨機森林, 隨機值, 情感分析
Deep-fried versions: “旗幟到喧囂”, “個人計算機化合作者”, “霧霾計算”, “信息庫”, “焦點準備單元”, “話語確認”, “均方錯誤”, “任意通行權”, “任意林地”, “不規則尊重”, “惡名考試”
In all seriousness though, I've experienced something similar before at a Japanese run American corporation as far back as the 90's. The problem was Japanese executives and executive assistants who didn't know American tech-jargon often resulted in accepting mangled suggestions by the spell-checker. A notorious example was the "Data Whorehousing" presentation, which somehow made it through several reviews and rehearsals before being presented to the entire American IT department at an all-hands meeting.
Clearly this made an impact as I remember it 23(ish) years later!
This type of manipulation and plagiarism may be partially to blame, but the academic writing style has also gone completely off the rails to the point that half the journal articles being published today read as if written by some kind of paper writing AI robot even when I am quite certain that that isn't the case. And no, I am not talking about cases where the author is writing in a non-native language.
I have a theory that it may have to do with imposter syndrome and a need to sound smart. The author, fearing that they don't really belong and at any moment will be found out, therefore never making tenure, starts jamming academic sounding words where they don't belong and stretching sentences with commas and semi colons until the whole thing is just as insufferable to read as it was to write.
There is also the possibility that there are just a lot of terrible writers out there.
Of course, writing good science is hard enough for native speakers. It is very difficult for the vast majority of people on the planet - no matter how good their research.
And just so we are clear: Not everyone can afford professional editing services at every point in their career.
We meet in English under the premise that it allows for universal communication. In this, we accept that English natives are almost infinitely more privileged in writing, speaking, conferencing and networking. We also have to accept that the level of English proficiency varies, and - especially English - is easy to learn and so difficult to master.
[1] https://en.wikipedia.org/wiki/The_Chicago_Manual_of_Style
and some of what is available under
[2] https://duckduckgo.com/?q=military+writing+guide
would be useful for american english and technical writing.
> And no, I am not talking about cases where the author is writing in a non-native language.
The actual problem here is fluent English that is written in a totally bizarre style only found in academic papers. I've found that academic-ese is less of a problem in good computer science papers (like the one this article is about), but it crops up in some fields a lot. A trivial and not very important example is the way minor things are routinely described as "novel", a word you rarely find in everyday English, but in the research literature everything is "novel".
Learning to write well in any language is difficult. English is not exceptional as a language. Its influence in economic activity is what gives it prevalance.
Of course, they're not wrong. It is written more like a blog post. Because the writing style used in blog posts is hands down better than the writing style used in scientific papers. Blogs talk about the real reasons you worked on something, they go through simple examples, and they mention where you struggled and what you found confusing and what you tried that didn't work. All of these things are very useful for understanding, and in my experience almost entirely lacking from papers. Or at least, in my experience they're lacking from modern papers. I think in papers from 100 years ago the authors tended to talk more about their worries and their excitement e.g. [1].
Surely they are and writing in a way that is easy to read and understand is an art in itself.
But I would agree, that the main reason is probably the intention to sound smarter, than they are. Whole scientific disciplines seem to live by that standard.
This is not limited to science though, I recall a german poet (I think Heinrich Heine) said about his fellow poets:
You only fly so high like the swallow, that no one can actually hear your singing.
The move from a structuralist account in which capital is understood to structure social relations in relatively homologous ways to a view of hegemony in which power relations are subject to repetition, convergence, and rearticulation brought the question of temporality into the thinking of structure, and marked a shift from a form of Althusserian theory that takes structural totalities as theoretical objects to one in which the insights into the contingent possibility of structure inaugurate a renewed conception of hegemony as bound up with the contingent sites and strategies of the rearticulation of power.
You just can't argue with that.
[1] Ironically.
https://infofranpro.wdfiles.com/local--files/19520101-on-coo...
The generated text was well over 50 pages, completely bypassed all known content/plagiarism checks and was even included in the Universities "exemplary examples". To this day, it is still there.
This is of significant concern as some of these GPT-3 based tools are now integrated within MS Word itself. Word 2021 allows for "add-ons", out of which I have noticed several third party content generation and paraphrasing tools.
I can imagine that plagiators use paraphrasing software quite extensively, though, and that it is a problem.
It was not all automated, there was a fair bit of manual intervention needed. I understand your concerns and they are valid and this is why I preface my statement with "anecdotal evidence". What I write is most certainly not the entire story and a fair bit of detail is left out.
It should be known that this is widespread across multiple industries and this will only become more of an issue in the future.
This is a US-based institution, fully accredited.
I don't doubt this at all, and I have no doubt that GPT-3 with a bit of human editing can spit out something better than the lower third of masters students at corn row colleges.
Masters degrees are cash cows, which is why no one in unregulated industries cares about them. People in regulated/unionized industries also don't actually care; even educators, who at least nominally see intrinsic value in education, go to borderline diploma mills to get that union-mandated raise at minimal effort.
Those doctorates don't require much more than taking some coursework and paying a boatload in tuition. Basically an expensive and length online masters program. Not worth the paper they're printed on, unless you're employed by the government or in a union job that mandates raises for education attainment.
As a general rule of thumb, PhDs from R01 universities that are paid for by the university through research assistantships or teaching assistantships are generally a good signal of at least minimal training in research skills.
Another good general rule of thumb is that paying for a PhD -- beyond perhaps some MD/PhDs or maybe nursing phds, stuff like that -- is always a good sign of someone who has both a meaningless degree and also poor reasoning/research skills.
But anyways, real doctorates outside of a few fields (e.g., pure math) usually come with a non-trivial publication record that speaks for itself. You don't even need to know that the person has a doctorate; you can just read their papers and a rec letter from an advisor describing the student's role in each paper.
(I'm excluding discussion of professional degrees like JDs, PharmDs, etc. which are technically doctorates but sort of their own class.)
Due to the current market saturation of math doctorates, any pure mathematics PhD worth the paper its printed on will also probably come with a non-trivial publication record. The exceptions I can think of are high-risk high-reward areas like cutting-edge number theory (I had a friend go eight years without publishing, which, yikes, but his thesis was semi-revolutionary (or so I'm told)) or, I guess, suitably abstract category theory (though the people I follow in this area seem to publish lots of interesting papers, like the Baez school or the homotopy type theory people; your mileage may vary).
It's really too bad. One wonders why we can't simply ax the entire advisor-candidate system (with all its myriad opportunities for physical, emotional, and even sexual abuse) and certify new candidates by saying: "You're a doctor of mathematics when you get five professors to sign off on 3-5 papers you've had published."
Or one big one.
Basically, take the "honorary doctorates" some Universities give out to people retrospectively to people who have made major contributions to their fields; do it more often; and then make it the only path to getting a doctorate, such that they're no longer "honorary" at all.
University of Phoenix does offer Ph.D's.
I would have never guessed.
First time I've heard the term "corn row colleges". Google's not bringing up anything that looks relevant.
I suggest picking something else. Given that "cornrows" are a predominantly black hairstyle, the term reads like a racial slur.
The name now includes small state schools -- usually branch campuses with lower enrollment and no major (R1) research output.
(NB: corn row colleges are also by definition non-elite, so small liberal arts colleges with billion dollar endowments which might otherwise count, don't).
Many such institutions have since started offering graduate (or at least non-bachelors) degrees and certificates that are somehow even more worthless than their undergraduate programs.
Apparently the name has a lot of different meanings these days -- see sibling comments -- but it has DEFINITELY never been meant as a racial pejorative. If anything, exactly the opposite, since most of those "crap-tier midwestern/southern colleges" cater to 99.99% WASP social networks (the P is even explicit).
Wow, coastal elitism much?
There are surely many degree mills and garbage universities, but to conflate their worth with their location is both injurious to discourse, and incorrect.
I didn't say anything about geographic regions.
I said rural.
The coasts of plenty of rural areas, and plenty of institutions that fit the "corn row college" mould.
I live in Wisconsin, and the state university system is chartered to serve the needs of the state. There are too many students to send them all to UW in Madison, so there are a number of smaller regional universities, many of which now offer graduate degrees, plus an even larger number of "commuter" and "satellite" schools, and an elaborate technical college and trade school system. Not everybody can get a degree at a residential college. Life gets in the way.
We can debate the relative prestige of these colleges, but I've worked with people who attended the regional schools, including many engineers and computer programmers. All I can say is, send me more.
The colleges that catered to pastors were largely private, and in my home state, there was one in every town. Some of them emerged as full service 4-year colleges with additional programs. My undergraduate college was nominally "Christian" but I got a secular science education there, and it was well ranked in science. It also adjoined a seminary where I never set foot.
I agree with your assessment that most of Wisconsin's land grants are quite good, btw. YMMV in other states, unfortunately.
People don't care about masters degrees engineering, law, business, art, etc. etc.? Try applying for many jobs without one, or with one from lower-ranking colleges.
The Chronicle of Higher Education article recently on the HN front page said that masters in some fields, they give the example of 'positive psychology', are indeed cash cows. But in the example, that degree was not part of the actual Department of Psychology, which is taken very seriously.
I haven't even heard of a Masters in Law (law degrees are doctorates), but I can't imagine it's worth the paper it's printed on.
MBAs are worthless unless they're from a few good places, and even then the brand and networking does a lot of the lifting.
Law degrees are called Juris Doctor but are professional degrees, like MBAs. You aren't required to publish original research (afaik) and in the US they were formerly Bachelor of Laws (LL.B.) and then renamed (as I understand it).
The doctorate is Doctor of Juridical Science (J.S.D.). You can also get a Master of Law (LL.M.).
So, uh, is Business Administration a regulated industry?
I don't mean this rudely, but it is attitudes like this which cause the CS interviewing process to be 100X more painful than the interviewing process in any other field: "I don't trust your credential so I demand you prove your competence to me on the spot and let's do 5 rounds of interviews just to be sure."
Well, yeah?
I've heard similar stories about generated phd theses and it is even more implausible. The reason is that writing a thesis is much more than just producing a hundred pages or so of prose. Any university student can poop that out in a few weeks. The main job of a thesis is coming up with a research question, conducting an experiment or a study, and describe the results and how it fits in whatever niche of the scientific world you are working in.
I regularly get dissertations with any or all of: barely readable English, useless empirics, half-baked research questions.
Presumably this is an arms race against things like https://www.turnitin.com/
Empower students ‘to do their best, original work’ and this is what you get. Though what the alternative is, I have no idea.
Doesn’t it stay published forever? Might be a shame for the someone during their career.
On the other hand, even a chapter of Mein Kampf was accepted in 20 journals, after replacing the old word with newer versions. Human reviews are hard. Maybe we should put computers in charge of reviewing papers, they’d recognize the work of AI quicker?
https://www.foxnews.com/us/academic-journal-accepts-feminist...
That's nice, how did it do the defense?
For anyone who has been a webmaster, one can immediately recognize it's an extremely common technique in the blackhat SEO scene for decades, used by content farms everywhere. One just copies articles from somewhere else, replace all words with dictionary synonyms to evade the search engine penalty, and fill the resulted websites with spam.
Perhaps it's not as popular in the English world, but common in China, and is a standard tool included in all blackhat SEO software. And no, it doesn't work well, the output is gibberish too in spite of the language differences. Oh, and the article says:
> A high proportion of these papers came from authors in China.
Exactly what I expected. The spammers found a new market, apparently. It's sad to see that some scientific papers and journals are literally becoming blackhat SEO spam and content farms.
The first one I found was about dog illnesses. They kept referring to dogs with phrases like "Your domesticated canine," and it was quite a chore trying to figure out most of the symptoms that they were listing. "Heart worms" was translated to "love snakes," which I thought was delightful.
In this case, sometimes you get lucky and can actually find meaningful information between the padding. But sometimes you just read an article that takes 5 paragraphs and 500 words to say "we don't know".
I don't have a specific example at hand, but it's typically an article a few paragraphs long with really strange phrasing, so strange that it's not explainable by the author not knowing English well.
In a handful of cases, I've managed to find the original source. Common phrases are systematically replaced by ill-fitting synonyms.
I suspect the motivation is to avoid accusations of plagiarism (though I don't know what benefit the posters derive from doing this).
Content generation was one of our bottlenecks, and as Google was already rather successful at detecting duplicate content, we were looking for a way to "uniqify" posts that would be used to stuff sites intended for googlebot, but not humans.
One of the methods that worked was taking source English content, running it through Babelfish, the Altavista translator to French, Spanish or German, and then using the same method to translate it back to English.
This resulted in texts that did not make much sense to humans, were full of precisely such "tortured phrases" but which were considered unique by Google.
Fake papers and plagiarism are the most blatant form of corruption. They tend to come from certain, let's say, large countries with less developed scientific cultures. Those countries need to put an end to it, because the rest of us keep having to work harder to suppress the racist impressions that we're bound to form of colleagues who look and sound like the cheats.
In more traditional scientific countries, the corruption is more subtle. Today, many groups publish every paper with half a dozen authors, and no indication of what each of them contributed. This enables the professors who run those groups to manipulate authorship more or less as they please, and have total control over who gets to have a career in science. It turns out that absolute power corrupts senior scientists as absolutely as it does other people.
No doubt there are more clever ways to game the system, that I haven't noticed. As long as million dollar grants and first-world citizenship keep being doled out for something as contrived as scientific paper authorship, corruption is inevitable.
Ceterum censeo Elsevier(um) esse delendum.
> The journal’s publisher, Elsevier, launched an investigation. This is still under way, but in mid-July the publisher added expressions of concern to more than 400 papers that appeared across six special issues of the journal.
I hate to open up this topic and I hate to pick on people that are trying to fix their mistake even more, but oh boy. Elsevier has been a pain in the butt for universities and researchers alike. They leech money from both sides of the community, they sue people trying to bring science forward and they gate scientific success. And their literally only reason to keep existing was to prevent exactly this.
I've never been a big fan of the current scientific publishing model. But Elsevier is a top publisher. It's pretty damming that they have one - highly overpaid - job and they don't even do it.
complex Hilbert space -> Complicated Hilbert space
Quantum gate -> Quantum door
> Unfortunately, fibroids are just one of many understudied aspects of health in people assigned female at birth. (This includes cisgender women, transgender men and some non-binary and intersex people; the term ‘women’ in the rest of this editorial refers to cis women.)
The article says it only refers to "Cis women" (presumably "cisgender women"), however the article continues to talk about rugby and brains, in which case the word "woman" would not only refer to "cisgender" women, but also to those people who identify as non-binary or transgender men, as surgery and hormone therapy (if that is undertaken by the individual) won't change brain axons, or the person's physical stature.
The article then talks about "male animals", not "animals assigned male at birth". There's no explanation given why animals are not similarly "assigned" a sex.
AFAIK today it has became necessary to "disguise plagiarism" even when you are not plagiarizing anything because bullshit "anti-plagiarism" software would detect many phrases similar to what somebody else already used. I believe the war on plagiarism brings little good in exchange for the hassle.
> In our strong opinion, the root of the problems discussed in this work is the notorious publish or perish atmosphere (Garfield, 1996) affecting both authors and publishers. This leads to blind counting and fuels production of uninteresting (and even nonsensical) publi- cations.
Here's "Microprocessors and Microsystems."[1] This is supposed to be about embedded systems, which is generally a no-bullshit field. I'd never heard of this journal. People read Electronic Design, EE Times, "Embedded.com", maybe Control Systems Journal, etc. Those have either articles about how to do something, or "why what we're selling is great" articles.
Now look at the article titles in Microprocessors and Microsystems.[2] Here are the first three.
- COPS: A complete oblivious processing system
- A perceptron-based replication scheme for managing the shared last level cache
- Efficient underdetermined speech signal separation using encompassed Hammersley-Clifford algorithm and hardware implementation
Now those might be legitimate, although what they're doing in an embedded systems journal isn't clear. They're all behind a paywall, so it's hard to tell if they're any good.
"Oblivious processing" is a security concept. That belongs in a journal on security and encryption, where the crypto people will know what holes to look for. (Microsoft was doing work in this area in 2013, but I don't think a product emerged. If you can make it work, some cloud computing company can use it.)
Cache management belongs in a journal on CPU design, where people who have struggled to make caches work will take a look. There are people using perceptrons for this, which makes sense; a cache has to guess which things will be reused. (If this works well, someone should be trying it in web caches such as NGINX to improve cache hit rates.)
Signal separation is an active field, but this isn't a journal where you'd expect to find articles on it. Wikipedia has a good article on signal separation. The history of that article indicates attempts to sneak in citations to sketchy articles. No idea if the Hammersley-Clifford algorithm is even relevant. (If it's a significant advance, there's commercial value in this in improving audio quality for conferencing systems.)
So these papers were all sent to a journal where the odds of getting published are good, and the odds that the editors have no idea about the subject matter is high.
Why is Elsevier even publishing this journal?
[1] https://www.sciencedirect.com/journal/microprocessors-and-mi...
[2] https://www.sciencedirect.com/journal/microprocessors-and-mi...
[3] https://en.wikipedia.org/w/index.php?title=Signal_separation...
Luckily, I'm fluent enough to recognise the particularly egregious examples, but finding good translations for technical words is hard!
One example that comes to mind is when trying to translate the phrase "data feed" which came back as "alimentation données" which ostensibly means "animal feed data".
If you're looking for a lot of English-to-French translations of technical terms, check out the theses any English University in Quebec (McGill, Concordia, etc..). They're made public online [0]. Can't vouch for the quality as I'm sure there are plenty that just use Google Translate, but everyone I know has their abstract edited by a francophone in their field.
A good way to validate translated technical terms is to just give them a quick internet search on e.g. DuckDuckGo or Semanticscholar.
[0] McGill's is https://escholarship.mcgill.ca/
But then again you can view _all_ solutions to social problems as inherently technological in the broader sense; I adhere to that paradigm.
They see it as a game that they’re playing and they’re doing their best to put as little effort as possible into the game while extracting as much reputation upside as they can.
We really need to make publishing fraudulent papers a career-ending move across academia and even the industry. The only reason this continues to happen is because it has a lot of upside but very little downside. Caught publishing fraudulent papers? Oh well, just leave them off your resume and apply somewhere else.
It's the next step in clickbait monetization. Why settle for low-effort content when you can have no-effort content?
My favorite example, go to google or DuckDuckGo and type:
“Xxx number hospitalized” or “yyy new cases”
You can type almost any number and get a ton of articles. Not exactly a reprint, but they all seem almost generated
Rewriting headlines that bots wrote and A B testing humans vs Software
You must read each paper to judge its merits. Lots of junk gets published in top ranked journals.
Lots of junk gets published in top ranked journals.
A lot more get published in vanity journals, so I use the impact factor as a first pass filter: I avoid papers from journals not listed the JCR¹ or those with a factor below 1.000.I assume, maybe naively, that if an important finding were to be published in such low quality journal, it would eventually get published in a more legit publication.
1- https://www.researchgate.net/publication/342623066_Journal_C...
University libraries who continue to pay Elsevier should know that they are propping up scammers and grifters.
Would that be considered good?
My favorite counter example is "A Draft Sequence Of A Neandertal Genome". The article was accepted by both Nature and Science before it was written. The authors chose to publish in Science, because Science offered more on the side: the title page and an unlimited(!) number of "contributed" (this means unreviewed) companion papers. The article itself was about 20 pages of drivel; all the substantial content was relegated to the 200(!) pages of "Online Supplemental Material". Nobody ever read, let alone reviewed, all of that.
After that, I can't trust either Science or Nature, which offered pretty much the same crooked deal. If those two aren't "highest profile", who is?
The highest profile journals (Nature, Science, The Lancet in medicine, ...) have some tendency to go for sensationalism. They want to publish radical, ground-breaking research more than there is actual new ground-breaking results happening. So they also end up publishing mediocre research presented as ground-breaking, and some less-than-accurate research where results are exaggerated to make them look ground-breaking.
(I'd like to see EV World or something else in that space reprint old articles as "1, 5, and 10 years ago in battery hype".)
Maybe that's where the term came from?
Though if they've published a convenient list of the bad papers, then, assuming other markers exist, that makes it easy for others to discover them.
'Colossal information' in place of 'Big data' ... wow - so wrong.
https://www.ctvnews.ca/health/offshore-firm-accused-of-publi...
That's why there's no real fix for this problem beyond defunding government science budgets. Any quick hacks you can come up with like running GPT-2 detectors over science papers are just treating the symptoms of the problem, not the cause. The root cause is that when governments eliminate the free market by subsidizing research they lose reliable signals of genuine utility that come from polling the market, so they have to use proxies that boil down to quantity-over-quality. The fix is for them to stop subsidizing research. The people who apply research to create new technologies aren't reading it anyway.
What am I missing?
> Out of 404 papers accepted in less then 30 days after submission, 394 papers (97.5%) have authors with affiliations in (mainland) China. Out of 615 papers of which editorial processing time exceeded 40 days, 58 papers (9.5%) only have authors with affiliations in (main- land) China. This tenfold imbalance suggests a differentiated processing of papers affiliated to China characterised by shorter peer-review duration.
How so? It seems quite related to me. Anecdotally, one would expect a pretty clear negative correlation between torturedness in the sense of this article and indicators of research quality.
https://www.seroundtable.com/auto-translated-content-spam-go...
Example-
Unknown Source -> stolen using Machine Translated English and Published -> Stolen again, Machine Translated language X -> Machine Translated English -> re-publish on another web site -
Result - https://m-eng.ru/en/plumbing/sinii-kit-vesit-150-tonn-zhivot...
So people are using machine translate to steal stolen machine translated articles.
Heck, we got a word for this sort of rampant plagiarism masking on Chinese internet — 洗稿 (manuscript (or blog post)-laundry).
OT: I do appreciate the funny phrase “elite figuring” for HPC. It’s kind of like how they translate things to Anglish.