[1]
[1]
The main problem is that there is an objective need (or desire?) by various stakeholders to have some kind of metric that they can use to roughly evaluate the quality or quantity of scientist's work, with the caveat people outside your field need to be able to use it. I.e. let's assume that we have a university or government official that for some valid reason (there are many of them) needs to be able to compare two mathematicians without spending excessive time on it. Let's assume that the official is honest, competent and in fact is a scientist him/herself and so can do the evaluation "in the way that scientists want" - but that official happens to be, say, a biologist or a linguist. What process should be used? How should that person distinguish insigtful, groundbreaking novel and important research from pseudoscience or salami-sliced paper that's not bringing anything new to the field? I can evaluate papers and people in my research subfield, but not far outside of it. Peer review for papers exists because we consider that people outside of the field are not qualified to directly tell whether that paper is good or bad.
The other problem, of course, is how do you compare between fields - what data allows you to see that (for example) your history department is doing top-notch research but your economics department is not respected in their field?
I'm not sure that a good measurement can exist, and despite all their deep flaws it seems that we actually can't do much better than the currently used bibliographic metrics and judgement by proxy of journal ratings.
Saying "metric X is bad" doesn't mean "metric X shouldn't get used" unless a better solution is available.
You trust in people, in consistency of physical laws, in coherency of your own mind, in constancy of temporal flow, in so many things because - as the Pyrrhonist refrains - nothing is certain, not even this.
Would both help solve the replication crisis, and resolve this problem.
Of course then you might have 10 000 studies replicating the same easy to do study... which is why the "score" should be reduced based on how many other times that study has been replicated.
Or, I dunno, paleontology or sociology or other stuff.
Citations can be a useful metric here, particularly if you can identify citations of people actually using the method (as opposed to people just mentioning it in passing, or other methodological researchers comparing their own methods to it).
An insightful study that's replicable (but has not yet been) is valuable. A lousy study that's been replicated five times (not because it's interesting, but because it was easy to do, and the replicators knew that they'd be rewarded for replicating anything) is not valuable.
A metric that says "number of studies" is IMHO even more arbitrary, more gameable, and more detached from actual value than citation count - which does have some notion that your study actually matters to other poeple; that it was worth writing that paper because someone read it.
Also, I really question your notion that people outside a field should be able to evaluate the quality of someone's work, especially in academia, where the whole point is to be well ahead of what most people can understand. That theory seems like part of managerialism [3], which I'll grant is the dominant paradigm in the western corporate world.
I understand why a managerialist class would like to set themselves up as the well-paid judges of everybody else. But I'm not seeing why anybody would willingly submit themselves to that. It's a commonplace here on HN that we avoid letting managers make technical decisions, however fancy their MBA, because they're fundamentally not competent to do it. That seems much more important for people doing cutting-edge research.
[1] https://en.wikipedia.org/wiki/Goodhart%27s_law
That’s not the case at all. Being at the leading edge of research should mean that you are creating new knowledge. That doesn’t imply that people cannot understand it. This expectation that laypeople cannot possibly understand science is one of the reasons so many papers are written so densely and obtusely. “They” can’t understand it anyway, right?
Feynman said if he couldn’t explain it to freshmen he didn’t understand it himself.
I do agree that researchers should be able to give decent "here's what I do" explanations to the general public. But that's very different than a member of the general public understanding the context well enough that they can judge the value of the work to the field.
It's about the question of resource allocation. Pretty much every subfield of academia is a net consumer of resources, i.e. someone outside of that subfield is funneling resources to it. That someone - no matter if it's a university, or some foundation, or a gov't agency, or a philantrophist - needs to make a decision on how to allocate resources. And, in general, they honestly want to make a good, informed decision on which projects and researchers to support; but nonetheless they have to make a decision according to some criteria. So there's no choice of "no metric", there will always be a metric and we can only argue that it should be better. And the answer to "why anybody would willingly submit themselves to that" is that duh, you don't get a choice - you can suggest a better method to fulfil their goals of allocating resources in a way that is (also in their opinion) fair and objective; but you can't get around the fact that scientists are generally funded by nonscientists. And they need(or want) to make decisions.
They could delegate that, but that doesn't solve the question about the criteria - if they delegate that to universities, they still have to decide on how to allocate between departments; if they delegate that to scientist councils uniting all the departments in the country working on some subfield, they have to decide on how to allocate between the different organizations. So no matter what, you have to compare not only quality of similar scientists, but also of dissimilar scientists working in different (sub)fields. And delegation doesn't absolve you from responsibility, so if the money is (or looks!) wasted, then that's a failure - so when you delegate, you want to require them to use objective criteria. Which is hard - I could tell you which researchers in my subfield are doing excellent work and which are useless; but if I had to justify these decisions, to demonstrate why they're not just my bias because of politics/liking certain methods/gender/ethnicity/etc then it would actually be tricky; and I think that I'd actually reach out for these metrics. And I'm quite certain that the metrics (for the people that I have in mind) would agree with my subjective opinion; on average, the great research gets cited much more and is in higher-ranking venues; while the lousy stuff gets no citations apart from the author's only grad student.
Also, there's a lack of trust (IMHO not totally unwarranted). You could get a bunch of experts who are qualified to evaluate who gets what amount of resources, you can't rely on them actually doing so - if we take spiders as a totally random example, in general you're qualified to distinguish which spider research is good and which is useless only if you actually work on spider research, most likely in one of these teams - and the expected result is nepotism, allocating resources based on purely (intra-field) political reasons. And who'd decide on how to split resources between spider research and bird research? Do you expect the spider guys and bird guys to reach a consensus? Or would it go to whatever field the dean is in? This is a big problem even currently, and a big part of why the metrics are being gamed - but at least metrics are something that require effort to game and can't be gamed totally; if we'd do away with them, then we'd be left with absolutely arbitrary political allocation, which would be even worse.
So at the end of the day "they" need some way to transform the only reasonable source of truth - actual peer-review - to something that "non-peers" can use to judge what the the aggregate of that peer review says. That need is IMHO not negotiable, I really believe that they do actually need it - they don't want to do resource allocation totally arbitrarily, they want to do it well, they need (because of external pressures) objectivity and accountability, and currently this (journal rankings, bibliometerics, etc) is the best what we have suggested to summarize the results of that peer review.
If I had to write a law draft for a better process of allocating resources, what should be written in it?
Again, since this is a tech community, let me use that for an analogy. It's a classic problem for non-technical founders to evaluate their technical hires. They aren't qualified.
The right solution is not to find some gameable metric of tech-ness, like LoC/day or Github stars. Instead one uses either direct experience-based trust or some sort of indirect trust, like where you have a technical expert you trust and have that person interview your first tech hires.
Yes, having expert humans make the decisions is imperfect. But it's not like a managerialist approach is either. And the advantage of using expert humans, rather than a gameable metric and managerial control, is that we have centuries of experience in how people go wrong and many good approaches for countering it.
Existing community/expertise based moderation and reputation systems might not be directly transferable or adequate. But it shows there are new ways to think about more decentralized measures of reputation that are new to this century and haven't been tried. New ways that may be preferable to a small group of kingmakers.
I think the biggest problem is leadership and cooperation of community to try something different. It's not just that there is no person who can mandate these things, it's that multiple constituencies have widely diverging interests, i.e. authors, universities, corporations, journals.
I also don't think it's a problem that different groups have different interests, etc. As I say elsewhere, I think that diversity is the solution.
>I also don't think it's a problem that different groups have different interests, etc.
I don't see how you can disagree that cooperation of community to try something different is not a major hurdle.
How many years has it been since important issues in the academic process were widely known? How much success in adoption has there been to date, regarding any fundamental changes?
It seems on its face to be crucial.
We fund academic work because we see value in it. But there are many kinds of value, and many different sorts of value. So I think it's appropriate that we have many different universities which have many different departments. Many different funding agencies and many different foundations. Each group has their own heuristics for picking the seed experts.
There are still systemic biases, of course, but that's true of any approach. And distributed power is much more robust to that then centralized power or a single homogeneous system.
The bureaucrat in this case would be another university professor working in the same field.
I disagree with the assertion that bad metrics should be used if there are no alternatives. Bad metrics give wrong answers, and only the illusion of meaningful information. The most common use of bad metrics is to lie to people, and it isn't the scientists using the metrics but the organizations that employ them.
Why should an easily measurable metric which has meaningful value exist? It doesn't seem obvious to me that it should at all. Determining the capability of a researcher is inherently a very complex intellectual task. The desire is to reduce that task to something which removes the need for the person doing the evaluation to read and understand the produced research, or to even understand the field of study in many cases. Perhaps, instead, those who are put in charge of things like awarding grant funding, granting tenure at universities, and deciding who to hire to teach ought to be expected and required to evaluate the research on its merits. This would greatly increase the intellectual sophistication and capability needed for people in those positions, but the alternative will always be fairly easily exploitable because it is easier to goose a metric than to do solid research.
We see the shortcomings of trying to reduce complex intellectual challenges to checklists or metrics all the time. And we simply ignore the alternative of relying upon intellectually capable people meeting the challenge. Personally, I don't understand why.
And for what it's worth, I've almost never heard impact factors discussed at NIH study sections, where investigator quality is explicitly on the agenda. Reviewers talk about relevant prior publications in the field, esp in marquee journals. [this latter feature is the reason we don't just put everything on biorxiv or equivalent and move on.]
We do need a metric imo, but I agree we don't have a perfect one yet.
[Link Test] [/Link Test]
Even better split it into individual contributors to give a count of researchers who have cited the paper?
Carl Bergstrom is a smart guy so I suppose the practical implementation of the above must have some wrinkles, but with enough brute force it seems tractable. What I despise more than anything is the gaming that takes place for “impact factor”.
I do OK by standard metrics but would very much like to know where I stand by less easily gamed metrics of influence.
In a very loose sense, PR is the same algorithm universities use, evaluate quality of some content based on the number of references to that content.
It is definitely gamed in similar ways. I'm surprised we haven't seen professors hire SEO firms to help increase citation counts of their research.
In fact PageRank was inspired by academic rankings in that aspect:
"PageRank was influenced by citation analysis, early developed by Eugene Garfield in the 1950s at the University of Pennsylvania, and by Hyper Search, developed by Massimo Marchiori at the University of Padua. In the same year PageRank was introduced (1998), Jon Kleinberg published his work on HITS. Google's founders cite Garfield, Marchiori, and Kleinberg in their original papers."
And suddenly Google is the authoritative source on literally everything in the world. I hope you like their political views, because they would become "the one".