But again, I haven’t read the paper! Just chiming in with some half-forgotten epidemiology. :-)
1,186 karma · joined May 27, 2009
But again, I haven’t read the paper! Just chiming in with some half-forgotten epidemiology. :-)
Baxter G, Rooksby J, Wang Y, Khajeh-Hosseini A. The ironies of automation: Still going strong at 30? In: Proceedings of the 30th european conference on cognitive ergonomics [Internet]. New York, NY, USA: ACM; 2012. p. 65–71. Available from: http://doi.acm.org/10.1145/2448136.2448149
Depending on your institution, and how they do their accounting, this can have serious negative effects for investigators. At an academic medical center, the majority of PIs are NIH funded, so all of the accounting and finance planning for research assumes a 54% indirect cost rate- so if, for whatever reason, your portfolio of grants doesn't fit that mold, you can have issues. I have known PIs who had multiple very large grants from private foundations, and so were not producing the expected amount of indirects and ended up being a net negative to their department's bottom line. This caused them (and their department chair) all kinds of problems.
Accounting and grant budgeting are two of the things that I wish I'd learned more about in grad school!
> Search systems, like many other applications of machine learning, have become increasingly complex and opaque. The notions of relevance, usefulness, and trustworthiness with respect to information were already overloaded and often difficult to articulate, study, or implement. Newly surfaced proposals that aim to use large language models to generate relevant information for a user’s needs pose even greater threat to transparency, provenance, and user interactions in a search system. In this perspective paper we revisit the problem of search in the larger context of information seeking and argue that removing or reducing interactions in an effort to retrieve presumably more relevant information can be detrimental to many fundamental aspects of search, including information verification, information literacy, and serendipity. In addition to providing suggestions for counteracting some of the potential problems posed by such models, we present a vision for search systems that are intelligent and effective, while also providing greater transparency and accountability.
Shah & Bender (2022) "Situating Search". In Proc. CHIIR '22 https://dl.acm.org/doi/abs/10.1145/3498366.3505816
And in fact, most/all CS papers _do_ have a "real" version _somewhere_ that has the date and other relevant info, but in CS, we have a bad habit of passing around links to a papers that point to a random version that the author posted somewhere, rather than to an entry in a proper bibliographic database or to the canonical/archival PDF of the paper. I.e., not the PDF that the author posted themselves somewhere (likely missing dates, publication information, revision history, etc.), but the version that came as part of the conference proceedings or journal that the paper was published in. Those typically (though admittedly not always, depending on the conference) have headers/footers added that include whatever would be needed to properly cite the paper. Google Scholar etc. try their best to link to the "right" place, but are often led astray and point at the "random early draft on the author's website" instead, which of course perpetuates the problem.
Incidentally, helping avoid this situation is the sort of thing is at least nominally a big part of the "value add" that traditional journal publishers are supposed to be adding- keeping track of citation/bibliographic metadata, assigning and managing DOIs, ensuring that there's a standardized layout and production process that includes such information in PDFs, providing archival/canonical URLs for papers, etc. It's also an important part of what professional societies that have publishing arms (ACM, IEEE) and libraries (like the NLM and its PubMed/MEDLINE services) contribute.
But is also available online as a preprint here: https://mlstory.org/
And the transformation itself is non-invertible, so it's not possible to recover the original values for about 7-8% of the rows in the dataset. The commit diff links to a thorough investigation of the data[1] in which the author takes a crack at linking up the ambiguous rows with the original 1970 Census data that supposedly went in to generating this dataset, and long story short, it looks like the original dataset's authors may have made some errors in their calculations on top of everything else.
1: https://medium.com/@docintangible/racist-data-destruction-11...
It remains a huge challenge from a budgeting standpoint, though, I don't want to downplay that too much. And there are also other challenges- until very recently, my university didn't have job categories for software-related research staff roles, and so we had enormous difficulties figuring out how to hire and appropriately pay such people. The "how do we hire them" part has been fixed, which is a big help- but we still have to figure out how to pay for them ourselves.
It’s a fabulous and under-utilized resource!
- Turing’s Cathedral, by George Dyson
- Black Software, by Charlton McIlwain
- Programmed Inequality, by Mar Hicks
The Dyson book is a rigorous and deep historical dive into the philosophical and practical origins of digital computing, and is really great.
The other two are equally great and deep but cover computing history through different lenses. The Hicks book in particular may be of interest for you, as its emphasis is on the history of computing in the UK. They’re less directly about how computers “work”, as such, and more about how computers and society have interacted with one another in interesting and non-obvious ways, and how those interactions have impacted the ways in which technologies have developed.
And yeah, "dysfunctional" only begins to describe the Columbia bridge debacle. But I will say, in defense of the project and its staff: that was a spectacularly challenging problem to solve. One of those fractal problems, where the more closely you look at it, the more complex it gets. The physical challenges were already going to be hard enough, and then the twelve-dimensional politics of the situation made it even more so. I don't pretend to know what the best solution would even be, but like everybody else in town I was disappointed to see it blow up after so much had been invested in it.
Like anything else, it's about tradeoffs. This series of planning decisions had a ton of downstream consequences, both positive and negative (and I'm happy to enumerate plenty of examples of either). I certainly don't enjoy the amount of traffic that has come with the city's growth --- I'm a lifelong Portlander, and have watched it happen --- but I firmly believe that more freeways would not solve the problem. Induced demand is a thing, and there is also geography to contend with. For example, one of the biggest traffic choke points is the freeway coming in to downtown from the western part of the metro area; since the 80s, there have been several major (one might even say "meaningful") projects that have widened it about as far as it is possible to go, but at this point there's literally nowhere else to put another lane. Of course, part of why it's a choke point is that after the tunnel it immediately intersects with another ill-sited freeway --- one of the last ones built before our "revolt," and one whose siting and construction caused logistical problems that the city is still dealing with, fifty years later.
Anyway, it's complicated, but the local government "betting the future on light rail" is not the reason traffic has gotten as bad as it has. Population explosion, geography, and seventy-odd years of path-dependent decisions about freeway placement and urban planning are the reasons.
I did think that the author did a good job of outlining (some of) the basic structural issues that make this a tough field to monetize, but even setting those aside, there's no substitute for actually knowing your users and what they need, and that's something the NLM is amazing at.
(Disclosure, my PhD was funded by an NLM training grant, some of my research is funded extramurally by the NLM, and I have a lot of NLM colleagues, so I'm maybe a little bit biased)
Also, the mechanical process of effectively reading a paper is highly non-linear, and is a skill in and of itself. In a lot of ways, it is more akin to high-level pattern matching than it is to more "normal" reading. At least at my institution, it is something that we actually teach our students to do in formal ways (the obligatory "How to read a scientific paper" lecture during the first term or two) and then make them practice over and over again for years (journal clubs, etc.). The original author eventually figured this out, which is to their credit.
The whole book is superb, and while some of the content is a bit long in the tooth at this point, the evaluation chapter has aged extremely well.
Oh, and if you’re interested in how to evaluate search engine user interfaces, the equivalent chapter in Hearst’s book on Search UI has you covered: https://searchuserinterfaces.com/book/sui_ch2_evaluation.htm...
> I think you all know that I've always felt the nine most terrifying words in the English language are: I'm from the Government, and I'm here to help. A great many of the current problems on the farm were caused by government-imposed embargoes and inflation, not to mention government's long history of conflicting and haphazard policies. Our ultimate goal, of course, is economic independence for agriculture, and through steps like the tax reform bill, we seek to return farming to real farmers. But until we make that transition, the Government must act compassionately and responsibly. In order to see farmers through these tough times, our administration has committed record amounts of assistance, spending more in this year alone than any previous administration spent during its entire tenure. No area of the budget, including defense, has grown as fast as our support for agriculture.
From this 1986 speech: https://www.reaganfoundation.org/media/128648/newsconference...
I actually agree with your immediate statement here, but that is not at all what I understood the OP to be saying. I read their "You should know what major type of finding..." argument as being one of starting with a concrete and well-formed research question, and carrying out carefully designed experiments, and thought that it was excellent advice.
In my little corner of computer science, I very frequently see people (at all stages in their scientific careers) start working on some new bit of research by a) getting a bunch of data, which they then b) feed into some nifty model du jour, and then c) spend a ton of time overcoming all manner of technical trials and tribulations, then finally d) get a number out the other side. They then e) find themselves totally stuck when it comes to actually interpreting their result, because before they ran their "experiment" they hadn't actually bothered to formulate a concrete hypothesis, and so it's not clear what they are supposed to _do_ with their shiny new number, or where to go next.
That's what people often don't get about science, whether it's wet or dry. The mechanical process of actually performing the experiment itself is usually the easy part, relatively speaking. The hard part is thinking carefully about the thing you're trying to study, formulating a theory, coming up with testable hypotheses, and designing experiments to perform those tests.
Part of that last phase involves planning ahead very carefully to what you're going to measure, what your control and intervention criteria will be, what specific statistical analysis you'll perform on the resulting data, and what your various interpretations will be. The more concrete and explicit you can make this, the better: "We're going to measure 'X' under conditions 'A' and 'B', because we think that 'X' will be a valid/useful measure of $PHENOMENON_WE_CARE_ABOUT, for reasons ____, ____, and ____. If X_A ends up being bigger than X_B, our interpretation will be ______, and if X_B is bigger than X_A, we will instead conclude ______'; if they are the same, that will suggest ______." [1]
Obviously you don't yet _know_ which of those conclusions you'll be drawing (if you did, it wouldn't be an experiment), but it is absolutely essential that you've gamed out the various possibilities to at least this level of detail _before_ you do the experiment. This is doubly true for exploratory analyses where you don't really have an intuition about what the outcome will be, as it helps keep the analysis from turning into an endless fishing expedition ("Well, what if I normalize this variable _this_ way? OK, what about _that_ way? ...").
In my experience, one of the best techniques for doing this is, yes, to actually write out blank versions of the tables that you think you'll need to tell the story of your experiment ahead of time, and to make dummy sketches of the various figures you'll need to help interpret the data. Not only will this help you clarify your thinking about what you are hoping to learn from doing the experiment, it has the added benefit of making sure that whatever code you write actually logs/outputs all of the needed data elements! More than once I've had to re-do an experiment because there was an important piece of data that I hadn't realized I would need until it was time to do the analysis. With just a bit more prior preparation, that poor performance would have been prevented.
To return to the OP's argument, they weren't saying that you should pre-specify your conclusions (which would be a terrible idea, for the reasons that you clearly spell out in your post). They were saying that you should have a plan about what specific experiments you're going to run and _how_ you're going to describe the motivation and results of those experiments.
And, if I may editorialize for a moment here, having a more structured approach to doing and writing about research can go a long way to helping to reduce the angst that comes with doing a PhD. I do very much think that many CS PhD programs are dropping the ball in terms of teaching experimental design and evaluation- but that is a rant for another time, as my TED talk today is already running long enough. :-D
--------------------------------
1: Very, very, very often, the process of formulating things this way takes several iterations, because usually once one is forced to write it out this explicitly, all sorts of little questions pop up- "Wait, is that actually what it will mean if X_A > X_B? What if means ____ instead? Hmmm... maybe I should be measuring X', instead? Oh, I'll need different data, in that case, because..."