Reasons blog posts can be of higher scientific quality than journal articles
daniellakens.blogspot.com
daniellakens.blogspot.com
There is short shrift given to the distortion that gets introduced in order to "get published". The "file drawer" effect is very real, and it gives us a very skewed view of what the actual scientific work being done really is; p-hacking and other statistical trickery are table stakes to anyone trying to survive in the academic world, as being published is a key metric to success yet it is surprisingly orthogonal to being right. As Brian Nosek's "reproducibility project" is uncovering, there is a non-trivial amount of "science" out there that is pure garbage, and much of it was produced in order to "get published."
In a world where "random hackers" can write twitter bots to parse scientific papers and uncover obvious mathematical flaws (which sometimes invalidate central claims of the paper), I don't understand how we aren't immediately gravitating to standards which promote more and more openness.
Lastly, the biggest issue we have to address is the fact that notoriety is now the chief indicator of success, and not genuine scientific discovery. In college I had world-famous faculty who have made a career, fortune, and placed themselves well on the spotlight on the backs of research which was later completely reversed -- and not in the way "science is supposed to work" sense, but rather in the "you purposefully fudged the math" sense. And yet there was virtually no reputational damage done to their career. Worse yet, I think it is clear that this lack of ethics actually is a direct cause of the success they enjoy today. Given that these are pressures faced by all in the academic community, true openness is the one avenue we have to counteract this sort of thing, and hopefully enterprising individuals come along and find a way to make "truth-seeking" reputation something researchers care about.
What I enjoy about learning about other subjects that make claim in research papers is that it is at least legitimized through a rigorous process. One can substantiate that y is better than x because tests show doing y requires less computing time than doing x. Or that doing y results in higher accuracy of the model. While getting a research paper published is an imperfect process where even high quality ideas get rejected for at times pretty trivial reasons, I can appreciate the process itself mostly results in a sharpened idea.
Parallax scrolling sites come to mind. They look pretty, but quite often are terribly dysfunctional to navigate and use for a customer who just wants to learn about or buy(!) your product.
Here's a study that disagrees with you, at least in terms of how many people have trouble with parallax (2 out of 43). http://uxpajournal.org/the-effects-of-parallax-scrolling-on-...
Not to say that study is a comprehensive judgment on the design pattern.
> We hypothesized that PS would improve UX, which is defined in this study as the emotions that are aroused when a user interacts with a product or technology. The PS website was perceived to be more fun than the no-PS website. With respect to perceived usability, enjoyment, satisfaction, and visual appeal, there were no differences between the PS website and the no-PS website.
Great that someone formally studied it, though I'm a little put off by their subjects demographics as compared with potential customers of any given business. That is, I feel like students at a university are more adept at using interfaces than "the average user" who is going to be your potential customer.
1/5 factors they tested did signal added value, the 1 being "fun". More fun, but not more usable. I'm not yet convinced that fun correlates to sales in this context as much as usability, ESPECIALLY once the initial wow factor wears off.
I assume by "trouble" you mean motion sickness. Two of them had motion sickness, not just run-of-the-mill confusion from the usability of the site.
All that said, the main point I was making was that often designers spend a great deal of time trying to get PS working just right not because they think it will bring in more customers or make their site more usable, but because that's what other designers are showcasing and they want to emulate that trend.
This gets ignored the second they get $800k in VC money but for some reason the same idea never translates to "real science" or academia.
The more open publishing becomes, almost by definition, the more dilution happens to the citation process. Want to know where this happened before? Trace back to Google's origins, where the "brilliance" of the "citation" technique made it so much better to begin with. But what happened afterwards? Was the interlinking just left as it was, pristine and still as valuable, just waiting for Google and other search engines to exploit it? Of course not. Any kind of system which can be gamed will be gamed. And what happened to the internet after Google became the sheriff? I would say it is now a fairly overcomplicated and twisted system (especially for the hobbyist) with the cat and mouse game going on between SEO companies and Google.
The distortions that are introduced today in order to get published, will be replaced by new distortions that will be introduced to get cited.
I understand your idea but the thing is, nobody has tone of free time to read an unknown hacker's blog (if he/she really does). People are busy for working so the best way is to read a peer-reviewed paper. This thing is like a brand, I guess.
My point with the part about the bot that checks errors is that, if we are being honest, in far too many cases that "peer review" is less than ironclad, to put it mildly. In that scenario, the journals actually do us a disservice because they attempt to signal that the content you are reading went through a strict bit of scrutiny before it reached your eyeballs. That is a dangerous assumption to believe in if it turns out that it really wasn't.
By all means, when you read a paper, be skeptical and go through it yourself, but in my field, referees do actually provide some quality control.
That's what we need to find a way to alter: If you are doing great science -- including the painstaking work of carefully reading a paper and following minus signs through calculation and checking integrals and whatnot -- that should be explicitly rewarded in the profession, and shouldn't just rely on intrinsic values a researcher holds in order to survive.
I don't see journals today being the leaders in driving that innovation.
Also, before you propose tearing down a system that has produced results for decades, but is flawed, you have to show that the replacement is actually useful and capable of at-least performing on par. I have no problem with casual discussion on HN (in fact I actively avoid the insane people who want "proof" for everything) so I don't really care to hold your comment up to a gold standard of absolute certitude, but you should probably give your ideas a bit more thought.
Do you have an example of this?
Good journals do have peer review, and they wouldn't have accepted this.
Also, one big disadvantage of blogs -- depending on where you host it, they often don't last long, and point 4 (easy to edit) is a blessing and a curse, particularly if people edit things without telling you they did it.
Of course, blogs have their place, I keep one myself, it's great for breaking news, explaining things in greater depth, sharing and understanding. And journals could learn from blogs / arxiv (make it easier to make corrections, allow discussions).
Had it been a blog post, I don't think anyone would have taken more than a cursory look at the analysis. Yes, anyone can fit data points and reach a conclusion with blogs, but isn't academic research fraught with similar problems already?
[1]: https://www.theguardian.com/technology/2014/jan/22/facebook-...
Academic literature has many problems but it is way above the blogosphere in that regard. And media coverage is going to suck no matter where stuff is published...
From the perspective of peer review, yes. But per google scholar that paper has been 66 cited times, including in many articles that are in at least somewhat reputable journals (published via Elsevier, ieee, Springer, etc) and apparently two books [0]. So I think it's pretty clear that being a "scientific paper from Princeton" was significant reputation wise regardless of whether or not it was peer reviewed.
[0] https://scholar.google.com/scholar?q=Epidemiological%20model...
I was also skeptical about the citations so I checked that two with freely available pdfs were actually this paper, they were
[1] https://pdfs.semanticscholar.org/4ebb/b654c59aff454ad97301e6...
Some papers are cited not because they are well written, but simply because they covered a topic which might not have been covered by other researchers. Literature review is a fundamental part of a paper.
I also know a guy who published preliminary results and musings on his blog, and was then refused publication because it was "unoriginal".
Journal editors seem to not be familiar with the subject matter either, given what does and doesn't get published - why only ever positive results?
Unless scientific publication changes completely science as we know it will die. It is already becoming vanity, egoism, orthodoxy. Why publish? So you can get better pay later in the private sector.
1. George Box has a very famous quote: "All models are wrong; some models are useful."
2. Even the model applies correctly for Myspace, it does not mean it is correct for Facebook. I would use a very common terminology - "survival bias/selective bias" for it. The only way to ensure the model is letting it through numerous amount of cases. There is not only one social network is active, can the author predict for them and for the old one? Here is the possible list: https://en.wikipedia.org/wiki/Social_networking_service
3. When they cited a paper, they are not supposed to agree with them. Some of paper even criticize each other. They may accept the concept of Facebook being cool down but not in couple of years (with different parameters). If you want to make sure they are all agree or not, please make a survey on this. Like I said, you (and probably me) are probably into survival bias.
4. There is a least one mistake here: "discussed in Section ?? [3],". So, was it well proof read?
Peer review, looking at it from the outside as a layman, appears to just be an enormous rubber stamp. I cannot believe that academics have time to rigorously go over the mountains of pages of dreck that make up most research papers, and even read the whole thing thoroughly, much less double-check the numbers. Not when they have a stack of papers to review a foot high, and another two feet of grants to write, their own research to do, not to mention what little teaching they haven't farmed off to grad students.
Most peer reviewers do not get to see the data in the article.
The conference software they were using had a sketchy form for uploading supplementary data, but it appeared that they expected it to be used for small data tables or something. It certainly wasn't going to be an option to upload 12 GB of data into their conference software; they mentioned nothing about code, particularly how to specify its dependencies and the computing environment it needed to run in; and it also certainly wasn't going to appear in any form that was convenient to review if I did.
How could you submit code to blind review anyway? Do you make alternate versions of all the dependencies that don't credit any authors that overlap with the authors of the paper?
In short, because of blind review, I was forbidden from doing anything that would make my paper reproducible at review time.
Of course, both are needed for great science. Nonetheless, your paper alone should provide a good enough description of the conditions and methods you used for the paper to be reproducible.
In fact, it can easily be argued that the ability to just run your code instead of re-implementing it according to your description is actually detrimental. To see how, consider that a bug in your code, re-used by other researchers, can easily lead to multiple derivative works finding entirely wrong results. In contrast, if those other researchers re-implemented your method there would be a much lower probability of them doing the same mistake you did, leading to incongruent results and hence raising alarms that probably lead to the discovery of your bug. Although re-implementing your algorithms is significantly more work, the overall quality of our research would benefit from doing so...
Too bad there's an 8-page limit, then.
On a personal note: I'm sick of smart-ass reviewers, unprofessional researchers, results over-selling (to put it mildly) and so on. I think the whole system is so corrupted that i just quit research in despair. Unfortunately, I don't see how blogs and/or just "publish the code" would magically solve all those problems.
Anyway, publishing your code is certainly a good thing to, so I applaud and encourage you to keep doing it!
This is a dated view of AI research.
Promising research appears in workshops, blogs, and/or arXiV. Finished research appears in conferences. Nobody's sure what journals are for.
Of course if you are in one of the Ivy league universities this may be different. Otherwise... yeah, you can feel research moves too fast for journals, but curricular evaluation practices moves even slower.
I agree with GP. The view that conference publications is less important than journal papers is inaccurate, depending on your field.
In my field, it's as you describe. Conference papers are not polished, and journal publications is what matters in your academic career.
For many of my peers in some disciplines in CS, it was the opposite. Getting into a highly regarded conference was much more valued than publishing in a journal.
Then I found this was not limited to CS.
It really just depends on your discipline's culture.
I wrote my bachelor's thesis last semester. You wouldn't believe for how many papers I needed to use archive.org. By the end of it, I made a small donation and recommended the company to do the same.
Blog posts might not be better, but hosted papers aren't that stable either.
I think is they key point here which the OP seems to ignore completely: I sort of agree with everything the OP states, but he also forgets some key aspects on blogging. For instance anyone can post whatever even on subjects he/she really doesn't know a lot about. Also, not every blog post does get reviews from all sides and might only get 'yes good' replies originating solely from confirmation bias while the lack of any criticism doesn't make everything right but for instance just indicates the critics didn't find their way to the blog post.
Whereas most scientific bloggers post things they are interested in and therefore feel have something to say.
You have a point but from all scientists I know myself there's not one for which this applies. Might depend on the field though and I'm not denying there's a bunch of impossible-to-reproduce-crap published and there are rotten apples everywhere. But what I see (in fundamental research which is usually not as publicly known as other kinds) doesn't come near what you describe, even though it's anecdotal of course.. Yes there is pressure to publish but their papers aren't 'scinetific-looking', they're properly scientific and they are just the length needed to describe the findings on the subject. And there is an abundance of motivation for reviewing mostly inspired by wanting to make sure everything is as correct as possible, for the sake of research.
These grants are aimed at building a working technology and strengthening national economy, so there's nothing wrong with the approach per se, but when your reviewers share your values they have little motivation to find every possible mistake in your paper, at least for second- or third-rate journals (which are still indexed in WoS so are perfectly enough for the funding agency).
Good networks concentrate intelligence by providing feedback and support that helps the network converge on useful, original insight.
Bad/malicious networks destroy it with noise and false information, creating a feedback network that suppresses reality-based original insight.
The process of blogging is irrelevant. So is the process of peer review.
It's the quality of the networks around each area of interest that defines the real value, not the process.
E.g. peer review works well when it's part of a high quality network. When it isn't, it's no better than random posting.
Free and gated idea exchange create completely different communities, with completely different problems and advantages.
We're seeing a huge amount of non-traditional scholarly activity (if you'll pardon the dry phrase) happening in blogs. Alternative metrics aka altmetrics (as distinct from altmetric.com) have been taking in to consideration the scholarly activity that happens around traditional publishing for a while. Crossref, the organisation who brought you DOIs for scholarly publications, thought that it would be a good idea to help collect this kind of data as a counterpoint to traditional publishing and citations. We're building Crossref Event Data, which is a free (libre, gratis) service for collecting mentions of articles on blogs and social media, so that it can be used by the community in all kinds of ways. Discoverability, recommendations, and yes, maybe more metrics. How you use it is up to you.
The article raises a good point about blogs as primary methods of publishing rather than, for example, as a venue for the discussion of traditionally published articles. Establishing an open 'citation'-graph-like-dataset of blogs is a good first step toward that.
We're heading into Beta soon, and you can read more about it https://www.crossref.org/services/event-data . The User guide is a work in progress, but might answer any questions you have: https://www.eventdata.crossref.org/guide . You can also contact me at jwass@crossref.org if you have any questions.
I find it extremely problematic. Sure, sharing data is necessary to verify that the conclusions are supported by it and they are not due to methodological errors. But to anyone? Out in the open? I would expect better data ethics, especially from a psychologist.
Your data can contain political opinions, health records, sexual orientation, contact information, the places a person has visited, and when. People sign up for these studies usually with the agreement that the information that can harm them cannot be freely shared, unless to people involved in studies with similar data protection systems. Data like I mentioned has sometimes to be stored in computers not connected to the Internet, to reduce the risk of data leak. Free access to this data paves the way to persecution and shaming.
If I were to sign up as subject to a psychology study and had a person with his ideas leading it, I'd withdraw immediately. I'd question if this person should be a psychology researcher at all. Sharing data is good, but protocols are there for a reason.
However, he put in the introduction of his post a warning about his field, as if he expects that his arguments applies especially to experimental psychology.
In many other cases, as pointed out, it's just not possible. I'm studying mobility patterns through cellphone metadata. Even if you strip out the actual phone number with a random ID you still know where a person is going, and thus re-identify them if you have other public data.
That leaves about 500 non unique individuals in a country of > 50 million inhabitants.
DE Denning, PJ Denning, M Schwartz, ‘‘The tracker: a threat to statistical database security’’, in ACM Transactions on Database Systems v 4 no 1 (1979) pp 76–96
A general tracker can always be found, unless the data released is extremely restricted. Almost anything is personally identifiable as it can be used to build a tracker into the database.
I am aware of this result from chapter 9 of Security Engineering (http://www.cl.cam.ac.uk/~rja14/book.html) by Ross Anderson, if you are more generally interested.
There are a few elements they emphasize.
One is what the blog format enables that traditional publishing doesn't support. Those are things like having real-time feedback and comments, being able to version and making the blog post interactive, rather than a static document. Another element to the format is a lack of gatekeepers, so it can be quickly disseminated and disseminated by anyone, so there aren't barriers to participation in the scientific discourse.
Another is norms and expectations. In blogging, it is more the norm that data and code are open. Open is still possible in traditional publishing; it just isn't yet the norm. A new format however, enables new norms and it's easier to set them from the start, than try to revise existing ones.
Finally, there's the element of 'correctness'. Going through peer review and being in a traditional journal certainly doesn't ensure that the paper is correct. You can just look at retractions to see that http://retractionwatch.com/2011/08/11/is-it-time-for-a-retra.... However it would be interesting to see more evidence around whether the blog format does ultimately lead to more 'correct' conclusions, on the whole, and not just for the posts that lead to a lot of discussion.
For anyone who is curious, some conferences share the peer review comments (even though they do not share the identities of author or commenter). Here are the comments from iclr2017 (https://openreview.net/group?id=ICLR.cc/2017/conference) and here are comments on Ian Goodfellow's GAN paper by Jürgen Schmidhuber and others (https://media.nips.cc/nipsbooks/nipspapers/paper_files/nips2...).
1. people that want to publish their research, theories, response, etc
2. a mechanism for having submissions peer reviewed
3. curating which submissions will be included in the publication
4. actually physically (or electronically) publishing and disseminating the papers to the interested community
5. people actually reading it and responding (possibly recursively)
6. reputation and status derived from authorship, being referenced by others, etc
7. archiving the papers - and hopefully the data
These features are traditionally part of the same process for practical reasons like the cost/time to physically publish before modern printing technology. This worked, creating resistance to any change. Now, I suspect that over reliance on the concept of a "journal" limits your thinking. This article demonstrates some of that with it's framing of blogs as an alternative or competitor to journals.
Instead, consider that modern computing and the internet make #4 very easy and almost free. We have various ways to archive (#7), which include actual verification (e.g. "git fsck", signed commits). We already curate (#3) as a separate step with specialty blogs and aggregators like HN.
I'm not saying journals are bad or obsolete. I'm just suggesting that there might be better ways to organize the process, and that different granularities of "journal"-like process can probably coexist.
If the hacker side of these challenges interest you, I suggest checking out the code4lib community: https://code4lib.org
No, they don't. This is conflating having an opinion with the scientific review process. The scientific review process can prevent publication of bad research in journals and require changes before publication, the blog "review" process cannot do either of these things.
It's not uncommon for an industry expert or creator of a product to Koolaid Man into the comment section around here.
The thing is, the very small exception is incredibly good.
- CreativeWork: http://schema.org/CreativeWork
- - BlogPosting: http://schema.org/BlogPosting
- - Article: http://schema.org/Article
- - - NewsArticle: http://schema.org/NewsArticle
- - - Report: http://schema.org/Report
- - - ScholarlyArticle: http://schema.org/ScholarlyArticle
- - - SocialMediaPosting: http://schema.org/SocialMediaPosting
- - - TechArticle: http://schema.org/TechArticle
Thing: (name, [url], [identifier], [#about], [description[_gh_markdown_html]])
- C: CreativeWork:
- - P: comment R: Comment
- - C: Comment: https://schema.org/Comment
A bit cynical, but a lot of the problems the writer suggests are more of a symptom of trying to get as many papers out as possible than the cause (and that one rarely reviews the supplementary material). If scientific blogs counted like publications many of the same problems would appear there.
It's similar to how Andrew Gelman always sarcastically refers to the Proceedings of the National Academy of Sciences as the Prestigious Proceedings, because that's invariably the adjective used in news reports, yet many of the articles are of the "female-named hurricanes cause more damage" type.
So you get this effect of very hot research that the professors can gloat over, relatively easy to read, well-publicized, but actually are lacking in much scientific data or detail when it comes to the actual text. In the field I work in I need access to DNA and protein sequences that should all be freely available. When I find a Nature paper the sequence is invariably missing, or presented in the supplement as a PNG screenshot of a microsoft word file. Again, it's not a detail a professor or casual reader would care about - but it's a detail that's absolutely critical to the science and for replication.
Most journals are neither so condensed to the exclusion of data/protocols, nor so well-read as to actually create issues when it comes to replicating, correcting, or retracting them.
This can be done, but I suspect it is rare whereas it's the norm with a journal article.
Ultimately, I think they're both good, but they do have different uses.
Archiving, for me, is a requirement if you think it should be used over more than a year or so.
Gwern has tackled link rot[0] in an exhaustive way, and may give you a starting point from where to continue.
There is no reason why this post should be focused on blogs in general, asides from generating discussion + clickbait. The key here is actually following the best practice, but do most blogs actually do this?
Alternatively: there is a high minimal requirement for research articles, but there is no minimal requirement for blogs.
(i) not written in clickbait format
(ii) impact is measured in terms of other research which sees it as a relevant study to cite rather than numbers of pageviews (see also (i))
(iii) one of the things a peer review process can flag up is that bold and implausible claims like "On blogs, the norm is to provide access to the underlying data, code, and materials" probably need some form of qualification for what is meant by "blogs" and/or quantitative evidence to support that claim
Blogs can be well reasoned and evidence-based, comments can (even more infrequently) be enlightening and journal articles aren't immune to flaws, but the linked blog article is a prime example of why arguments expressed in blogs are seldom afforded the respect of arguments advanced in journals.
Really? I find them to be incredibly clickbaity whereas blog posts are much more matter-of-fact. Blogs are almost never a commercial entity/enterprise and while no writer would shun search engine traffic, they usually write for peers, not for publicity like journalists or scientists, both whose funding depends on this.
Scientific writing in closed-access journals is the definition of writing for peers, and never includes AdSense, affiliate links or digital tip jars.
(See HN passim "Every attempt to manage academia makes it worse")
There's alot of value to the existing journal system, but perhaps it would be appropriate to have a more formal process and a more informal/collaborative process. Maybe that would address the problems with having so many journals.
I don't know academic publishing well enough to compare. I think that as far as controversial claims that would make it to somewhere like Hacker News, the effect is that if the controversial statements are wrong, then the comments will shoot them down, and the story will be buried. Furthermore, the rest of the comments will pile on the first comment that points out that the blog post is mistaken.
Perhaps an effect of this is that you can tell from the tone that some blog posts take, that they are being extremely careful to lay a rock-solid foundation, and kind of "convert" their readers who might otherwise vehemently disagree with them. For a recent example, check out this blog post:
https://gregfallis.com/2017/04/14/seriously-the-guy-has-a-po...
I actually just noticed something really interesting! The very first words in this blog post are: "I got metaphorically spanked a couple of days ago."
Would any academic article on any subject start in this way?
So the effect of the backlash against controversial statements, on the writing style of blog posts, is interesting and in many cases highly visible.
It is hard for me to compare this with academic writing. (But it certainly wouldn't start with words like that.)
In terms of error correction, there are errata that are published if someone finds an error in a major study and in journals like the physical review are linked back to the original article.
The open data problem is hard. In my particular field, our raw data is available on the web, but the problem is that the meta-data needed to interpret is not. And that's nontrivial. For example, I may perform several operations on my data (for example, background subtraction) before fitting it to a model. We are experimenting with dataflow languages where we embed the series of "filters" that we apply to the data, but if we want someone to be able to reproduce this 20 years from now, then there's a whole ecosystem that would have to be maintained--for example, not just my code, but all of the libraries that it depends on to run. We can describe the basic process in the paper, but for true long term reproducibility, it's a hard problem...There are groups that are working on open-data and reproducible research, so I don't think that there isn't interest, just that it's not as easy a problem as you might think. But, for reduced data, I agree, it would be nice to have that available for papers in a machine readable format...
Open access. This is also a hard problem. Increasingly, funding agencies are requiring that publications be made available in an open format after some embargo period. I think that may be the best we can hope for. Just paying for competent editors requires funding. If we want additional features like data attached to papers (for decades), it will take more funding. That money has to come from somewhere. For typical open access journals, the author pays, but that seems to create difficult incentives--not to mention that it makes it difficult for poorer funded researchers to publish. I'm not sure what the right answer is, because I agree that the public should be able to see the results that they paid for--perhaps an embargo period is the best solution...
I think a blog is a great way for communication of ideas and for education, but I do not think that it is able to replace publication in a refereed journal.
"Why can... "
Uh, shouldn't that be "How can... "?To be fair, some fringe journals have no scientific value either, because they do not peer review properly, but they are easy to spot and everybody in the scientific community knows them.
On a side note, originality is one of the key requirements for publication in a reputed journal and I have not ever seen a blog post that made any scientifically original claims. But maybe that's because I don't read blogs very often.
Bunch of mathematicians collaborating through blogs.