Reading academic computer science papers
stackoverflow.blog
stackoverflow.blog
https://jeffhuang.com/best_paper_awards/
These are papers that were deemed "best papers" in that year, though obviously may not have turned out to be as influential in retrospect, i.e. they're likely not the papers we consider "best" when looking back today.
Also, do you think ACM opening their archives will have a high impact on which ones you recommend in the future?
ACM archives -- not really, I haven't added any new conferences because it makes each year even more work to update. And I find that nearly all CS papers are accessible through various sources that you can find in Google Scholar or Semantic Scholar (e.g. author homepages, course websites, arXiv, etc.).
Yep. I didn't realize so much has changed with the rest of the site There was a bit of slowdown I presume in mid-2010s where the table felt a tad behind schedule. I did think of personally writing you an email to help in updating if necessary. In fact, I secretly wanted to just mirror your publication page - but stealing someone's thunder wasn't my ballgame :)
Links to the papers: https://docs.google.com/spreadsheets/d/1wS6O7-ZoFL7Cfjgt-kdh...
I've been slowly reading through them - some easier than others. (Some I've pretty much skipped - way too long for my interests).
I've found that the original papers are always super dense and the material has usually evolved to be more explainable, especially when someone has put in the time to compile it as part of a course.
I have not finished it yet, just went through a few chapter, but I can definitely recommend it!
This collection is great - thanks so much for your effort!
Question though, have you got a process for reading individual papers? In college my old CS professors insisted on reading the abstract, then the conclusions, then the references, and then the rest of it.
Summary at - https://derekchia.com/how-to-read-a-research-paper-3-pass-ap...
More Advice from S.Keshav - https://svr-sk818-web.cl.cam.ac.uk/keshav/wiki/index.php/Adv...
I have heard that advice before, but I generally disagree. Unless you are doing a literature search, reading the references is a waste of time. And most conclusions aren't really worth reading either. They typically read as though the author has completely exhausted themselves by writing the rest of the paper, and simply restart the last paragraph of the introduction in new words.
If you are trying to get up to speed on a new field, a good introduction can really help with you reference search.
Part of "how to read a paper" depends on what you are trying to do. If it's a seminal paper that are are just going to read, that is a very different thing from reading a paper in search of a solution (or trying to figure out of your idea is novel).
In this second case, your main task is to decide if the paper is even worth reading. IMO, this takes a lot practice. Fully reading a paper to the extent you really understand it can be extremely time consuming. It's important to be willing to throw out the paper if at any point (no matter how much time you've put into it) it becomes clear it doesn't work for you.
I will generally skim the abstract and/or the last two paragraphs (or so) of the introduction. If it still sounds promising I look through the next section or two, which are usually some kind of "problem setup" and "proposed solution scheme". I skip things I don't immediately understand. If my interest is still piqued, I look for the results section where the plots are (if it's a practical paper). If I'm still interested, I go back to Section II and start reading more carefully. I'll spend more time with tricky math, but not too much. Save staring at the same four equations for two hours for at least the third pass.
Oh, and if a paper is tricky (for me) and seems worth my time, printing it out single sided and laying the pages out side by side on my desk can be really helpful.
Anyone with the money and general scientific interest should consider subscribing to Nature or Science. It’s a fun browse every week and you never know what fields interesting finding might fancy you until you look at the articles.
1. https://pubmed.ncbi.nlm.nih.gov/16904174/ 2. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1213120/?report...
The concepts are complicated enough, but then the writing style in papers is just really strange and unnecessarily verbose and over-complicates even simple concepts.
I end up getting most of my knowledge through blog posts and slides which cover the papers vs. the papers themselves.
I think the issue is that I have a hard time learning things without doing myself. And generally trying to reproduce stuff from a paper is really hard.
Have you ever contacted the authors to request their data? I personally have not.
Specialized terminology and complete descriptions can sometimes help to make writing more precise and correct. They can also make writing less clear - intentionally or unintentionally.
Why would the writer of a paper want it to be less clear? Perhaps to artificially inflate the importance of the work (and/or the length of the paper) or to make it seem non-obvious and "novel." Or perhaps to "fit in" with a certain style of writing or language ("academese" or "paperish") commonly used in the discipline.
My advice is: be as clear (and as simple) as possible without sacrificing precision or correctness. Clear papers with good and significant work are incredibly valuable and are likely to be more influential than unclear papers, even if the underlying work is good.
Since so many papers are unclear and poorly written, you may also find that clarity and good writing can help to differentiate your work.
When regular people publish books, they get proofread and you actually get feedback from people who are trained in proofreading. In academia this not something you can expect from a publisher. And that's why we have lower standards for published academic writings.
I wish there were a culture of (a) always writing a blog post to accompany any paper, and (b) rewriting historical* papers to make them more accessible. On the latter point, undergrads should probably be doing this for at least one paper per semester, despite the general expectation that they ordinarily do not deal with "research".
* Doesn't have to be distinguished ones; just any paper you come across and feel is too hard for you/your peers to follow without more effort than should be necessary, or any paper you thought was interesting but didn't get its due in public, or ones that did but have fallen out of cross-generational memory.
This isn't necessarily the writer's fault at all. I have had reviewers complain that the language in my paper was too informal. I wasn't using slang or something like that. But in so many words, the reviewer wanted _more_ academese. It was the last paper of my grad school career and I was sick of academese. In so many words, I told them to pound sand. That was my only paper to never get published.
Most papers about new results are published in early stages of research. The authors barely understand the results enough to be convinced that they are correct. They can't explain the results well, because they don't understand them yet. If they continue working on the same topic, they will probably come up with a good explanation after a few years and two or three complete rewrites.
1. Many CS papers are presented at conferences, and many of these talks are recorded and available.
2. Look for a paper citing the one you are interested in. The “related works” section often contains brief summaries, which are written with the benefit of hindsight.
Somewhat related, an earlier comment of mine on how to acquire copies of a paper without resorting to unauthorised copying [1]
Motivation is to get more people into it, and also give enough of a flavor of the paper for those who aren't used to reading academic papers, or don't have the time.
If the concepts are new and bleeding edge, the reading papers is the only way to become familiar to them. Otherwise, time is better spend reading books and courses, where the said concepts are explained in relation with the others, forming a bigger picture.
[0] https://github.com/siraben/fp-notes/blob/master/Papers.md
So I recently created a small tool to read research papers together with other people with annotations and comments.
I was researching a lot about NASA's Mars helicopter - Ingenuity and found couple of research papers, so I though I'll publish the list on my blog. List would include title of papers with appropriate links to PDFs. The PDFs are hosted _somewhere_ and they are easily reachable if you just google the research title. I gave up eventually since I'm not sure if this is legal or not.
A known counter-example was someone condemned mostly because of admitting browsing up the directory and realizing that the files were supposed to be protected by a login + password :
https://arstechnica.com/tech-policy/2014/02/french-journalis...
https://medium.com/@parttimeben/how-not-to-read-scientific-p...
Is there a site that gives free access to these research papers?
Anything USENIX falls into this category. In my area (networking/distributed/operating systems) the other (non-USENIX) major venues are also open access (SOSP and SIGCOMM).
Old papers that are not publicly accessible on the web are very likely just a pile of craps and not worth reading; even if it was worth reading, its information should have been compiled in better ways in textbooks/blog posts; if neither were applicable, then you are supposed to be some researcher working in academia and can ask your institute to offer the access.
Short answer: scihub
Some contexts have larger research communities. For example, there isn't nearly as many papers on real-time path planning for agent mutable environments vs static environments. I assume this is because we still don't have Boston Dynamics robots in people's homes. If we could get the cost low enough it may be more profitable to send mining robots to mars than people, but I guess there are other applications as well.
I spent some months trying to find, understand, and implement the state-of-the-art algorithms in real-time path planning within mutable environments(Minecraft). I started with graph algorithms like A*[0] and their extensions. For my problem this was very slow. D* lite[1] seemed like an improvement, but it has issues with updates near its root. Sample based planners came next such as rrt[2], rrt*, and many others.
I built a virtual reality website to visualize and interact with the rrt* algorithm. I can release this if anyone is interested. I've found that many papers do a poor job describing when their algorithms perform poorly. The best way I've found to understand an algorithm's behavior is to implement it, apply it to different problems, and visualize the execution over many problem instances. This is time consuming, but yields the best understanding in my experience.
Sample based planners have issues with formations like bug traps. For my use case this was a large issue. Moving over to Monte Carlo Tree Search(MCTS)[3] worked very well given the nature of decision making when moving through an environment the agent can change. The way it builds great plans from random attempts of path planning is still shocking.
Someone must incorporate these papers' best aspects into novel solutions. There exists an opportunity to extract value from the information differential between research and industry. For some reason many papers do not provide source code. A good open source implementation brings these improvements to a larger audience.
Some good resources I've found are websites like Semantic Scholar[4] and arxiv[5] along with survey papers such as one for MCTS[3]. The later half of this article is what gets me excited to build new things. I would encourage people to explore the vast landscape of problems to find one that interests them then look into the research.
[0] https://en.wikipedia.org/wiki/A*_search_algorithm
[1] https://en.wikipedia.org/wiki/D\\\*
[2] https://en.wikipedia.org/wiki/Rapidly-exploring_random_tree
[3] https://www.semanticscholar.org/paper/A-Survey-of-Monte-Carl... /c37f1baac3c8ba30250084f067167ac3837cf6fd
[4] semanticscholar.org
Although the method uses a variant of A* and might not be that “fancy” in academia terms, it’s astonishing how far it can achieve (see demos like [1] and [2]) and might actually be far more useful to study it closely instead of more theoretical papers.
If not, do you believe there is value in education beyond optimizing TC?
I'm speaking from experience. I got a PhD while working full-time. I enjoyed it and it gave me a lot of perspective. But - if I'd spent even 1/6th the time LC'ing I'd probably have a much higher TC at this point :). No regrets - but its an amusing thought that kinda lurks in the background.
What's TC in this context?
It really annoyed me to not know what it meant, so I spend a bit of time finding that
Putting TC aside (don't most programmers make plenty anyways?), isn't it just cool to learn new and cutting-edge things? Personally, I think so. Curiosity is central to what it means to be a hacker.