Inciteful: Using Citations to Explore Academic Literature
inciteful.xyz
inciteful.xyz
The biggest hurdle was the speed of the graph creation. Basically taking a 250,000,000 paper/2,500,000 citaiton db and creating graphs that could be up to 200k papers and 3-4mm citations. For that I ended up learning/using Rust (which was a great experience).
The plan is to keep it totally free and hopefully get some institutional support once I get a better handle on demand and costs.
Ask me anything!
EDIT: As you are going through the site, be sure to use the purple "+" buttons to create your own graphs centered on the topic of your choice. That combined with the in-graph keyword filters are probably the most powerful ways to quickly zero in on the most relevant literature.
But long story short, I end up doing most of the graph analysis by passing in the citations, using PyO3, to graph-tool in python then returning the data I need about each paper. I am planning on moving that over to Rust. But not being an academic I wanted to get feedback on the quality of the results before making it difficult to quickly test different types of algorithms.
Do you have any plans to add a graphical visualization of top/central papers?
https://metacademy.org/graphs/concepts/bayesian_logistic_reg...
and src code for the graph view is here: https://github.com/metacademy/metacademy-application/blob/ma...
they do some clever hiding of edges so graph is not overwhelming, but still only O(100) nodes.... for O(100k) nodes you'll need to do some selection for sure ;)
I love this idea btw, I’m going to use it to find some holiday reading!
This one in particular had some very nice features, some of which are present in Semantic Scholar (my current favourite) but some which are certainly not.
Recommending papers based on citation graphs is a good way to very quickly get up to speed with fields I'm not to familiar with, but I'm always wary that I'll end up back in the feedback loop of very few popular papers rising to the top while perfectly good papers go unseen because they weren't well cited in the year they were written.
So I'll certainly keep an eye on this and give it a try, but I'm certainly still in the market for a "serendipity" slider on such recommendation engines.
1. Find a paper you like in a field you want to learn.
2. Use the keyword filters to filter down to papers that match your criteria.
3. Add a bunch of the interesting ones to a new graph using the purple "+" buttons.
4. On the "new" graph page, check out the similar papers section. If any of them are interesting, add those to the graph.
5. Repeat until you don't find anything else that is interesting.
The similar papers section uses a link prediction algorithm that basically says, if two papers cite a bunch of the same papers, rank them higher BUT if the paper they cite, is cited a bunch of times, don't give that connection much weight. The net effect of this is that it doesn't really matter if the paper was highly cited, only that it cites the same niche of papers as the ones you just chose. Also, because of the temporal nature of academic literature, the papers it brings up tend to be the newer and harder to discover papers.
The results are pretty great and it's as close to the "serendipity" slider that you'll get right now.
EDIT: Formatting
It is a nice user interface and the reference material is useful and well presented. But when you get down to it, it's linear lists about the characteristics of a semantic graph. As I've said many times, this is like describing a tree with a tabular catalog of its leaves. Graph navigation needs a graphical representation, because things like branchiness (node out-degree) and other factors are more easily shown than described.
The basic problem with graph representation/ navigation/ traversal is that there are many valid ways of looking at the graph and it's hard to render them all. Maybe try using a gutter to allow users to temporarily pin certain graph characteristics and render accordingly. In this context, sometimes I might be interested in the latest research that cites a paper, other times I might be looking to see who picked it up first, or to apply some sort of windowing function to a large spectrum of citations.
But I want to see the graph, even if I am looking for a particular leaf.
I like the idea of interactive filters to allow the person to explore it visually. I'm hoping to have something people can play with in the next few weeks. I hadn't really expected the site to take off the way it has and so it's not really feature complete yet :)
But what you're discussing is a frequent problem with graph visualization - it's very easy to end up with 'hairball charts' that may be meaningful to the person who generated them but only because of familiarity with all the steps it took to get there, and the more inclusive the graph, the more time eaten by clustering algorithms, the more clusters produced, and the more of a cluster...well you get the idea.
As you're at an early stage of development, perhaps this technique, which trades away completeness for clarity and is relatively novel, might let you leapfrog some of those problems: https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=8896846
Two things that might be tweaked:
* The search didn't behave in an ergonomic way: I typed a query ("graph neural networks") and great relevant stuff came up immediately in the dropdown. When I hit enter, however, I got an error that read "Invalid search: Check your spelling, enter a DOI, or another paper identifier or." I would have expected my action to take me to a search results page that listed what I saw in the dropdown (which I regard as a preview of the top hits) so that I could peruse the selection carefully.
* I wanted to load a paper to take a look at it and it took me a while to realize that I could click the "Yes" above "Open Access" to download it. Since one of the big use cases for a site like this is the eventual consumption of these papers, I suggest making a "read/download paper" call to action more explicit.
Can you tell me how you're managing copyright? In my understanding, you're effectively "republishing" abstracts, possibly altering or remixing them too. Presumably the site isn't a commercial concern yet, so there's probably little issue. If the site starts to generate revenue, even if it's just to cover costs, or if it becomes a registered business, does this change the copyright situation for you?
I see that you're using several sources of open data, so perhaps all the data you're using is free from copyright, or has highly permissive copyright e.g. CC0.
For those wondering about open-access, in my understanding that's about _reading_ papers. Putting papers on a website might make the website provider subject to copyright. This applies even to abstracts.
As a minimum, it might be necessary to have a button on your site, weishuhn, to report any paper that has restrictive copyright. You've already noted one example of misclassification, and I know that Crossref and other sources provide their data with a big caveat that it might have inaccuracies. Responsibility cascades to you, unfortunately.
This inevitably leads to a trade-off between completeness of your data vs usefulness of your service. It's a problem I'm wrestling with too!
None of this is criticism, just thoughts from someone working in a similar area. As others have said, it's always good to see innovation like this in literature searching. Good work!
I would like it if the bibtex entries had meaningful cite keys as opposed to long numbers. as is, it would be pretty difficult to actually write a paper using these bibtex files.
“on it's head” should be “on its head”.
“not only with” should be “with not only”.
“analysis'” should not have an apostrophe and should possibly be “analyses”.
https://inciteful.xyz/p/186039733?&keywords=hello&maxDistanc...
I checked a paper, 10.1097/RHU.0000000000000563 for example; which I know for a fact is open-source marked as non open source.
Otherwise, very nice tool. Will be using this regularly. Awesome stuff.
https://api.unpaywall.org/v2/10.1097/RHU.0000000000000563?em...
According to them it's closed but I definitely see it as open here:
https://journals.lww.com/jclinrheum/fulltext/2017/09000/omeg...
I'll look to see if I can find another data source for OA papers to add more coverage.
> Otherwise please use the original title, unless it is misleading or linkbait; don't editorialize.
> Don't solicit upvotes, comments, or submissions.
But I don't think the language is relevant for the direct inciteful.xyz site itself. Better to submit both links separately than trying to combine them as they have different audiences.
Regarding moderation, its a thankless task which I don't envy and its hard to draw a line over nitpicky article titles when one has been voted in already.
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...