Zotero: An open-source tool to help collect, organize, cite, and share research
zotero.org
zotero.org
- as I read new PDFs in the browser, the PDFs are passively downloaded typically in a Downloads/ folder.
- This results in thousands in papers lying in Downloads/ or elsewhere.
- The command p from the script [1] let me instantaneously fuzzy-search over the first page of each pdf (The first page of each pdf is extracted using pdftotext, but cached so it's fast). The first page of academic PDFs usually contains title, abstract, author names, institutions, keywords of the paper; so typing any combination of those will quickly find the pdf.
What is particularly convenient is that no time is spent trying to organize the papers into folders, or importing them into software such as zotero. The papers are passively downloaded, and if I remember ever downloading a paper, it's one fuzzy-search away. Of course it does not solve the problem of generating clean bibtex files.
[1]: the script: https://github.com/bellecp/fast-p
[2]: an illustration in GIF: https://user-images.githubusercontent.com/1019692/34446795-1...
edit: The script has been moved from the gist to the public repository https://github.com/bellecp/fast-p
Basically, a personal search engine with a passively gathered corpus of my experienced content - maybe even filtered at times as in your case where you limited it to academic PDFs (to keep the knowledgebase focused). Kind of like an extension of our human memory.
Consume -> add as extension of knowledgebase -> recall.
Thank you for sharing your workflow - simple and ingenious!
Have you had any issues or thoughts for future enhancements? I can think of a number of other helpful things you could do with the corpus you’ve built, for yourself.
I recently used the command with some combination of airport/city/airline and the only match was the boarding pass I was looking for. It could probably be used for receipts from hotel or whatnot, as soon as pdftotext can retrieve the text. It should find tax returns and related PDFs by querying "IRS + SSN".
A current issue that I would like to fix is the preview window that does not always highlight the query in full if a single match was found before the full query was typed. It is linked to how fzf handles previewing. I do not have plans for any big enhancements.
edit: I created a public repo to replace the gist. Feel free to post your thoughts or suggestions in the issues!
http://www.lesbonscomptes.com/recoll/faqsandhowtos/IndexWebH...
I'm trying it out now -- I have way too many pdfs, and I could really use the extra features (like context previews).
However, to be fair, you can follow a somewhat similar workflow with Zotero in combination with the Firefox plugin: download pdf by adding it to the Zotero database in Firefox and Zotero takes care of the indexing. Zotero misses the fancy interactive fuzzy searching you have in your workflow thanks to fzf, but I've added it as a feature request for Zotero [1].
You don't have to organize your papers into folders (or collections in Zotero parlance since a single item can appear in multiple collections). For most academic papers the Zotero plugin will also grab the pdf's metadata as a bonus without additional costs.
I’ve enjoyed having access to the database from my laptop as well as iPhone and iPad. It’s definitely been a workflow I’ve cobbled together. This seems to be working out better though.
At that point I realised that my BibLaTeX file was really a mess. Better BibTeX created tons of needless double curly brackets {{like this}} in the BibLaTeX file, making searching it directly a pain. And it created lots of @misc entries, the BibLaTeX entry of last resort, when it should have made @url's and other things.
Zotero has massively benefited my work, but it's also been something of a training wheel which in the longer term slowed me down. Emacs/AucTeX/RefTeX does everything Zotero does (at least according to my use) but faster, cleaner, more holistically, and with more features (eg crossref'ing entries). And two years on I'm still cleaning up that Better BibTeX's BibLaTeX file.
[0] https://retorque.re/zotero-better-bibtex/
* edited because I confused AucTeX with RefTeX.
I just call up RefTeX by hitting ctrl-x + (or "C-x +" in emacs parlance). This calls up a list of all possible references and a fuzzy search box which lets me narrow it down. When I've found my reference I have ten options, including:
f1: Open associated pdf, url or doi
f3: Insert the citation into my thesis and then be asked for what kind of note I want to add (footnote? Title? author name? Year? Just a bibliography entry?), then any text to add before the reference (eg, "This argument has been made by,"), then after (p. 67). -- This answeres your question
f7: Attach pdf to email
f9: Show notes, if there's an .org file of the same name.
f10: Add pdf to library
If RefTeX can't find my query, not only does it ask if I want to add a new one, it asks whether to search various databases for it including arXiv, DBLP (computer science bibliography), Google Scholar, Bodleian Library, HAL, Library of Congress, British Library and others.
Adding entries to the database is as simple as editing a plain text file, and BibTeX provides quick shortcuts for all the different kinds of entries, checks them for you and provides a sane and customizable key.
Everything is completely integrated - my writing, my pdf library, my org notes, my bibliography database, my email as well as my thesaurus (wordnut), my pdf reader (pdf-tools) and my git repo (magit). It's just brilliant. All together it makes emacs the ultimate writers tool, as far as I'm concerned.
I love Emacs (not as much as you do it seems ;) but Zotero is amazing for ingesting academic references including fulltext and bibliographic data from the web. One click of a button in your browser, and everything is ready to read and cite.
(Source: Wrote quite a number of papers and co-authored one textbook, all using Zotero. My workflow was Zotero for ingestion -> per-publication bibtex files for authoring)
So either you add them manually or you query Crossref, arXiv, DBLP or HAL (French open archive) within emacs and then can copy their bibtex entry from there. No pdf or web reference importer as far as I'm aware. Yet!
I mean it is nice that Zotero has an import browser plugin (which, btw, Mendeley does as well), but once you have a substantial library of pdfs, it is just too time consuming to re-import it all again, type in names, years, journals etc. ugh..
Also, Mendeley can open pdfs within the app and has great features for mark-up and comments. Reviewing a paper is really nice this way. In fact, Mendeley is the program with the best mark-up features on all of Linux, in my opinon. And all this is synced across devices.
I'd love to use Zotero but tbh Mendeley is just better.
(I don't think it would work with papers too old to have DOIs but haven't checked lately.)
The metadata retrieval process is a very handy feature, especially paired with Zotero's web browser plugin.
Zotero also keeps a repository of all the citation styles, which are curated and administered by the Zotero team and stored in a Github repo [2].
It automatically added 90-95% of PDFs I imported, most directly from Sci-Hub with no editing from me.
What is the proper way to store PDFs that I annotate (say, using Okular in Linux) on Zotero with the ability to send them to other researchers and then update them?
I don't mind paying money for cloud storage, but I gotta be able to work with the pdfs.
Then you can annotate the file and save it. If Zotero creates a copy of the file when you save your annotations, you might need to use Show file by right clicking the article in Zotero and make changes to the file in Zotero's storage. In Linux, that'll be in ~/Zotero.
[1] https://forums.zotero.org/discussion/1977/changing-the-defau...
When I say a bit of a pain, I mean it - I switched to Mendeley with great sadness because its mobile application means I can read and annotate papers on my tablet and have it synchronised perfectly with my desktop.
On a semi-related note: the Mendeley mobile application makes selecting text a joy on a touchscreen with a kind of magnifying glass. I really wish it was built into Android system wide.
What kind of annotations are that? I think GP refers to document annotations, which are stored in the PDF file itself.
If they are stored in the file itself, that's news to me. Last time I used it, Okular didn't modify the actual PDF and considered that a feature.
Thanks!
It does that, and it can actually look up metadata from pdfs. But: It can NOT update already existing entries from such a lookup. You have to delete the pdf and re-import. Why? No reason.
It's this sort of really bad usability decision that makes Zotero just not very good imo.
Disclosure: Zotero developer
If not, with Zotero you have a chance of getting a missing feature added. Not so with the closed-source Mendeley. Especially now they have started doing user-hostile lock-in stuff like gratuitously encrypting users' data, making it no longer accessible to the user (see downthread).
With Zotero, in principle you could hack together something using `bash` and `sqlite3` that batch-updated documents; you may not even need to look at the source code to the official app. It's ad-hoc munging like that that Mendeley has recently gone out of its way to prevent, stopping people from working effectively with their own data.
Zotero can not update metadata for a PDF you already have in your library
Hence, adding an existing library means you have to do it manually for each pdf This is bad design, because the functionality exists, but it's implemented in a way that makes everything complicated for no reason at all.
Why disable the "lookup metadata" button for a pdf which is already in a library? WHYYYYY
So far, Zotero's metadata lookup seems about the same quality.
I'm slowly realising that my colleagues for whom this "works" are never going to change. Our shared folders are littered with "Copy of FINAL final + comments 2.7.18-my-copy.docx.docx". It doesn't matter how slick my git + markdown + pandoc workflow is when conversations go like this:
"Just use track changes."
"But..."
"JUST USE TRACK CHANGES."
Comments might be extractable, I'm not familiar with docx format but it's zipped X(?)ML type data so there will be parsers. Or a conversion to an intermediate format that's more amenable to computer processing, perhaps.
The problem remains in reverse, though: it's expected that I will produce Word docs full of track changes edits.
Zotero's cross-platform plugins meant I could use any device for reading and not worry that if I followed important links I'd never find them again.
The API is also great and I've been using it to automatically sync PDFs of papers I want to read to my reMarkable tablet.
I currently have an iPad Pro but I really don’t like reading PDFs on it. I’d love to stop printing out papers and an eink display would be nice.
EDIT: just read your review linked in the post
My two biggest annoyances are currently:
1) No web app, so the only way to export PDFs is via the desktop app which is only available for Windows and Mac. The Android app is also pretty terrible. I think not starting with a web interface was a bad decision.
2) The PDFs exported from the desktop app are huge (a PDF that started out ~400KB ended up >100MB after export). Support tells me they're working on this.
I abandoned Mendeley because it would not permit relative paths, and I have not looked back.
The one downside of this approach is that multiple people can't modify the same PDF at the same time. The upside is that you can use whatever PDF tools you want and annotations remain accessible in the file even if you stop using Zotero, which goes with our philosophy of leaving people in control of their own data. (Mendeley stores annotations and highlights in its own encrypted database, and you can't even export PDFs with annotations in batch. If you want to get a PDF with your annotations out of Mendeley, you have to do it one file at a time.)
Disclosure: Zotero developer
For LaTeX users I recommend the Better Bibtex plugin, making my life a lot easier!!
Mendeley has started encrypting users' bibliography files (from version 19), claiming variously that encryption is required by GDPR (?!?!) and/or that it improves security on a multiuser system. (Neither of these excuses holds water.) The keys are not available to the users whose data were encrypted. The encryption is completely proprietary and there are no tools available for letting users work with the now-encrypted databases.
Previously, users could access their data by running `sqlite3` on a plain sqlite database in their profile directory.
Not only did users take advantage of this, but a small ecosystem of third-party tools and scripts had sprung up to help people take control over automating repetitive tasks in managing their bibliographies.
Well, now Mendeley has encrypted everything, that's the end of that. No more tools and scripts. No more user control over one's own data and workflow.
There's now only the limited (for my purposes, useless) "export database" facility in the client. It only exports a limited subset of the fields in the database. Alternatively, I could register as a developer, get an API key, and develop a full web-style app to get hold of my own data. ("Did you just tell me to go fuck myself, Bob?")
The Zotero people suspect that Mendeley's move was retaliatory for Zotero's implementation of import-from-Mendeley [1]. The Mendeley twitter account has dismissed this - literally - as being "fake news" [2]. Super weird.
I reverse-engineered the encryption they'd put in, in order to export my own data and migrate to Zotero. I wrote up the instructions for others to follow [3], but it's really not easy, not portable, not reliable, and not something users should ever have to do just to get access to their own data!
Anyway, at this point, I don't trust Mendeley as far as I can throw them, and I thoroughly regret ever recommending to my friends that they use the tool. After nine years of using Mendeley, though, I've switched to Zotero, and it's at least as good, and in some ways better. Plus it's properly open.
[1] https://www.zotero.org/support/kb/mendeley_import
[2] https://twitter.com/mendeley_com/status/1006919608471818240 and several others. Really, the official twitter responses from Mendeley to people discussing this issue were bizarrely dismissive and mocking.
[3] https://eighty-twenty.org/2018/06/13/mendeley-encrypted-db
"The data subject shall have the right to receive the personal data concerning him or her, which he or she has provided to a controller, in a structured, commonly used and machine-readable format and have the right to transmit those data to another controller without hindrance from the controller to which the personal data have been provided [...]"
That said, if I had to collaborate in a larger group I'd surely choose Zotero instead of locking everybody into some proprietary ecosystem. But for personal use I definitely don't feel that I'm enabling or supporting Elsevier somehow in using their freeware.
If instead of using Elsevier's freeware you were to file bug reports - not code just reports - on Zotero then you could be helping improve the software.
And when you suddenly _do_ need to collaborate you'll find that you are using the same tool you have gotten used to, instead of the dissonance of using a tool that is similar-but-not-quite what you know.
Go ahead and use Zotero, check out the two or three leading Android clients, and file bugs. You've got these clients currently: https://play.google.com/store/apps/details?id=com.gimranov.z... https://play.google.com/store/apps/details?id=computer.benja... https://play.google.com/store/apps/details?id=net.ezbio.zote...
[0] https://play.google.com/store/apps/details?id=com.mendeley&h...
It'd be nice to have the database synced, though. Funny how "synchronization of replicas of information", for all its centrality to modern internet-based life, isn't something that our operating systems provide as a system service.
Just FYI on that front, macOS has iCloud, which does this. I put something on my Desktop, it appears on all my machines. I can access it from my phone, etc. etc. - it's basically a built-in Dropbox.
I subsequently tried contributing but got lost in the weeds of the somewhat unwieldy codebase. In part because it's interface is based on XUL and I don't know wtf is going on there. I think they're slowly migrating to either electron or something else. If you know XUL and have some free time, please consider contributing.
But since I am a LaTeX/LyX person, I also wrote an extension that lets Mac users automatically add newly saved citations to BibDesk: http://mackerron.com/zot2bib/
Is there a way to make it work with SciHub? Asking for a friend... ;).
https://www.reddit.com/r/scihub/comments/7ilwzv/scihub_look_...
https://github.com/bwiernik/zotero-tools/blob/master/engines...
- metadata extraction/automated adding really doesn't work right with my research workflow (a lot of google scholar searches, a lot of humanities and social science sources)---lots of inaccurate or incomplete info, lots of downloading RIS files and then manually importing and separately manually importing the PDF.
- documentation for things that would be useful like hooking up to academic library proxies is nonexistent. Take a look at the chain of empty links when you try to get proxy info: https://www.zotero.org/support/proxies
- no better bibtex for zotero 5... Although maybe this has changed recently? Which would be amazing.
I'm not sure exactly what you mean, but the primary way of adding items to Zotero is with the Zotero Connector browser extension, which lets you save high-quality metadata and PDFs with a single click from a huge variety of sources (certainly including humanities sources, since Zotero was created by historians). No other tool comes close to Zotero's abilities here. Metadata quality does vary by site, though — Google Scholar specifically only provides limited metadata, so you'll often get better results by clicking through to the linked article and saving from there. We have plans for functionality to flesh out incomplete metadata retrieved from subpar sources.
Zotero can also automatically retrieve metadata from PDFs you drag in, which should work for the vast majority of recent PDFs and many older ones with DOIs assigned, though that's not meant to be the primary workflow.
> documentation for things that would be useful like hooking up to academic library proxies is nonexistent
Current proxy documentation is here [1], and that's what's linked from the main documentation page. I've fixed the outdated page you pointed to — thanks.
Note, though, that the proxy functionality is meant to work automatically for the popular academic proxies, so most people don't need to configure anything to use it. (And as far as I know other competing tools don't offer anything like this.)
> no better bibtex for zotero 5
BBT has worked with Zotero 5 since last year.
Disclosure: Zotero developer
[1] https://www.zotero.org/support/connector_preferences#proxies...
Incidentally, one thing that would be really helpful for a documentation standpoint would be a series of articles about best practices, like which metadata sources work best.
On the whole, the transition was pretty effortless and I am pleased with Zotero (plus the Better BibTex plugin and webdav syncing). I like having the rss feeds inside the client. The firefox add-on is far superior to the Papers version. And finally, because it's an open source project with a reasonable ecosystem, many of the things I found annoying about Papers have been solved (like customizing which fields to exclude from a bibtex export).
Major differences seem to include cloud based Zotero's advantages for sharing, collaborating, and Citavi's better fine-grained document quote/cite tools.
https://www.quora.com/What-is-the-difference-between-Citavi-...
I've used Endnote, Mendeley, and Zotero before and found them all to have their own issues. Paperpile is not perfect but it shines in a few places and I really like it. The chrome plugin adds an 'add to paperpile' button to places like google scholar making it easy to add citations.
It also is designed for writing in Google Docs and makes it very easy and quick to add citations. For some, Google Docs may be a dealbreaker, but if you need to collaboratively write/edit a manuscript it's much better than the alternatives in my opinion (esp. because everyone can write/edit at the same time).
I've heard of other alternatives being better but shrug Zotero works well for us.
ETA: The firefox integration of Zotero is great, and something I missed from Mendeley. Also, I should say, before I switched to Zotero, I was a Mendeley user for about nine years, and did my PhD dissertation's bibliography using it. I'm quite confident Zotero would have been just as good.
That said, Zotero as a research tool for organizing papers by topic,etc,adding notes, is just fantastic. I really do love Zotero and it's potential.
However due to the literary suite paradigm I would recommend Jabref (which is integrated in Docear) instead [2], if you are coming from Zotero/Mendeley, which is what I Did. Docear alone is capable of much more. I was not aware it is not really active these days. It is really unfortunate.
[0] https://www.scss.tcd.ie/joeran.beel/blog/2014/01/15/comprehe... [1] https://gtpedrosa.github.io/blog/apresentando-o-docear/ [2] http://www.jabref.org/
[1]: http://git.27o.de/dataserver/about/Installation-Instructions...
They have a super simple interface to do this from firefox, but I also configured my qutebrowser via a userscript:
I'm slightly confused what that means, is that just meta data in the pdf file, that isn't visually visible?
For a while I've wanted to make something that can extract the title, authors, and bibliography visually from a pdf. Is that what zotero can do also?
For extracting metadata from a formatted bibliography you can use AnyStyle [2], which is a separate service written by a Zotero developer.
[1] https://github.com/zotero/recognizer-server [2] https://anystyle.io
Update: I wonder why this comment got downvoted any without explaination. HN is not same, as it used to be.
Because this is a freeware program that is community-driven and you're blaming the developers when in reality it is the community that is failing to provide translations; largely the Indian language community, I might add.
Still operating.
Could you use this like a digital "commonplace book"?
2 GB: $20/yr 6 GB: $60/yr Unlimited: $120/yr
I've had the paid 2 GB plan for a while, I've got about 800 MB stored with them, and I've been very happy with the service.
Between the Firefox connector, the automated PDF OCR metadata lookup, and the ability to pull in metadata from DOIs, Zotero makes indexing things insanely easy. It's like MusicBrainz for articles. I'm always surprised how many obscure PDFs it correctly recognizes and indexes for me.
In the options under the Sync tab, I select WebDAV as the File Syncing type and for the URL I put:
my.nextcloud.server.com/remote.php/webdav
with the appropriate credentials.Works flawlessly.
Zotero is the full package, GPL, the only premium service is extra data for syncing with their server (not necessary if you only have one client). None of this Orwellian hinting about keeping researchers from using libgen.
Reading about Elsevier now and I understand the objection. Thanks.
https://www.theguardian.com/science/2017/jun/27/profitable-b...
I'd rather just use JabRef http://www.jabref.org/