> if you are interested, we are looking to work closely with more researchers, to make sure it saves time for you in particular on a daily basis.
Sure, I can sign up. But can you clarify data retention and usage policies? As researchers we're always worried about being scooped. Despite complaining about the system and incentives in other comments, we're bound to the game we're forced to play to achieve scientific progress (if it weren't for this stupid game I'd research completely in the open). I can't find information on how data is used. This is especially important if you're discussing institutions like work, as many of them can't even use standard GPT4 accounts do to privacy concerns.
I'll provide some comments about what I personally would like but clearly personal preferences. I'll say that there's a lot that can be done in this space that doesn't even need AI.
For literature review:
Let's say I am a ML researcher (true) who focuses on 3D semantic image segmentation (not true) or some other appropriate niche (we should be able to get much more narrow btw). How do we deal with this literature search? As a simple example, let's say I wanted to get introduced to a topic that my colleague is working on. Finding a proper survey paper is often a surprisingly cumbersome task. Making this searchable (especially with a ranking not just by citations by some coverage metric -- too many 10 page "surveys") would be exceptionally useful. In addition, I'd love to have graphs that I can see paper or author connections through different sorting methods. A timeline is often highly useful to understand the progress of a research topic and especially helpful to making new surveys and onboarding new grad students (or anyone into a new topic).
Citation and Reading List:
I love Zotero, but it is so fucking limited it gets to me. Here's my issues:
There's no good way for me to create a prioritized reading list (which should be topic based). Sorting papers that come across my desk is a cumbersome task and they just don't get sorted. But I am able to assign prioritization values to these works (especially attached with a certain project or domain). But Zotero does not make this a realistically viable task. Realistically I have multiple browser windows and store arxiv links in tabs or use Obsidion (more cumbersome). This is actually something ML could help with, at least for the automatic classification (we can discuss other metrics too if you want).
Discovery is a challenge, especially in a new niche, but the bigger problem is actually wrangling the papers I've already discovered.
What also sucks is that I can't markup these works. Preferably on my iPad (OSX would be fine since easy to share). Adding notes is also a bit cumbersome. I often want to add notes to papers to categorize them as citation references to certain topics or place them into importance for certain niches, and this should be different from a reading list. And then comes the problem of sharing. Not great.
I said there's a lot that can be done without ML, and that is because I don't even have good tools to sort and catalogue the papers that come via natural discovery. This comes through multiple devices too, and idk if I'm going to get that on my computer or phone (twitter is still a great place for paper discovery, and comes with author comments/summaries. No tool catalogs this).
Another great pain is citations. Authors aren't helpful here, but grabbing the proper journal/conference citation is surprisingly a non-trivial task. They aren't always indexed by Semantic Scholar or Google, who both have quite large delays. But there's definitely no automatic correction in Zotero for when I add a arxiv paper (via plugin) that replaces the arxiv bib with the proper bib, even if that is indexed by GS/SS. Every time I write a paper this task has to be done and my zotero list gets updated.
So kinda in short, I don't want these fancy ML tools to summarize and write papers for me (I don't trust them for that, especially as an ML researcher. I'm also afraid it might only exacerbate our existing issues with publication, especially in the ML shitshow[0] (side note: useful tool could be finding similar papers and potential similar works. Especially flagging for plagiarism. The community does need this. Especially the blatant kind)). I more need classification models for tagging/indexing and a nice UX that allows me to prioritize reading, sort importance values on a highly niche domain, and dealing with document markup and note taking. It is crazy that good tools don't exist in this area. I'd pay for such a service, even a not great one, especially if it allowed for storage (which is easy on your end. Plus you can lessen the load by central storage to prevent dupes across users as well as automatic version tracking (arxiv)). Even better if you integrated a latex editor. Hell, the only reason I use (pay for) overleaf is because at some point I need to do pair editing, and I pay because I need the version tracking on specifically pair editing (sub-authors make many mistakes, but do need to comment), otherwise I'd far rather use vim and github. (idk an "experimental" editor that is just vim + Zathura would be great. If you hooked in multiple user cursors you'd solve this problem for me. See CoVim, vim-bookmarks, etc. The super feature would be allowing pdf markup with like an ipad onto a working latex document. Jesus fuck, I want this so much!). The experience of vim+github(+zathura) far more pleasant and I can write faster (add an emacs and sublime option and you got all your CS researchers hooked).
Note: I'm 30 year old ABD, not an old tenured professor. While I'd like fancier tools (though most aren't that "fancy"), some aspects also have to be usable by my advisor or boss who isn't as willing to learn new tools. So shifting between advance and simple is a weird challenge, but this is also why things like Zotero can be so bad, because momentum (and no other challengers).
I won't even get started on experiment tracking. That's a whole other rant. But both with experiment tracking and literature wrangling, you'll notice that everyone has their own system. That demonstrates a market need. I highly suggest going out and talking to a lot of researchers and trying to read between the lines of their workflows and their needs. Idk that we have them well defined because we're in a fucking rat race and jumping from task to task with little time to contemplate on these things (I'm sure many of these people would tell you they've wanted to design a system but never had time. I'm included and I doubt far from uncommon)
[0] If you really want to discuss how to solve this problem, we can discuss my anti-{journal,conference} dream and how to create a realistic peer review system. tldr answer is allowing comments on works (linked with profile) and providing links for replication attempts (should be incentivized, even with imaginary points). There is no good central repository for literature discussion, so keep in mind that this can be a great way to build a large userbase quickly. Tons of discord groups and stuff for this kind of stuff but it's a disorganized ad hoc mess.
P.S. How many people on the team have a PhD? I ask not for prestige value but rather as a loose gauge on how close the team's experience is to doing deep academic research. You're not going to experience much of this as an undergrad or even masters. Industry research will have similar and different problems (I've done both). But I've seen too many projects try to guess what we need without having experienced our issues. If you don't have the personal experience it means you need closer collaboration to ensure market fit or you'll misalign (already hard not too with how noisy this environment is).