Elicit – AI Research Assistant
elicit.com
elicit.com
We're glad you're enjoying it.
Edit: After looking at the examples on the front page "What are the benefits of taking l-theanine?" this seems geared for the general public, so maybe it wasn't the right test.
When that works for me, I am probably weak on the subject material myself. eg writing quirky love poems to my wife in different styles.
For research tasks, because the AI is not deeply self-reflective, it can output inconsistent and incoherent results. What it does is present text that only *looks* as if it confidently knows what it is talking about.
For domains where high rational quality doesn’t matter like love poetry, it is amazing. For other domains, be wary. If you can’t tell the difference between what is actually good and what merely looks good superficially you will be in trouble.
“ I am probably weak on the subject material myself. eg writing quirky love poems to my wife in different styles.”
“ For domains where high rational quality doesn’t matter like love poetry, it is amazing.”
So, self-described weak at love poetry, but confident that it is a domain that LLMs excel at. That is an interesting take. Perhaps the LLM is just as weak at liberal arts as it is hard science, but it is just more difficult to measure since you aren’t in the domain. Most poetry I’ve seen from LLMs has been pretty rote and boring although as you say, not a “rational quality” I suppose.
If you don't read a lot of poetry, what LLMs output look like poems, but almost always lack wit, a through-line, coherence and poignancy. It usually contains the individual parts, but never fitting as a whole.
Instrumental vs terminal values, I guess. Makes me think of coding - the overlap between good code, and code that makes money, is nearly empty.
The end result of good code is to do a thing. That might result in money, but the instrumental value is in the thing it does rather than for the joy of coding.
Many people will assume that artists create things strictly for the terminal value of just doing the thing because the prose doesn’t “do” anything so the artist must have just enjoyed making it. But the artist usually wants to make an impact - communicate an idea, change a mind, etc. Not just “make a thing that looks like a poem so that I can sell it.”
I say this is an important frame to look at the problem in because there is pretty much zero instrumental value in generating LLM output beyond the kind of gatcha-style fun of putting in words and seeing something pop out the other end and the terminal value is always measured in money or “time efficiency” because what else is there to measure in?
Did I actually craft a love poem that communicates my true feelings in prose with all of the little flaws and personalizations that only we know in one of our most intimate relationships? Did I choose my voice or was it someone else’s? What is the value I’m getting when I fish something out of a generative AI? Did I really get the same value fishing with prompts through and AI’s output than making the thing? Maybe, maybe not.
There's probably a term for that which I'm forgetting, so let's provisionally call it tangential mediocrity - when the work is mediocre by general standards, but quite good for purpose it was made for.
We now know though that the perchlorate detection may have been an instrument error and satellite imagery constraints water content of the RSL to below what would be expected from brine flows. It's not conclusive though and there is no consensus whether the RSL are caused by liquid or dry processes or some combination of both.
[1] https://meetingorganizer.copernicus.org/EPSC2015/EPSC2015-83...
You might not have given the "right test" in terms of the actual userbase, but it is absolutely the right test in terms of Elicit's marketing claims. Elicit might be implicitly geared for the general public, but they are explicitly marketing to scientists.
I suspect a lot of Elicit's target users want to use scientific knowledge in their personal/professional lives, but without doing the hard work of gaining scientific understanding. However, they're not going to spend money on a product that says "we use AI to create sciencey bullshit that sounds plausible in conversation." They want a product that Real Scientists would use. (Similar to how purely decorative Damascus steel Bowie knives are gussied up by an outdoorsman pretending to use the knife to gut a fish or whatever.)
It sounds kind of like toy marketing: want to sell a toy to 5 year olds? Show 7 or 8 year olds playing with it, even if they'd never actually choose the toy in real life.
How did you get all those blue chip orgs give you permission to use their name? None of our blue chips clients allow us to do so.
Do startups just roll with it and use the logos without permission?
I would start with Hanlon’s Razor, with a 10% chance of malice.
Too bad the internal paperqa system at scihouse isn't available for public use...
edit: various typos
Oh no
That's basically the same as the percentage of people who read news stories when responding to or sharing the headline
This link: https://insights.rkconnect.com/5-roles-of-the-headline-and-w... says "Only 22% only read the headline of an online news story, according to data from the Rueters Institue for the Study of Journalism."
But following the link to that study gets me to https://reutersinstitute.politics.ox.ac.uk/sites/default/fil... which… if it supports that claim, I can't seem to find where :P
Tbere is also https://www.researchgate.net/publication/323202394_Opinion_M...
and there was yet another one but I can't find it
Accuracy and supportedness of the claims made in Elicit are two of the most central things we focus on—it's a shame it didn't work as well as we'd like in this case.
I'd appreciate knowing more about the specifics so we can understand and improve
Actual quote from the abstract: “ No tryptamines were detected in the basidiospores, and only psilocin was present at 0.47 wt.% in the mycelium.”
It does not differentiate between psilocin and psilocybin, those are two different molecules.
A 90% accuracy rate seems like the sweet spot between "an annoying waste of time" for honest researchers and "good enough to publish" for dishonest careerists.
I don't like disparaging the technology experts who work on these things. But as a business matter, 1/10 answers being wrong just is not good enough for a whole lot of people.
In my experience it takes quite a bit longer to falsify GPT-4's incorrect answers than it does to a Google search and get the right answer. It might take 30 seconds to check a correct answer (jump to the relevant paragraph and check), but 30 minutes to determine where an incorrect answer actually went wrong (you have to read the whole paper in close detail, and maybe even relevant citations). More specifically, it is somewhat quick to falsify something if it is directly contradicted by the text. It is much harder to falsify unsupported generalizations or summaries.
As a specific example, I recently asked GPT for information on arithmetic abilities in amphibians. It made up a study - that was easy to check - but it also made up a bunch of results without citing specific studies. That was not easy to check[1]: each paragraph of text GPT generated needed to be cross-checked with Google Scholar to try and find a relevant paper. It turned out that everything GPT said, over 1000 words of output, was contradicted by actual published research. But I had to read three papers to figure that out. I would have been much better off with Google Scholar. But I am concerned that a large minority of cynical, lazy people will say "90% is good enough, I don't want to read all these papers and nobody's gonna check the citations anyway" and further drag down the reliability of published research.
[1] This was a test of GPT. If I were actually using it for work, obviously I would have stopped at the fake citation.
Now? I’ve seen people argue positions that are demonstrably wildly wrong in unusually creative and often subtle ways and there’s no way to figure out where they went off the rails. Since the LLM is responsive, they can use it to come up plausibly sounding nonsense to answer any criticisms collapsing the debate into a black hole of bullshit.
I'm not sure I agree that those rule-of-thumb statistics are "arbitrary" or "fictional"… I guess it depends on what you mean by that. I can say that on our part they're a good faith attempt to help users calibrate how best to use the tool, using evaluations of Elicit based on real usage.
Definitely accept that the tool can work better or worse depending on your domain or workflow though!
One way we do try to distinguish ourselves from vanilla LLMs is that we provide sources for all of the claims made. I mention this because we hope our users can approach the falsification process you mention for Google. We want to show people where particular claims come from such that we earn their trust.
Walking citation trails and verifying transitive claims is something we've talked about but need more people to implement! (https://elicit.com/careers)
Sorry for the confusion: I meant that fragmede's comment was arbitrary and fictional, not the 90% figure. I was talking about these numbers:
if it takes 1 hour to get one answer by hand, but only 20 minutes for the machine, and 20 minutes to check the answer, the user still comes out aheadIF the machine actually got it right.
1) spend 10 hours doing all of them by hand
2) spend 3h 20 waiting for the machine, 3h 20 checking the machine, and 1h replacing the machine's mistake with a hand-written version, for a total of 7h 40
(I never trust marketing claims, so I doubt 90% accuracy; but also it generally takes LLMs a few seconds rather than tens of minutes to produce an output to be checked).
I think it's tempting but oversimple to focus on "output" and "time saved generating it," but that misses all the other stuff that happens while doing something, especially when it's a "softer" task (vs. say, mechanical calculation). It also seems like a mindset focused on selling an application rather than doing a better job.
I don't think that captures what I'm thinking about, which is more skills atrophy ("Children of the Magenta") problem than a conflict of interest problem.
https://www.computer.org/csdl/magazine/sp/2015/05/msp2015050...
> William Langewiesche's article analyzing the June 2009 crash of Air France flight 447 comes to this conclusion: “We are locked into a spiral in which poor human performance begets automation, which worsens human performance, which begets increasing automation” (www.vanityfair.com/news/business/2014/10/air-france-flight-447-crash).
> ...
> Langewiesche's rewording of these laws is that “the effect of automation is to reduce the cockpit workload when the workload is low and to increase it when the workload is high” and that “once you put pilots on automation, their manual abilities degrade and their flight-path awareness is dulled: flying becomes a monitoring task, an abstraction on a screen, a mind-numbing wait for the next hotel.”
I still think it's a concern at least as old as writing, given what Socrates is reported to have said about writing — that it meant we never learned to properly remember, and it was an illusion of understanding rather than actual understanding.
(That isn't a "no", by the way; merely that the concern isn't new).
This is a good point! (Hopefully) obviously, if we knew a particular claim was fishy, we wouldn't make it in the app in the first place.
However, we do do a couple of things which go towards addressing your concern:
1. We can be more or less confident in the answers we're giving in the app, and if that confidence dips below a threshold we mark that particular cell in the results table with a red warning icon which encourages caution and user verification. This confidence level isn't perfectly calibrated, of course, but we are trying to engender a healthy, active, wariness in our users so that they don't take Elicit results as gospel. 2. We provide sources for all of the claims made in the app. You can see these by clicking on any cell in the results table. We encourage users to check—or at least spot-check—the results which they are updating on. This verification is generally much faster than doing the generation of the answer in the first place.
I assume it is difficult for Elicit to give specific numbers because they lack the data, and confabulations are highly dependent on what research area you are asking about. So the "rule of thumb" is a way of flattening this complexity into a usage guideline.
I think it's a little early to bring AI to research field which need enough accuracy and rigorism.
What I'm really excited about is a tool like Elicit using the new Google Gemini 1.5 Pro/Ultra models with the 2 Million token window sizes, filtering down the papers using traditional search and high quality meta-data, then critically, prompt/activation caching to make the tool economically viable.
Maybe it won't work better, but I'm willing to bet it'll find those really specific ideas/needles in the haystack a lot more often than vanilla RAG will.
The basic problem is that scaling up understanding over a large dataset requires scaling the application of an LLM and tokens are expensive.
Our main focus is a little different to SciSummary actually. We're focussed on understanding researchers broader workflows, and providing a research assistant (i.e. rather than a particular narrow tool for summarisation or search).
The workflows we're most excited about at the moment are literature and systematic reviews: we think we can make these orders of magnitude faster and higher quality.
transitive verb
1: to call forth or draw out (something, such as information or a response) her remarks elicited cheers
2: to draw forth or bring out (something latent or potential) hypnotism elicited his hidden fears
We do use LLMs, but the secret sauce is an approach we call Factored Cognition which we wrote about here: https://ought.org/research/factored-cognition
(Elicit the company and app was spun out from Ought the research lab).
We do joke internally about the homophone (in fact, IIRC we did a little joke on our CEO by rebranding for his birthday in 2022) but I'm sorry to report that we're all careful, ethical, and well-behaved people :(
No interest in such a product of any kind.
For example, every patent is 50 or 100 pages of what amounts to bland, unreadable background. The stuff that is actually new in the world is usually a tiny fraction of all the words.
In a world where everyone's AI can generate all those extra words, their value goes down. (Yes, this is bad news for people in my line of work who get paid to write all those words.) So maybe the profession moves in a direction that values conciseness and brevity, to re-add value to what the lawyers and agents do, that the AI can't do.
It's a fantasy. But the same kind of thing applies to academic publishing. Maybe the future of publishing is very short, useful documents with rich, tiered, AI-generated hyperlinks on every word.
I hope I'm wrong. Eventually, when only AI produces and consumes text, it could be more brief, but at that point will it matter? Eventually a less ambiguous format could be used, something more like code, that computers and AI can consume and apply.
I don't feel that jargon, legalese and newspeak are going to vanish. On the contrary, the value of these words may actually increase due to the appeal that they possess by way of gesturing toward authority, virtue and the like. There is a likelihood that AI will only contribute to their proliferation, just in the same terse manner that you are describing. Densely packed compositions of inarticulate gobbledygook that serves none, yet is served by many.
We are on the cusp of a fissure in the field of knowledge work led by AI tools. I'm inclined toward the more manual labor that favors a dextrous approach to "information overload" but there are opportunities to leverage these tools as a sort of valve for copious amounts of information as well. The technology isn't going to go anywhere, so hopefully this is Elicit is a useful product that will serve the interest of the conscientious.
I don't understand that. If people bother to generate those words when they're expensive to write, why would they stop when they become cheap?
My contention (that I see some people don't like!) is that they don't have much value, regardless of who writes them.
My hope is that AI will cause the patent legal community to undergo a paradigm shift and place value on conciseness and brevity.
In the US, this would require Congressional and judicial buy-in. LOL!!!
Notice that I used the word "fantasy" above.
It does not matter whether a human or an AI produces a high-signal synthesis from a lower-signal body of work. The signal boost is value creation.
The underlying debate is whether AI is capable of signal-boosting arbitrary works more effectively than humans. I'd say presently, it is not. Even with the most powerful models, consistency and reliability across work types and lengths are far from business-grade.
If that were to change, the fact that an AI boosts the signal does not globally deflate the value of text. The value of text is determined by the signal, not the signal producer.
The main issue right now is that AI is not reliably boosting information signals, and therefore, most serious professionals probably still prefer to read original works.