Edit: After looking at the examples on the front page "What are the benefits of taking l-theanine?" this seems geared for the general public, so maybe it wasn't the right test.
Edit: After looking at the examples on the front page "What are the benefits of taking l-theanine?" this seems geared for the general public, so maybe it wasn't the right test.
When that works for me, I am probably weak on the subject material myself. eg writing quirky love poems to my wife in different styles.
For research tasks, because the AI is not deeply self-reflective, it can output inconsistent and incoherent results. What it does is present text that only *looks* as if it confidently knows what it is talking about.
For domains where high rational quality doesn’t matter like love poetry, it is amazing. For other domains, be wary. If you can’t tell the difference between what is actually good and what merely looks good superficially you will be in trouble.
“ I am probably weak on the subject material myself. eg writing quirky love poems to my wife in different styles.”
“ For domains where high rational quality doesn’t matter like love poetry, it is amazing.”
So, self-described weak at love poetry, but confident that it is a domain that LLMs excel at. That is an interesting take. Perhaps the LLM is just as weak at liberal arts as it is hard science, but it is just more difficult to measure since you aren’t in the domain. Most poetry I’ve seen from LLMs has been pretty rote and boring although as you say, not a “rational quality” I suppose.
If you don't read a lot of poetry, what LLMs output look like poems, but almost always lack wit, a through-line, coherence and poignancy. It usually contains the individual parts, but never fitting as a whole.
Instrumental vs terminal values, I guess. Makes me think of coding - the overlap between good code, and code that makes money, is nearly empty.
The end result of good code is to do a thing. That might result in money, but the instrumental value is in the thing it does rather than for the joy of coding.
Many people will assume that artists create things strictly for the terminal value of just doing the thing because the prose doesn’t “do” anything so the artist must have just enjoyed making it. But the artist usually wants to make an impact - communicate an idea, change a mind, etc. Not just “make a thing that looks like a poem so that I can sell it.”
I say this is an important frame to look at the problem in because there is pretty much zero instrumental value in generating LLM output beyond the kind of gatcha-style fun of putting in words and seeing something pop out the other end and the terminal value is always measured in money or “time efficiency” because what else is there to measure in?
Did I actually craft a love poem that communicates my true feelings in prose with all of the little flaws and personalizations that only we know in one of our most intimate relationships? Did I choose my voice or was it someone else’s? What is the value I’m getting when I fish something out of a generative AI? Did I really get the same value fishing with prompts through and AI’s output than making the thing? Maybe, maybe not.
There's probably a term for that which I'm forgetting, so let's provisionally call it tangential mediocrity - when the work is mediocre by general standards, but quite good for purpose it was made for.
You might not have given the "right test" in terms of the actual userbase, but it is absolutely the right test in terms of Elicit's marketing claims. Elicit might be implicitly geared for the general public, but they are explicitly marketing to scientists.
I suspect a lot of Elicit's target users want to use scientific knowledge in their personal/professional lives, but without doing the hard work of gaining scientific understanding. However, they're not going to spend money on a product that says "we use AI to create sciencey bullshit that sounds plausible in conversation." They want a product that Real Scientists would use. (Similar to how purely decorative Damascus steel Bowie knives are gussied up by an outdoorsman pretending to use the knife to gut a fish or whatever.)
It sounds kind of like toy marketing: want to sell a toy to 5 year olds? Show 7 or 8 year olds playing with it, even if they'd never actually choose the toy in real life.
We now know though that the perchlorate detection may have been an instrument error and satellite imagery constraints water content of the RSL to below what would be expected from brine flows. It's not conclusive though and there is no consensus whether the RSL are caused by liquid or dry processes or some combination of both.
[1] https://meetingorganizer.copernicus.org/EPSC2015/EPSC2015-83...