ArxivGPT: Chrome extension that summarizes arxived research papers using ChatGPT
github.com
github.com
Then again, perhaps LLMs could simply be incorporated into the peer-review process, where after submitting your paper, you'd have to answer the AI's basic questions. As a reviewer, I could imagine a structured AI report for a paper being helpful in guiding discussion: "The paper compares to recent approaches X, Y, and Z. And the work is indeed novel."
I'm a big believer of using all these new AI services/programs as tools to enhance my workflows, not replacing them.
The response was: The sound of cracking knuckles is actually a release of gas in the joint, but there is little scientific evidence to support the idea that it leads to arthritis or other joint problems. In fact, several studies have shown that there is no significant difference in the development of arthritis or other joint conditions between people who crack their knuckles and those who do not.
I wonder if OpenAI used that paper to fine tune ChatGPT
I don't think we know exactly what the improvement consists of.
That's only partially correct at best, because joint cracking is not a single phenomenon. For one, there is disagreement over whether the sound comes from the formation of the gas bubble, or the collapse of the gas bubble (cavitation). It is likely that both are partially true.
Furthermore, it doesn't explain the sort of joint cracking in which people can crack a joint repeatedly with no cooldown. I can crack my wrists and ankles by twisting my hands and feet around in circles, cracking once per rotation as fast as I can turn them around. 60 cracks a minute easily; neither of the gas bubble hypothesis can explain this, it is almost certainly a matter of ligaments or bones moving around against each other in a way that creates a snap sound, like snapping your fingers makes a crack sound by slapping your middle finger against the base of your thumb.
It’s literally monkeys with typewriters pressing keys randomly.
Until we get new models which have true understanding, they will never be truly useful.
If you actually follow the literature, you'll find that there is tons of evidence that the seemingly "simple" transformer architecture might actually work pretty similar to the way the human brain is believed to work.
It's also clear that RHLF models can be given instructions to reduce this issue. And in many production LLM models, something called few shot is used - in context learning - where you provide several examples and then ask for a new case. The accuracy is again improved this way, because the model "deduces" that you are being serious and humourless, when asking about mirrors breaking, and are not asking in the context of a story.
It's also one of the only datasets that fails to increase with scale (There was a big, cash-paying high incentive challenge to find other datasets, and it pretty much didn't find any that can't be worked around or are not experienced by RLHF models like ChatGPT). So it doesn't represent a wider trend of truthfulness above human accuracy being impossible (dated/data drift on the other hand, clearly remains an issue).
Or same question in a different way, what sort of workflow would be enhanced by an innacurate summary?
I would not suggest anyone to use ChatGPT outputs for actual knowledge at this point.
> "Discovering Latent Knowledge in Language Models Without Supervision" Existing techniques for training language models can be misaligned with the truth: if we train models with imitation learning, they may reproduce errors that humans make; if we train them to generate text that humans rate highly, they may output errors that human evaluators can't detect. We propose circumventing this issue by directly finding latent knowledge inside the internal activations of a language model in a purely unsupervised way.
https://arxiv.org/abs/2212.03827
In other words the model already tries to predict the truth because it is useful in next token prediction, but we need to find a way to detect the 'truth alignment' in its activations.
aws cloudfront update-distribution --id <distribution-id> --distribution-config <new-config> --no-reset-origin-access-identity
Took me a while to figure out that the parameter --no-reset-origin-access-identity was not only not working. But it did never exist on any version of the cli tool.
But for this problem (CF-Distribution lost OAC settings when updating the root file), all the google fu in the world did not help me. It turned out that I had to update my aws-cli and my problem went away. Apparently no one else on the internet had that problem, so only my gut could help me figure it out.
Meaning, if all the facts exist in the prompt then the likelihood of synthesizing fiction is diminished.
There are a number of ways to use the principle of analytic augmentation to add most or if not all of the facts required for a truthful response, ranging from simple “prompt engineering” to evaluating code to document embedding in latent space.
For example, if you use prompt engineering to k-shot a task to turn math word problems into executable JavaScript, meaning LLMs are only translators and the computations are done by a software interpreter, then the results are much more likely to be truthful.
Sampling from a number of variations on a prompt can lead to a more accurate outcome if say 1:10 times the translation attempt has a different answer.
This is a conflict of incentives. Whereas ArxivGPT has no reason not to tell the problems first.
This says more about you than the hypothetical "people" you are talking about.
Edit: [*] 'we' in my comment here is indicating the HN community, not entirety of humanity.
But on the other hand, this is what the guidelines say
> Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something.
I don't think these type of comments teach us anything. But these are guidelines and not rules for a reason.
Comments like the one at the root of this thread teach people that the current/approximate generation of LLM’s are poorly suited for certain kinds of tasks and remind us that many people don’t yet seem to understand what their systemic limitations are.
LLM’s are statistical models over their training data and aim responses mostly towards the most dense, data-rich, and redundant centers of their corpus. Summarizing novel, esoteric or or expert material is something they’re poorly suited for because it inherently has poor representation in that data.
Scoping is very constructive feedback.
My own opinion is that the cats out of the bag so whatever's going to happen is going to happen wrt LLMs. But trying to shut down all criticism of a new thing just because you think it's cool is itself not cool.
And those little tots sure do look cute spinning the 'saws right?
That seems useful, prudent, and completely in line with the spirit if a community like HN.
There’s no more reason that every critique should come with a “proposal” than that every cheer should come with some kind of admonition. As a community, multiple points of view are expressed and developed simultaneously.
Of course, some of points of view might personally frustrate you or leave you feeling like you don’t know how to respond to them. But is that so bad? Does it need to be squelched just because you don’t enjoy it?
But that's not how any of this is discussed. That's where my scope comment comes from, the scope of every project cannot be "for everyone and 100% safe from the beginning". In this way there is no encouragement to discuss/make better things, just discouragement. I, personally, hate this.
If I can't, I'm gonna have to skim the paper anyway. But even that could be pretty quick.
Quickly reading something to gauge relevance is something I can confidently say I do much faster than GPT can.
And I don't have a paper feed. I look up papers relevant to what I'm working on at the time.
"Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples. By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do. Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine- tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning. with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic. At the same time, we also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora. Finally, we find that GPT-3 can generate samples of news articles which human evaluators have difficulty distinguishing from articles written by humans. We discuss broader societal impacts of this finding and of GPT-3 in general."
The summary:
- If you train a computer on a lot of words, it can do things a lot better
- In this case, the computer has learned how to translate languages, answer questions, and unscramble words
- It still has trouble with some things
- This is very interesting because it shows computers can do things which are very hard.
So here, it drops a lot of information (cloze tasks), and it adds some (that last point). But now I know what the paper is about in 10 seconds.
I go back and see that oh, "a lot of words" really means 10x the previous. I'm now hooked on what problems it has trouble with, and what problems it solves for humans.
I didn't know a thing about LLMs when I first read this. If I tried to read it top to bottom, I'd get stuck on "task-agnostic, few-shot performance" then "state-of-the-art fine- tuning approaches" then "an autoregressive language model". They're big words, but turns out they're not the interesting parts, and understanding what the paper is excited about helped me to understand the basics.
2. What a useless summary! This summary is so dumbed down it could describe literally any paper on LLMs. This would give me zero information on whether the paper is worth further reading.
With this new class of products based on crafting prompts that best exploit a GPT's algorithm and training data, are we going to start seeing pull requests that tweak individual parts or words of the prompt. I'm also curious how the test suite for projects like this would look for specific facts or phrases to be contained in the responses for specific inputs.
Edit: btw, congratulations on the release. This is the kind of stuff I think should be explored more using LLMs. Great choice on making a chrome extension, it's great UI for this kind of thing.
[1] https://github.com/hunkimForks/chatgpt-arxiv-extension#how-t...
https://huggingface.co/ml6team/keyphrase-extraction-kbir-ins... is a decent tool to explore the constant stream of publications. The last mile still is left to the human.
Are you using your own API key and pay for the usage? How can you justify operation of programs that produce high costs but no income? Isn't the API publicly exposed to the client-side and possible subject of theft and abuse?
I think the extension is only useful for someone who spends a lot of time doing research on arxiv.org, (and if the quality of the summary is good enough, the jury's still out on that one)
Disclosure: I work at MSFT but not on Bing