What's the difference from sending the system prompt in the api call, as usual?
Edit: Oh, missed that: "We’re working on connecting Agents to tools and data sources."
154 karma · joined April 13, 2023
What's the difference from sending the system prompt in the api call, as usual?
Edit: Oh, missed that: "We’re working on connecting Agents to tools and data sources."
But, as per the first graphic, CriticGPT alone has better comprehensiveness than CriticGPT+Human? Is that right?
Wait, they can do that? Assuming weights have copyright, shouldn't the finetuning be a modification of the original work and so have the same license?
Specifically, the author is investigating (possible) changes in the system prompt and tools available to the model in the chat interface of ChatGPT Plus. That tells nothing about the model (GPT-4).
What should you do when someone, in a serious discussion, says that the Earth is flat, with a straight face?
Are they mocking you? Perhaps mocking the debate at hand? Are they trolling? Or maybe they are just naive and don't know that they are embarrassing themselves?
Maybe just treat them as a troll? But the thing is that when they appear just as some naive, maybe young person, the generous take would be to explain the ridiculous thing they are saying.. but I always feel so fool when I do that. Plus, I don't have the patience anymore.
I think it's asking for the old yubikey, but neither the site nor firefox give me other option or a link about what to do.
Also, maybe A-GPL could be a good license here. It adds a provision that if the user accesses the code remotely (as on a server), you should share the code too. The default GPL only requires that if you distribute the binary.
PS. not a lawyer, would be happy to be corrected if something I said was wrong
Seems that running docker in an old android repurposed as a server is still too much niche, unfortunately.
Bonus if it can run KVM with hardware support, but I don't know if it's possible yet.
That's no small deal, but in the grand scheme of things that a hot superconductor can give us.. I mean, this can (possibly, with decades of research) give us fusion, quantum computing, etc.
I hope that AI hype lead us to more memory and more memory bandwidth, because they are really lagging behind computer power increase from like 15 years already.
So we are talking about 4x your numbers per specialist model:
180GB * 4 = 720GB. If you count the greater context, let's say 750GB.
Anyone remember how many specialists they are supposedly using for each request?
If it's 2, we are talking about 1.5TB of processed weights for each generated token. With 4, it's 3TB/token.
At 0.06 for 1k tokens we get
3TB*1k/0.06 = 50 petabytes of processed data per dollar.
Doesn't seems so expensive now.
Exactly! Almost every weak point that Knuth commented is fixed in GPT4 answers.
Maybe OP feed Knuth's observations to the model?
If that ins't the case, I'm really impressed.
Actually, there's some experimental evidence that GPT4 have a Theory of Mind as good as humans, maybe better.
https://arxiv.org/abs/2304.11490
> GPT-4 performed best in zero-shot settings, reaching nearly 80% ToM accuracy, but still fell short of the 87% human accuracy on the test set. However, when supplied with prompts for in-context learning, all RLHF-trained LLMs exceeded 80% ToM accuracy, with GPT-4 reaching 100%.
Anyone know why this can't be a silver bullet for most tumors? Like, a way to target an intrinsic characteristic of cancers that in theory could be very effective on most of them.. sounds too good to be true. So, what's the catch?
Transformers perform qualitatively better than other architectures, and GPT2 (the most advanced public model at the time) shows near 100% accuracy. The best correlate of performance in the experiment is the next-word prediction accuracy of the model. Other AI performance metrics don't appear significant.
The conclusion is that this is strong evidence that the brain processes language using the same predictive algorithm as transformers. And GPT2 may have an architecture very similar to the language processing areas of the brain.
That said, I don't think the questioning of GP was malicious, just a natural curiosity. Yes, a little suspicious, but, well, we are in the internet after all. In the least, it's good to point when someone does the extra work to make a great presentation.
Anyway, great work riter!
I'm speculating, but I think that the few who refused to serve were met with the same outrage that someone like Weinstein receives today. It's no surprise that most eighteen-year-old boys could not do that, even if they were strongly against the war.
And getting hid of the NC clause of the original llamas too, of course.
As of right now, there's trouble replicating the eval results of the paper, for example.
It's like generating code in a language that you know nothing about. You should check for bugs, but you can't.
There are multiple options for ever one of the three programs os this setup, but it takes some work to make all run well together. I can check the exact binaries that I'm using or send the scripts, if there's someone interested.
Edit: to make it work well enough with wifi, I had to make some tricks to deal with the tablet's wifi powersave behavior. On usb it works perfectly though.
Some points addressed in the paper:
- Is the model emulating the reasoning of the 2-shot examples?
> The Davinci-3 and GPT-4 models experienced increases in ToM performance from all of the classes of CoT examples that we tested: Photo examples, Non-ToM Inferential examples, and ToM examples. The mean accuracy increases for each model and each type of CoT example are shown in Figure 4, while the accuracy changes for individual ToM questions are shown in Figure S.1. Prompting with Inferential and Photo examples boosted the models’ performance on ToM scenarios even though these in-context examples did not follow the same reasoning pattern as the ToM scenarios. Therefore, our analysis suggests that the benefit of prompting for boosting ToM performance is not due to merely overfitting to the specific set of reasoning steps shown in the CoT examples. Instead, the CoT examples appear to invoke a mode of output that involves step-by-step reasoning, which improves the accuracy across a range of tasks.
- Is the test data included in the training?
> The LLMs may have seen some ToM or Photo scenarios during their training phase, but data leakage is unlikely to affect our findings. First, our findings concern the change in performance arising from prompting, and the specific prompts used to obtain this performance change were novel materials generated for this study. Second, if the model performance relied solely on prior exposure to the training data, there should be little difference between zero-shot Photo and ToM performance (Figure 2), as these materials were published in the same documents; however, the zero-shot performance patterns were very different across Photo and ToM scenarios. Third, the LLM performance improvements arose when the models elaborated their reasoning step-by-step, and this elaborated reasoning was not part of the training data. Therefore, although some data leakage is possible, it is unlikely to affect our conclusions concerning the benefits of prompting.
Other highlights:
- With the prompting techniques 3.5-turbo got human level performance.
- In zero-shot GPT4 is already near human performance (90% of the human score)
- In zero-shot regime 3.5-turbo is worst than davinci-3, but much much better than it with prompting. This happens because turbo is too much cautious by default and often refuses to draw conclusions.