LLaMA2 Chat 70B outperformed ChatGPT
tatsu-lab.github.io
tatsu-lab.github.io
I'm a staunch believer that it would be foolish to rely on GPT4 for quality comparisons, and it has been mind boggling to see so many people do it and treat it as perfect proof of anything.
It would be slightly more understandable if there was a study to see how human and gpt4 preferences compare, but I'm unaware of any such thing.
I wound up abandoning the experiment because getting enough reliability would have required screwing down the standards so tightly that it would have ruined the underlying point of the thing.
I've done a lot of work in audio synthesis, which is notoriously difficult measure. The gold-standard is human ratings of audio quality, but it is tough to design good tests (easy to fatigue raters) and the iteration time waiting for results is quite long.
Instead, there's now some projects which use neural networks trained on human ratings to predict audio quality, such as ViSQoL: https://github.com/google/visqol
This opens up fast iteration - scores going up generally corresponds to higher quality - followed by human testing at major milestones (eg, releasing a paper/model). VISQOL has a harder time comparing 'unrelated' models, IMO - ends up being not so great for comparison of different techniques, but excellent for measuring incremental improvement or catching regressions.
But, in the end, yes - you can use NN's to measure the quality of other NN's, so long as you're careful about it and make use of human raters from time to time as well.
The problem of test data getting into the training data seems to be an especially pernicious issue with LLM's, which isn't really arising in the audio synthesis space.
AudioLM puts a language model on top of the compression tokens, and thus can generate speech or other audio. There's piles of recent papers pushing that approach into music generation. Mulan is a name that comes to mind.
Or maybe your interested in audio separation to get at the isolated instruments? There's lots of great work on that, as well. Like MixIT, which is an unsupervised audio separation system.
I even read a classic book the other day and had a great discussion with ChatGPT about moral relativism, the different schools of thought and how it fit into philosophy as a whole. For students this technology is incredible, I wish I had it for all my classes.
Even sometimes comments I'll make on here or Reddit I'll pass through ChatGPT first to see if I made any mistakes in my logic.
There is no such thing as an average person.
https://www.thestar.com/news/insight/when-u-s-air-force-disc...
For whatever reason*, it is particularly bad at discussing philosophy I find. When I was grading philosophy 101, I would have probably given it a passing grade against the overall curve, but that's about it. Philosophy is a discipline of careful, sometimes jargoney, and always very couched assertions that can be easily misunderstood and appropriated. This is probably its greatest weakness, and in many ways this weakness is the progenitor of philosophy itself in the Western world, with Plato at the start (i.e. with the figure of the sophist, the paradox of a false wisdom).
- Maybe one reason: there is a huge amount of, lets say, "armchair philosophy" on the internet, compared to other disciplines. Many blogposts and tiny manifestos of people really excited by some out of context quote from Spinoza or whatever. And you start to really feel this part of the dataset when you ask about philosophy. Many strange takes and misunderstandings.
Philosophy
Ethics/Moral Philosophy
Meta-Ethics: The study of moral thought, language, and properties
Moral Realism: Belief that there are objective moral facts
Moral Anti-Realism: Denial of the existence of objective moral facts
Moral Relativism: The belief that moral judgments are true or false only relative to some particular standpoint
I thought that was pretty good, what do you think?I don't doubt it can do, like, Wikipedia type classification ok, but that's not like really getting to the substance of anything! And, either way, its not like there is one decided-upon hierarchy of concepts like this people consciously work within. This is a fine picture to some, but others might contest, perhaps, that Meta-Ethics is the "study of moral thought, language, and properties." What is "moral language" anyway? Why is it meta relative to Moral Philosophy writ-large? Or perhaps one might argue that we need to think of meta-ethics as a sibling rather than child. The whole discipline is a mess of different thoughts and possible rebuttals and grand intellectual overturnings that will not be captured here. Maybe just try pasting that back into the prompt and asking "what's wrong with this picture?".
But like I said, its fine in that its fairly comparable to Wikipedia for utility, (with IMO a worse interface, but I get why people like it more).
After all, LLMs and ChatGPT in particular are are indistinguishable from productised Gell-Mann Amnesia.
Edit: rewrote to be more neutral, sorry
Keep in mind that due to the nature of the data and the RLHF training, it's more like a weighted average, something like
0.5*(average American view) + 0.4*(average WEIRD view) + 0.1*(average human view).
(where WEIRD = Western, Educated, Industrialized, Rich, and Democratic, standard terminology in psychological research.)This may or may not matter depending on the questions you're asking, just something to keep in mind.
The world has long been divided into two camps: People who think computers can make mistakes; and people who think computers never make mistakes, and blame the humans that program them.
Well, now the computers are programming themselves. And clearly they're making mistakes.
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
*FreeWilly2 is a Llama2 70B model finetuned on an Orca style Dataset
EDIT: actually, impressive:
FreeWilly2 GPT-3.5 GPT-4
ARC 71.1 85.2 96.3
HellaSwag 86.4 85.5 95.3
MMLU 68.8 70.0 86.4
TruthfulQA 59.4 47.0 59.0
So reasoning (ARC) is lagging behind, but the other evaluations are at GPT-3.5 level and closing the gap with 4.Source for GPT-3.5 and GPT-4.0 values (but mind it might not be the same # of shots)
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
It beats out ChatGPT in every category except SAT-Math. We definitely need harder benchmarks.
So far, there's BIG-Bench Hard https://github.com/suzgunmirac/BIG-Bench-Hard and just published, Advanced Reasoning Benchmark https://arb.duckai.org/
So first step in the pipeline is cleaning up the transcripts for incorrectly transcribed words or sentences. Using the context of the rest of the transcript, it is able to fix the vast majority of them. Then we add punctuation and format it with paragraphs. Then I have another check over the whole transcript for any remaining issues.
After all of that, I have a relatively clean transcript that represents the original audio very closely. From there, I am doing things like: 1. creating question/answer pairs from the transcript 2. creating a document of additional context that fills in details about what the speaker is talking about but may not have explicitly said 3. creating a summary of the transcript identifying the main purpose 4. creating a knowledge graph from the transcript with nodes and edges 5. creating an annotated version of the transcript using that knowledge graph
I plan on putting some of this data into a vector database, and some of it will be used for fine tuning LLaMA2 models on specific tasks (like knowledge graph creation, annotation using a knowledge graph, and writing using a knowledge graph to keep track of events)
My only experience with transcripts is in the context of transcribing short interviews. I used Whisper and it was pretty good. I mostly work with quantitative data, though.
In terms of the disambiguation of speakers, I haven't done it, but I remember blind signal separation discussed in a signal processing seminar I attended. There is also this paper, in case you haven't seen it already: https://enk100.github.io/speaker_separation/
Thanks again!
Also, I haven't tried using Whisper for getting a transcription from the audio. I went the route of downloading the automatically generated transcripts from Youtube for a set of videos. An audio processing pipeline is definitely something I could add later though as an additional input channel for the overall pipeline.
Here is an example: https://gist.github.com/Tostino/f6f19e88e39176452c1a765cb7c2...
Here is the transcript that I created that knowledge graph from, and then annotated with the knowledge graph for training purposes: https://gist.github.com/Tostino/e64524437848fbb3aebe52056df8...
Edit: I am using symbolic IDs intentionally. Reason for that was this paper: https://ai.googleblog.com/2023/07/symbol-tuning-improves-in-...
The ranking I linked and quoted in my comment's is much better. See the About tab. It has 4 evaluations and it doesn't use GPT-4 to evaluate.
Also the top one is a tuned Llama 2. Also clarified in my original comment.
Disclaimer from the site:
> Caution: GPT-4 may favor models with longer outputs and/or those that were fine-tuned on GPT-4 outputs.
> While AlpacaEval provides a useful comparison of model capabilities in following instructions, it is not a comprehensive or gold-standard evaluation of model abilities. For one, as detailed in the AlpacaFarm paper, the auto annotator winrates are correlated with length.
Which is not close, because performance is logarithmic in training compute. Each additional percentage point of performance requires exponentially greater investment in compute during pretraining. Llama 2 was pretrained on 2 trillion tokens -- a significant investment in compute, for sure, but still not enough to get close to GPT-4.
Surprisingly, LLaMa 2 won 5-0 for me.
At least in my examples, the llama output was more verbose/comprehensive. Sometimes ChatGPT didn't expand enough, sometimes Llama missed the mark entirely (eg explaining the Eiffel's architecture.)
They both gave great answers overall though.
That is, each additional percentage point of performance requires exponentially greater investment in compute during pretraining.
Llama 2 was pretrained on 2 trillion tokens -- a significant investment in compute, for sure, but still not enough to get close to GPT-4.
And this is only one benchmark.
Paul Gauthier made this benchmark [1] to measure correct git diffs. If you ask GPT-4 for help with your code, it can output a change in a git diff more reliably than 3.5.
My hope is that we can do that with Llama 2.
It would be awesome to have all this running on a laptop in a completely offline mode.
But if you really want a portable offline thing, sure.
I have a whole host of personal pdf's and documentation that I would love to be able to ask questions about.
https://www.sematic.dev/blog/tuning-and-testing-llama-2-flan...
It's the most straightforward explanation I've found so far. I'd love to hear if anyone's found something better though.
Not only possible but quite easy. Inference for 70B can be done with llama.cpp using CPU only, on any commodity hardware with >64GB of RAM
And if it doesn't you need to do some workarounds with compiling and it gets a bit harder to run.
But, well, I am currently not hyped enough about it to actually try.
Lets just hope that there won’t be any embarrassing vulnerabilities coming out of this when someone could prompt the model to reveal its own environment variables or API keys or the internal prompt that it is using.
But it seems the $0 free AI models are eating OpenAI’s lunch and Meta so far is winning the race to zero.
While Llama2 is an improvement over LLaMA v1, it's still nowhere near even the best open models (currently, sans test contamination, WizardCoder-15B, a StarCoder fine tune is at top). It's really not a competition atm though, ChatGPT-4 wipes the floor for coding atm.
This was the HumanEval contamination one dev measured: ``` replit_glaive: 56.71% replit: 7.32% wizard: 4.88% ```
From the WizardCoder paper https://arxiv.org/pdf/2306.08568.pdf you can see that it hits SOTA (for open models) in not just HumanEval and HumanEval+, but also MBPP and DS-1000 as well, so it's not a one off.
For those interested in reading more about various considerations for coding models, I highly recommend reading the MSR phi-1 paper: https://arxiv.org/pdf/2306.11644.pdf
Looking forward to if they ever publish code/model/dataset since it has extremely strong performance trained on a very small number of tokens very manageable 1.3B and 350M parameter models.
and how do you know all these benchmarks not leaked? I think they all scrapped from web sites, the same as training data for LLM, so risk of contamination is extremely high.
The best way to measure this is through synthetic datasets, which generate new tasks every time and model can't memorize them during training. One example is BigBench has multiple such tasks, but researchers usually(always) not regenerating those datasets.
*edit: oops, my brain inserted "by" in the middle of "outperformed chatgpt". I'll leave my wrong comment up as a testament to shame.
Not to mention GPT4 at 95% and ChatGPT at 89% - I use chatgpt(3.5turbo)/gpt4 daily for work, and I rarely ever bother with 3.5turbo because of how unreliable it's answers are compared to gpt4.
So whatever this is effectively measuring is useless for comparing these models, especially across work types.
(But maybe it's a good filter: if someone is talking about "ChatGPT's" performance, they probably don't have anything useful to say.)
GPT-4 level models that regular people can run with a reasonable hardware budget are going to require innovations in optimization and model efficiency beyond just quantizing weights. Rumor has it that GPT-4 is a "committee" of ~220G models, which would require ~128GiB VRAM at 4-bit quantization to run each model.
It looks like the eval is open sourced so you could easily build a version w/ your own questions for blind testing...
You can easily opt out of the data sharing.
and without handing a whole bunch of data to a 3rd party and hope they're securing it properly
Also there is nothing about that attack that makes it iinherently only applicable to self hosted models.
> without paying API fees or relying on an unstable dependency that's constantly being tweaked
I see no evidence that this part will be possible with OpenAI. Usage fees will always be a thing because that's how they make money, and based on what I've heard from people who have actually tried to build on their APIs, I would not trust them to keep the model stable. There are always new safety features they need to add, and those changes break things.
I’d imagine OpenAI might run experiments on ChatGPT that they wouldn’t on the API, to avoid breaking 3rd party applications unannounced.
For example, if I were to ask how to do something with burp it will just answer instead of going into the "as an AI" monologue.