Llama 2 is about as factually accurate as GPT-4 for summaries and is 30X cheaper
anyscale.com
anyscale.com
> We used a 3-way verified hand-labeled set of 373 news report statements and presented one correct and one incorrect summary of each. Each LLM had to decide which statement was the factually correct summary.
The problem with this approach is its assumption that if an LLM can recognize an accurate summary, it'll be able to reliably produce accurate summaries. We know very little about the inner workings of LLMs right now, and what we do know suggests that they work highly counterintuitively, so I think there's no basis to make this assumption.
I think that's really significant. It's arguably a case for fine-tuning models for specific tasks. That's great for teams with ML experience. But for product engineering teams without ML engineers, they can just use a foundational model and get great performance for a low cost.
We've shared the code openly, you can reproduce it yourself if you want.
Time and time again we see benchmarks like this where 0 effort into actually trying to get anything good out of the model.
Of course you can make extremely specific changes that essentially game the test, so to me a fair way to test this is to establish a realistic token budget across all the models, maybe even adjusted for cost depending on the goals of your testing. Then for each model, use your budget to get the highest quality result possible.
_
You're asking the GPT models for example, in the worst possible way. You've moved detailed instructions into the user prompt, despite their newest updates focusing on steerability via the system. Meanwhile from tinkering LLAMA 2 pretty much doesn't care about the system prompt vs user prompts.
You're also not allowing for any form of chain-of-thought. Giving the models a few hundred tokens to form a conclusion instead of spending those tokens trying to force out a mathematically induced order bias (which I wouldn't expect to do much) would have been much better.
There's also no way that GPT 4 should have struggled to give you a well defined output: All of the models would have probably benefited from a well defined output format in terms of a schema, rather than asking for a single letter response. The model rarely has to produce a single letter answer and will struggle with that. Something like asking for { "output" : "A_IS_MORE_FACTUAL" | "B_IS_MORE_FACTUAL" } would have been better
Overall I think there's no open dishonesty, and obviously there's no objective correct amount of effort to put into the test setup here... but there's a conflict of interest that would have made me want to see more effort put into it. I think there were trivially low hanging fruit ignored here that I'd expect people selling LLMOps to have pick up on.
If someone wants to try it, they can, but I think it’s based on some unwarranted assumptions that LLM’s are likely to have balanced performance, just because there are some that do seem to be pretty balanced?
I've gotten great performance from llama2 derivatives. Out of the box performance is not near GPT4 but it is still very strong in its own right. And, if you are able to break down your problem so precise logit control coupled with guidance from forward or backwards chaining on knowledge graphs is applicable, you can easily exceed gpt4's reasoning ability for your domain. No fine-tuning necessary either.
I've been getting useful things out of LLMs since the days of roberta and raw T5, when Large stood for hundreds of millions of parameters. I am flabbergasted when people say a 7B parameter model is no good for them.
What do you mean by this?
https://github.com/guidance-ai/guidance
I'm going to butcher this explanation - after you've generated your selection of logits but before you sample from them, you check which ones conform to your schema. If you want the only two options to be "true" or "false", then you take any of the logits that would provide invalid answers and lower their probabilities manually.
Another example is structures like JSON can be validated so when your sample is "{'name':'Carl'" you lower the probability of "{" since that would invalidate the json. In fact the only valid ones you'd likely have left would be ",", " ", and "}"
I've actually tested Llama 2 for summarization but haven't blogged about that yet, and across multiple domains, Llama 2 is pretty good.
I do see some differences -- Llama 2 doesn't follow instructions as cleanly, and is more verbose.
It's not a plugin replacement -- in the article itself I point it out. GPT 3.5 largely followed instructions to return A or B. Llama 2 didn't so I had to use another LLM to post-process.
I'm not saying you should always use Llama 2 or ChatGPT. It's that for some use cases, you can save a lot of money by using an open source alternative.
We dedicated an entire team of 6 for two months on evaluating LLMs and while Claude 2 was the final choice, we found Llama 2 70b to be absolutely great (for summaries and structured data generation from free text).
We chose Claude because running Llama on an A100 or H100 comes with a baseline cost that doesn’t go to zero when you don’t need it (you could spawn a new instance but right now GPUs are so rare everywhere except for expensive cloud providers that it’s possible you don’t get one).
That said, we found the smaller Llama models to be so hilariously bad we have an internal slack channel where Llama 7b writes jokes and the funny part isn’t the jokes but how utterly stupid and random they are.
Stretching aside, how does one follow from the other?
this is the age (or year) of token price arbitrage
For OpenAI - that is very interesting though. The 700M cap means the big FAANG+ companies that Meta competes with - Google, Apple, Amazon, Netflix, even Microsoft - can't use Llama 2.
But OpenAI has less than 700 million users? So theoretically, if Llama was better than GPT models, OpenAI could replace their GPT engine with Llama :)
Their original compute platform for running arbitrary ml workloads will become obsolete as the industry consolidates around LLMs.
1. a model for sentiment analysis
2. a model for summarization
3. ...other NLP tasks
other tasks:
4. a model for object detection in images
5. a model for face recognition in images
Whereas now the LLM does all of the above better than the previous state of the art. This will continue and eat more fields of machine learning. It will happen for images and video. I argue that it will even extend to things like time series analysis
imo the industry will consolidate around LLMs for early prototypes and as part of the workflow for building a corpus to train domain-specific models.
Source: team went through this process. LLM cost ~1% unskilled human labelers, then we finetuned BERT to bring the cost down to < 0.001% (savings $120k/y compared to just using the LLM)
The tech consolidator works by: if it reasonably can, it will.
PS5? Gaming machine, 4k disc player, streaming box, etc.
Smartphone? Internet, phone, texting, camera/photos, storage, shopping, identity, gaming, music, video, directions. Soon we'll add highly useful AI bots to the list.
Windows? Pretty much non-stop consolidated software from the industry into itself, across decades.
Also known as eating the ecosystem. GPT & Co. will eat their best plugins.
The point is that a model can do everything, not that it is the most cost effective or best way to do everything.
There are tradeoffs with all tech, and there always will be.
That's for sure. The main problem is how to put them together. There are several ways. First, tokenizing the media and inserting it in the stream. Second, a media controller which accepts verbal instructions generated by LLM. They can be combined. Interesting variation is controller with (immediate) feedback. It can be anything, for example database access.
> I argue that it will even extend to things like time series analysis
AFAIK transformers have been used for low level robotic control. And LLM for high level have been reported by MS and Google.
Ray (https://github.com/ray-project/ray) is available for anyone to use. You can also use Aviary (https://github.com/ray-project/aviary) to serve any of those models yourself.
Our prices are competitive (starting at 25c per million tokens) because of the tech we've built that maximizes GPU utilization. That's what we do better than anyone else.
I still think GPT-4 is best LLM out there. It's just very expensive and you don't need all that horsepower you can save bucketloads of money.
OpenAI had a first mover advantage. ChatGPT has a good interface and GPT4 still leads on many benchmarks. But it continues to get worse and worse and others are catching up with significantly lower parameter counts. Their moat is shrinking very quickly.
Llama2 is an open source meta model, btw. Anyscale's ray makes it much easier to leverage the super fast pace of development in the OSS community (or released by large companies as OSS).
FWIW GPT-4 is not one single model as you're aware and it evolves over time. In our experience it has objectively degraded in quality over time.
As for anyscale, ultimately it's more like databricks (and funny enough lots of ex databricks folks there). Anyone can just run hosted spark. Databrick's ecosystem has set them apart and it looks like that's what anyscale is building towards.
With GPT4 being OpenAI's year ago tech.