I think that's really significant. It's arguably a case for fine-tuning models for specific tasks. That's great for teams with ML experience. But for product engineering teams without ML engineers, they can just use a foundational model and get great performance for a low cost.
We've shared the code openly, you can reproduce it yourself if you want.
Time and time again we see benchmarks like this where 0 effort into actually trying to get anything good out of the model.
Of course you can make extremely specific changes that essentially game the test, so to me a fair way to test this is to establish a realistic token budget across all the models, maybe even adjusted for cost depending on the goals of your testing. Then for each model, use your budget to get the highest quality result possible.
_
You're asking the GPT models for example, in the worst possible way. You've moved detailed instructions into the user prompt, despite their newest updates focusing on steerability via the system. Meanwhile from tinkering LLAMA 2 pretty much doesn't care about the system prompt vs user prompts.
You're also not allowing for any form of chain-of-thought. Giving the models a few hundred tokens to form a conclusion instead of spending those tokens trying to force out a mathematically induced order bias (which I wouldn't expect to do much) would have been much better.
There's also no way that GPT 4 should have struggled to give you a well defined output: All of the models would have probably benefited from a well defined output format in terms of a schema, rather than asking for a single letter response. The model rarely has to produce a single letter answer and will struggle with that. Something like asking for { "output" : "A_IS_MORE_FACTUAL" | "B_IS_MORE_FACTUAL" } would have been better
Overall I think there's no open dishonesty, and obviously there's no objective correct amount of effort to put into the test setup here... but there's a conflict of interest that would have made me want to see more effort put into it. I think there were trivially low hanging fruit ignored here that I'd expect people selling LLMOps to have pick up on.
If someone wants to try it, they can, but I think it’s based on some unwarranted assumptions that LLM’s are likely to have balanced performance, just because there are some that do seem to be pretty balanced?
I've gotten great performance from llama2 derivatives. Out of the box performance is not near GPT4 but it is still very strong in its own right. And, if you are able to break down your problem so precise logit control coupled with guidance from forward or backwards chaining on knowledge graphs is applicable, you can easily exceed gpt4's reasoning ability for your domain. No fine-tuning necessary either.
I've been getting useful things out of LLMs since the days of roberta and raw T5, when Large stood for hundreds of millions of parameters. I am flabbergasted when people say a 7B parameter model is no good for them.
What do you mean by this?
https://github.com/guidance-ai/guidance
I'm going to butcher this explanation - after you've generated your selection of logits but before you sample from them, you check which ones conform to your schema. If you want the only two options to be "true" or "false", then you take any of the logits that would provide invalid answers and lower their probabilities manually.
Another example is structures like JSON can be validated so when your sample is "{'name':'Carl'" you lower the probability of "{" since that would invalidate the json. In fact the only valid ones you'd likely have left would be ",", " ", and "}"
I've actually tested Llama 2 for summarization but haven't blogged about that yet, and across multiple domains, Llama 2 is pretty good.
I do see some differences -- Llama 2 doesn't follow instructions as cleanly, and is more verbose.
It's not a plugin replacement -- in the article itself I point it out. GPT 3.5 largely followed instructions to return A or B. Llama 2 didn't so I had to use another LLM to post-process.
I'm not saying you should always use Llama 2 or ChatGPT. It's that for some use cases, you can save a lot of money by using an open source alternative.
We dedicated an entire team of 6 for two months on evaluating LLMs and while Claude 2 was the final choice, we found Llama 2 70b to be absolutely great (for summaries and structured data generation from free text).
We chose Claude because running Llama on an A100 or H100 comes with a baseline cost that doesn’t go to zero when you don’t need it (you could spawn a new instance but right now GPUs are so rare everywhere except for expensive cloud providers that it’s possible you don’t get one).
That said, we found the smaller Llama models to be so hilariously bad we have an internal slack channel where Llama 7b writes jokes and the funny part isn’t the jokes but how utterly stupid and random they are.