GPT-4, without specialized training, beat a GPT-3.5 class model that cost $10M
threads.net
threads.net
We're not in the business of training models. We will never be as good as OpenAI / Anthropic etc.
Where the real value in applications is smarter prompting techniques and RAG. There is a lot of room at the bottom in doing "dumb" things and simply feeding models with the right context to deliver customer value.
Just last Friday I took the contents of the 2024 folder of one of the teams at the company I work for, for which we use RAG at the moment. I dumped the text index, concatenated it and used Google’s API to return the token count, to see if it would fit in Gemini’s 1M context window; turned out it was 5.7M tokens. And that’s less than 3 months worth of documents for that team.
So yeah RAG is not dead yet, although I do question its usefulness, but that’s a separate topic.
The claim that RAG is dead is obviously wrong.
The internet is full of blog posts about this. That doesn't mean they're actually good - I'd love to be pointed at one that has proven itself useful for someone (and definitely isn't just LLM blog-spam).
I don't care if it's trivial to fine-tune and get crap results - I care about fine-tuning where the result was worth the effort.
For the record, my favourite guide to fine-tuning is the section of this Jeremy Howard video that shows how to train a text-to-SQL model: https://www.youtube.com/watch?v=jkrNMKz9pWU&t=4850s
The person asked for citations, leave it be, stop dudexplaining how Internet works for you please.
I call bs on your 2 comments.
Even when people sincerely ask for a citation on a debatable topic, on an internet forum, it's effectively saying "I won't be hear any opinion that doesn't match my own unless it's as water tight as a law of physics". Another form of this is "show me the data".
Disclaimer: I work for Google.
Fine-tunes lead to catastrophic forgetting.
RAG is only irrelevant if you’re completely disinterested in cost and latency.
We also don’t have enough data to gauge performance of models >200k context window size when reasoning over inputs of that size, much of which will be irrelevant to any particular user. Multiple random needles in haystack tests work flawlessly, but rarely applies to real world activity.
LLMs are a commodity
You can run a dozen of these LoRAs atop the same base model on the same infrastructure for a dozen specific use cases.
The inference quality, performance and cost can all be substantially better than GPT4 with prompting.
You have to know when to RAG, finetune, or RAG+finetune.
Just curious about what the usecase is for a 7b model in a business context - ie. what does it do?
greatly outperform GPT4 *for* just a prompt
your overfitting to training data convinces no-one that you created a "better GPT4"I mostly work on AI, so I know if I'm overfitting or not. It performs provably better in it's domain (a niche programming language). GPT4 can barely write a hello world for it.
I'm not creating a "better GPT4" general chatbot. I'm finetuning for a specific task.
If you write about your experiments with that in detail I guarantee you'll get a lot of interest. The community is crying out for good, well documented, replicable examples of this kind of thing.
In short, an 8B model could degrade almost 10x after being finetuned on robotics tasks, while 500B model experiences a very minor degradation (~4%) and there's a hope that with a larger-sized model, it could become either zero or even negative (adding more experience improves general purpose reasoning).
What I am saying is that while GPT-4 could beat a finetuned GPT-3.5 class model, I predict good things about finetuned GPT-4 class models, when they become practical outside of OpenAI/Google.
The best single work of fiction ever created about LLMs' capabilities (and, perhaps, dangers) is Colossus by Jones. Although I think the film is even better than the book, only the latter mentions how, despite being created specifically for US national defense, Colossus is also fed unrelated data including Shakespeare's sonnets, because its creators do not know if it could be important.
ALOHA-2: https://aloha-2.github.io/
..and plan to do an updated version soon for much of what's been released since. I've also done work related to LLM and robotics integration, also on that site.
Happy to chat about it.
> Technology’s largest leaps occur when new tools are provided to those that want to make things.
I love this sentence. And the general attitude of curiosity of your post.
If you want a great overview of what a modern robotics stack would look like with all this, https://ok-robot.github.io/ was really good and will likely make it into the article. It's a VLA combined with existing RL methods to demonstrate multi-tasking robots, and serves as a great glimpes into what a lot of researchers are working on. You won't see these techniques in robots in industrial or commercial settings - we're still too new at this to be reliable or capable enough to deploy these on real tasks.
Finetuning was the best option for weaker LLMs with lower context windows (e.g. the original GPT-3): both problems have been solved nowadays.
The cost economics are much better with few-shot prompting to modern LLMs too: input tokens are super cheap (especially with the recently-released Claude Haiku), so giving a lot of examples per call will still end up cheaper than finetuning.
Meanwhile, a finetuned ChatGPT costs 4-6x of normal ChatGPT usage.
You are an expert copywriter. Write five distinct blog post titles optimized for high clickthrough for Hacker News for the article the user provides.
Your response must follow the style of these titles:
- The ü/ü Conundrum
- Why isn't preprint review being adopted?
- Majority of web apps could just run on a single server
- Weather Planning for Eclipse Day
- PSChess – A chess engine in PostScript
Then provide the blog post as the user message input.I just ran one of my blog posts (https://news.ycombinator.com/item?id=39495476) with the workflow through Claude Haiku and got this:
Here are five distinct blog post titles optimized for high clickthrough on Hacker News for the article provided:
1. Tipping ChatGPT: Does Offering Monetary Incentives Improve AI Text Generation?
2. Quantifying the Impact of Incentives on Large Language Model Performance
3. Carrot or Stick? Exploring the Effects of Positive and Negative Prompts on ChatGPT
4. Gamifying AI: Using "Generation Golf" to Test ChatGPT's Ability to Follow Length Constraints
5. The Curious Case of ChatGPT's Motivations: Can an AI Be Incentivized Like Humans?
Not bad titles, although more verbose than the 5 input examples I gave. I only gave 5 for simplicity: my main point is that you can give it a lot more than five and/or be more aggressive with constraints, like the blog post linked incidentially.intuitively, prompting like this to get an answer seems basically like the first part of a fine tuning process (more exemplars).
what is your thought here behind why reinforcing good output via a loss optimization is worse than the one shot example? does the model start to over fit at some point towards some local minima? and this is avoided in this scenario?
It can make sense if you have a particularly simple use case.
For that alone I commend him.
“I have 1000 flops of compute” - works.
“I have 1000 flops of computation” - doesn’t work.
“That compute failed” - doesn’t work.
“That computation failed” - works.
They’re different.
…but ‘computation’ is also uncountable, and your second sentence seems to be perfectly fine to me.
Your examples do not constitute an argument. You haven’t articulated the (purported) difference between the two words; you’ve just decided arbitrarily that some sentences don’t work, and not elaborated or explained at all.
I can make up words too, and provide example sentences: “karrotz are delicious” works. “carrots are delicious” doesn’t. “inside the karrotz” doesn’t work. “inside the carrots” does.
I don’t actually think there is any difference. The above comment about ‘brospeak’ was snarky but I do think it’s more of a cultural phenomenon than a semantic one — unless someone is willing to kindly explain the difference rather than just rolling their eyes!
What exactly is wrong with the sentence ‘this would require huge amounts of computation’? Saying ‘compute’ seems more to be a synonym of ‘computation’ that’s caught on recently than a useful gap-filling addition to the language. Again: reasoned arguments please. Or just ‘we think it sounds cool so we use it’ — that’s fine, too.
EDIT: pondering briefly, perhaps one could argue the difference is something like ‘you can own compute, but you can’t own computation.’ ‘Compute’ is the capacity to carry out computation. …although ‘compute’ seems to be used to refer to the ‘abstract’ computation being done as well as the computational resources, so I don’t know.
I’m stretching it. To be honest I’m not sure it’s a useful (or even real) distinction. I think it’s a matter of fashion, and that’s fine and normal.
https://web.archive.org/web/20240321091803/https://www.incom...
"All work and no play makes GPT a very dull AI"
Guess. Stop trying to shape the NN. And let it learn on its own.
The best single work of fiction ever created about LLMs' capabilities (and, perhaps, dangers) is Colossus by Jones. Although I think the film is even better than the book, only the latter mentions how, despite being created specifically for US national defense, Colossus is also fed unrelated data including Shakespeare's sonnets, because its creators do not know if it could be important.
But I believe that most of the data stored in foundation models are just useless for some particular domain. So it's better to forget something, getting really useful info instead.
What Bloomberg did for $10M was not finetuning..
That's a big claim - can you back that up with any examples?
And in practice, most tasks people are using GPT-4 for in production are more like the latter than the former.
(Disclaimer: building https://openpipe.ai, which makes it super easy to productize this workflow).
I saw some comments here say to check out Claude. From what I can tell, Claude hasn't figure out yet how to do the whole "generate Python code and run it in a Juptyer notebook" for math yet.
Correctly prompted, even Mistral-7B can write and run code in response to questions, and it's a model that can run on laptops from half a decade ago, with two or three orders of magnitude fewer parameters that GPT-4.
By default, the ChatGPT "model" knows to not try to do math and instead write code to do the math then run it. I get that it's set up infrastructure wise to be able to run it, but why is Claude's main chat UI not trying to instead respond
"hey, do this calculation on your own since I can't" or something of this nature instead of responding to math incorrectly
For example, I'm able to get Claude-3-Opus to write Python and call a Python interpreter of its own accord when questioned about time series data in some of my data analysis workflows, though I haven't glued together a pretty GUI for it yet (e.g., plots are simply saved to disk). While I haven't run into any problems around calculations yet, I'm sure it wouldn't be too hard to further refine the system prompt and ensure that all calculations are performed or checked using Python.
Although based on other tasks, overall, GPT-4 seems to be the best, but by a very small margin, so I cancelled my subscription. Although the native mobile app is really great.
so when looking at it that way, the real question is what do you need? all I need is Mixtral 7x8B Q5 in an 8,000 token context window, at the moment
I think there are plenty of other people that can design their applications and problems around lower fidelity tools, or just pursue something else
Debugging a Python function this morning. Claude 3 Opus failed completely. GPT-4 found the bug, as well as two others I hadn’t even been looking for.
As always, your results will vary based on your personal prompting style. My style apparently works great with Opus.
Here's one example: GPT-4 gave me code that was missing some async/await keywords: https://chat.openai.com/share/117fb1ad-6361-41e2-be59-110f32...
Claude 3 Opus with the same prompt got it right the first time: https://gist.github.com/simonw/2002e2b56a97053bd9302a34e0b83...
Once Opus has the ability to run a code interpreter, it'll really be an exciting time.
https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
The one mixed-bag weak spot I've found is in coding -- It tends to make more "d'oh" mistakes while coding, but comes up with more creative solutions at the same time ¯\_(ツ)_/¯
GPT-4 didn't figure that out, either; that’s just tooling built around the model, not something the model “figures out”.
A lot of models claim to be GPT3.5 class that clearly are not in the first place.
Not true. Most prompt techniques that work on current modern LLM models will work on different or future models, although it will require a QA pass for any regressions.
I think that a whole lot of what I do in prompt engineering is what's necessary to fully specify the output that I want.
A newer model may be less finicky, so I have a higher chance of getting it to work on the first try (and it's more reliable afterwards), but it's hard for me to imagine it needing a whole lot less prompt.
I don’t think RAG is going away, at least not because of this. But I expect new techniques to become available fairly regularly.
But not doing it is an opportunity cost. You don’t built skills, tooling and experience, and you don’t get feedback on what works and where you should go.
It’s like computers in the 1990s: there’s always a better one 6 months away, so if you wait for it to stabilise, then you don’t do anything for a decade. Just enjoy the ride, bearing in mind that things change very fast and some things will be obsolete next year.
On the other hand, GPT-4 actually did worse on the NER task - labelling and tagging terms used in the text - vs their finetuned model. I assume the finetuned model was better at using the specific labels they were targeting.
It is. https://www.adalovelaceinstitute.org/resource/foundation-mod...
I can't tell from the OpenAI docs whether it's possible to access GPT-4 without the ChatGPT fine-tuning. If so, that'd make this result more meaningful. Otherwise, I just don't think you can draw any great conclusions from this.