Claude 3 surpasses GPT-4 on Chatbot Arena for the first time
arstechnica.com
arstechnica.com
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make
./server -m models/7B/ggml-model.gguf -c 2048
I don't think it'll take you the whole weekend :)"Tomorrow" your desktop computer might be twice powerful but at the same time the "good model of tomorrow" will be four or ten times larger - I'd expect that the gap between what can be done locally versus what is offered as a service will grow, not shrink.
That doesn't mean you'll be able to run the best model, but I'm relatively optimistic about the gap not growing out of control.
Honestly the only thing that keeps me to sometimes prefer GPT-4 now is the UI. I like being able to Edit my messages, and to Stop the model if I gave it the wrong prompt. Please improve Claude's UI!
The interoperability between LLMs right now is amazing. When I write a program I can quickly test it with each of GPT, Claude and Gemini to see which work better for what I'm doing. Here's to hoping nobody figure out how to create a moat any time soon!
Each of them do some things better.
Edit: although on some level the training only gets you the general capabilities of the model. How you fine tune it to specifically be a useful bot is a very important element. That’s not really model complexity so much as design thinking and experimentation
It just seems to do a much worse job than pasting your code into the chat UI.
Like, it's answers are just profoundly bad in comparison.
I've noticed that in visual studio (IDE) copilot gives better answers if I physically view an interface or implementation, then I get "okay" results. But it stuggles with larger more abstract projects.
Vscode is better for sure, but usually I'm working in smaller projects or interpreted languages.
The vim copilot extension is probably the better bit again I'm not working with dotnet in vim.
You can't rely on IDE auto context since the entire codebase is too large to feed into LLM (maybe Claude 3 200k context token can take it, but too expensive). And RAG is not smart enough to figure which part of code is relevant.
It's not feeding in the entire codebase but whatever you selected or whatever files you have open, at least that's what the UI suggests. So I ask it "what does this line do" and I get an answer that uses the whole file to explain what the line does.
But if you use it to do something across multiple files (DB schema, service, controller, HTML, JavaScript), then it becomes less accurate (precision vs recall problem) as like you said it uses your open windows or some heuristics to decide what's the context.
With an IDE as an interface, it is just not intuitive, UX wise, to "open tabs" to signal to GitHub Copilot that the files should be included in the context.
Definitely not GPT-4, otherwise it would not be less than $10 a month for constant usage.
I figure if they do this, they have to throttle or nert it somehow since it is cheaper than ChatGPT Plus which also gives access to GPT-4.
Copilot Chat uses 4, but it's suspiciously free of confirmation that is also used in the more contextual Copilot (no-chat).
Their moat right now is developer tooling. They allow fine tuning + an easy API to use their llm.
Noone else does that right now. By the time others do, so much tooling and infrastructure has been built around OAI that the switching costs will be a lot.
It will get to the point that if you want your llm to beat oai on the market it won't be enough that you are as good or even better. You need to be vey very significantly better than OAI. For an extreme example of this see Windows. The network effects keeping it together are so strong that the platform becoming an abandoned adware hasn't been enough to push users to significantly better platforms like Mac.
Now I've fine tuned the hell out of gpt-3.5 and I'd love to see how my app would performed on a fine tuned Opus. I went to their website and I can't seem to fine tune their model yet. Meh. My guess is that by the time they make it available I won't have a strong reason to even try anymore.
I think you are correct, for chat. But for audio, video, 3D stuff, it will never be that easy for a newcomer.
I think Claude has better writing style, and it's refreshing not having to fight with the thing to get it to give me full code snippets. I also hate how difficult it is to get the arrow keys to scroll up or down on chatgpt.
I've found Claude 3 massively better than ChatGPT4 for my use cases.
ChatGPT4 will just get the question entirely wrong the majority of time on high levels of complexity.
For example, here's this 1k lines of code. Find this bug, and fix.
ChatGPT4 totally gets it wrong, Claude 3 gets it right, or at the very least finds the area to get it right.
This is repeatable over and over all day.
Edit: The funniest part of my workflow is when one LLM gets an idea wrong, I then put it to the other LLM.
It's a great way to get unstuck. Both LLMs get stuck on certain things, but they are usually different things.
Claude3 opus (and often Sonnet) actually fills in the full dependency paths, actually makes tests, and just overall seems to "know what I want from it".
(To their credit, they count anything where the 95% confidence intervals overlap as a tie)
> Vote won't be counted if model identity is revealed during conversation.
(Also, though, free MS Copilot only allows off-peak use of GPT-4/GPT-4-Turbo; paid MS Copilot is requires for peak hours use of those models.)