Claude 3.5 Sonnet
thezvi.substack.com
thezvi.substack.com
I spent some time yesterday experimenting with Projects, and, like Artifacts, it looks really useful. I like the idea of being able to have multiple projects going simultaneously, each with its own reference materials. I don’t need to use it in a team, but I can see how that could be useful, too.
The one problem I see is that the total context window for each project might start to seem too small pretty quickly. I assume, though, that Anthropic’s context windows will be getting larger as time goes on.
I wonder what other features Anthropic has in the works for Claude. My personal wish is for a voice interface, something like what OpenAI announced in May but has now put off until later this year.
Or did I miss something unique there?
The fact that Claude 3.5 Sonnet also seems smarter than the other current flagship models makes the Projects feature that much more attractive.
Maybe if I were smarter I wouldn’t find much use for Projects.
Also, you may be underestimating how buggy Swift and the rest of Apple's stack are. It's hard to get those bugs resolved unless you happen to work at Apple. Thus, a lot of time is spent working around bugs up the stack. So I don't find it surprising that a company moving fast like OpenAI ships _some_ bugs. The mac app just came out this month? Give it time.
It's odd that things like that are what they're missing.
As an experiment, I produced a set of bindings to Anthropic's API pair-programming with Claude. The project is of pretty good quality, and includes advanced features like streaming and type-safe definitions of tools. More than 95% of the code and docs was written by Claude, under close direction from me. The project is here:
https://github.com/cortesi/misanthropy
And I've shared part of the conversation that produced it in a video here:
For example, I’m an experienced backend programmer but have been using Claude 3.5 Sonnet and GPT-4 asking questions about a frontend I’m building in TS using Svelte (which i am not very proficient in). The LLMs frequently confuse themselves with TS/JS, server/client side approaches, recommend old and deprecated approaches, and mixing patterns from other frameworks (e.g. react) when an idiomatic approach does exist. The biggest risk is when, in my ignorance, I do not detect when this is happening.
It’s been functional enough to push a hacky prototype out (where it would take me probably months longer to learn and do this otherwise), but the code quality and organization of the project is certainly pretty low.
What I'm seeing with Sonnet 3.5 is a night-and-day step up in consistency. The responses don't seem to be that different in capability of opus / 4o when they respond well, it just does it with rock-solid consistency. That sounds a bit dull, but it's a huge step forward for me and I suspect for others.
PS. Consistency is everthing sometimes.-
Considering cancelling my subscription with OpenAI as I was previously using GPT-4 quite heavily as a multiplier for myself, guiding it and editing outputs as required, but GPT-4o feels significantly worse for this use case. It is certainly better in many other areas, but its coding ability is not great.
I tried to revert back to standard GPT-4 but it is now so slow to respond (higher load?) that it breaks my mental flow, so I'm exploring other options.
I wasn't worried about how this would affect our industry a few months ago, but this has me reconsidering. It's like a junior engineer that can do most tasks in seconds for a couple of cents.
Coding can be similar to playing an instrument, if you have mastery, it can help you be more expressive with the ideas you already have and lead you to new ones.
Whereas if we take away the craft of coding I think you end up with the type of code academic labs produce: something that purely starts on a “drawing board”, is given to the grad student/intern/LLM to make work, and while it will prove the concept it won’t scale into long term, as the intern doesn’t know when to spend an extra 30 minutes in a function so that it may be more flexible down the road.
I see this sentiment a lot regarding gen AI. An I get it, we need to learn our tools. But this seems like it's saying the only way to learn problem solving is the way you learned it. That's just not true. Everyone learns problem solving differently and the emerging field of gen AI will figure out it's own way. It's a different way of thinking. I see my niece using ChatGPT to make projects I wouldn't have even imagined taking up at her age. Her games work. Who am I to say she isn't learning problem solving? In hindi we say "pratyaksh ko praman ki kya avashyakta" (what's right in front of you doesn't require proof).
I broke my hand 2 months ago and Claude 3.5 sonnet has been writing all my code for me. It's been awesome
Currently you are apparently paying for:
- Everything in Free - Use Claude 3 Opus and Haiku - Higher usage limits versus Free - Create Projects to work with Claude around a set of docs, code, or files - Priority bandwidth and availability - Early access to new features
But what are the usage limits? Higher than free by how much?
Having an invisible limit on a paid product really rubs me the wrong way. Maybe some rate-limiting after a certain amount would be better than a hard cutoff, but even then I'd like to know what the limit is before I pay, not when I accidentally hit it in the middle of something important.
I’ve been paying for GPT since 3.5 debuted and I know what I’m getting - full, unlimited use of the best model. Period.
Anthropic needs to figure out what the hell they are selling.
FWIW I regularly hit my ChatGPT Plus limits, and I think the “dynamic” limiting is regularly in place. I’ve only once hit my Claude Pro limit. I now use Claude more than ChatGPT.
From this page:
https://help.openai.com/en/articles/6950777-what-is-chatgpt-...
As of May 13th 2024, Plus users will be able to send 80 messages every 3 hours on GPT-4o. and 40 messages every 3 hours on GPT-4. The message cap for a user in a ChatGPT Team workspace is approximately twice that of ChatGPT Plus.
Please note that unused messages do not accumulate (i.e. if you wait 6 hours, you will not have 80 messages available to use for the next 3 hours on GPT-4).
In certain cases for Plus users, we may dynamically adjust the message limit based on available capacity in order to prioritize making GPT-4 accessible to the widest number of people.
[0] https://support.anthropic.com/en/articles/8324991-about-clau...
Seems to also be limited by tokens. It’s still quite obscure.
> Please note that these limits may vary depending on Claude’s current capacity.
Fine for the free tier of course, but not great for the paid version.
For longer writing,I really like going for a walk for 45 minutes and brain dumping on a topic, and transcribing it. Then I write a brief outline and have Claude fill it out into a document, explicitly only using language that I used in the transcript. Then edit via voice
Interesting. I felt GPT4 was virtually useless and GPT3.5 was the best, then came along GPT-4o and it instantly became the only version worth using.
I find GPT-4o to be extremely powerful and useful.
What don't you like about it?
But gpt-4o actually got a working solution in a couple prompts -> https://gist.github.com/thomasdavis/fadbca01605fb3cb64911077...
Though the new artefacts is really damn handy, you can describe the most detailed projects, and it does a really great job of what you asked for, and I found it delivered beyond what I wanted.
I am now paying for both -.-
- it's hard to rank which codes better, but I think claude has better abstractions
- sometimes I paste the output of the other, and continue solving on the other
Would love to see your workflow sometime, my experimentations have been small tbh but I was quite impressed with Claude.
Funnily, I think (modulo costs) feeding outputs of one model into another once you hit a brick wall is quite powerful, but feels clunky to do at the moment
PS Anybody is welcome to jump in with a description of their setup.
You can pass the current file, or multiple files or your whole project as context to the models.
Anthropic tech is vastly better at code.
You get to create a set of default instructions that all project prompts to adhere to, like: Use Linq whenever possible. Consider Fluent capabilities when it will improve developer experience.
Then you can add all of your existing code (though as has been pointed out, connecting to a repo would be ideal).
And then you can have multiple conversations within that project, so you can separate your lines of thinking.
This is the problem that LLMs have not yet solved.
I haven't used the normal website in a year, it is a way better experience:
- It allows tweaking of the model / temperature.
- You pay only per token used
- It does not use your conversations to train their models
- You can edit both your and the assistant message in the conversation or delete it.
- Allows for easy message replay if the output is not great
Be mindful that you are paying per token for every message in the thread, so if the thread becomes long it gets expensive, but still far more value than subscribing.
------
USER
What is the maximum number of tokens supported by gpt-4o?
ASSISTANT
As of my most recent update in October 2023, GPT-4 does not have an officially designated variant called "gpt-4o". However, OpenAI's GPT-4 model supports different maximum token limits depending on its variant. There are two primary context lengths for GPT-4 models:
1. *GPT-4 with 8K context length*: This variant supports up to 8,000 tokens. 2. *GPT-4 with 32K context length*: This variant supports up to 32,000 tokens.
These limits encompass both input and output tokens in a given interaction. If you intended to refer to a specific or newly introduced variant, such as "gpt-4o," it is essential to look up the most current resources or official announcements from OpenAI, as models and their capabilities can evolve.
If you think you can save money by paying per token in the playground, you should know that you're not getting the same results.
For example, if I ask Gemini the identical question, it gives a reasonable answer, even though Gemini was released before ChatGPT was released. It can look up the current OpenAI documentation and summarize it:
USER
What is the maximum number of tokens supported by gpt-4o?
ASSISTANT
GPT-4o has a context window of 128,000 tokens. This means it can process up to 128,000 tokens between your input and its output combined. However, the maximum number of output tokens (the AI's response) is limited to 4,096 tokens.
I’ve also had worse results with 4o than 4. It’s faster but it vomits output.
I can’t tell if I need to go tweak memory or if it’s just like this.
Where will all the power and computers come from that will replace the tens of millions of software developers?
If OpenAI comes out with a stronger model I’ll prob switch back, I don’t have much brand loyalty in this regard. I could see some features making usage more sticky (memory, projects, etc).
Cody's chat UI seems pretty good for making comparisons like this. You can set it to whichever LLM you want, including GPT-4o and Claude 3.5 Sonnet.
I haven't used Claude enough to do to a detailed comparison, but GPT4o and Claude 3.5 seem roughly similar for my coding questions.
It is surprisingly good and helpful. I am still exploring the limits.
Writing doc comments or test cases is much faster and more fun with this kind of tool, but you still have to double check everything as it inevitably make mistakes, often small and not obvious…
My experience with Claude is very positive when it comes to programming and planning out infrastructure. My only gripe so far has been some ethical constraints that didn't exist for ChatGPT, and those are a big one for me since I don't need Anthropic employees to act as my moral compass. For a specific example, asking about securing code through making decompiling or reading harder is a no-no for Claude, but a-ok for GPT.
Claude: subjectively sounds more human to me, and really nails data questions that 4o is lackluster at
4o: far better assistant logic reasoning. I can trivially break Claude's assistant (system prompt) instructions within the user prompt, where 4o succeeds in all of these tests.
Pricing and output speed, for our purposes, are functionally identical. Exciting to have a competitor in the space already who stands to keep openai honest.
The graph does not look like it is accelerating. I actually struggle to imagine what about it convinced the author the progress is accelerating.
I would be very interested in a more detailed graph that shows individual benchmarks because it should be possible to see some benchmarks effectively be beaten and get a good idea of where all of the other benchmarks are on that trend. The 100 % upper bound is likely very hard to approach, but I don't know if the limit is like 99%, 95% or 90% for most benchmarks.
The same problem could well be present in other benchmarks as well.
If you take the upper bounds at any given point in time, the rate of increase of the best models over time is accelerating.
+1 OpenAI Subscription -1 Anthropic Sonnet->sudden-death-automatic-review-system
Unfortunately, by the time Anthropic's next model is out you are likely to be banned by an "automatic review" again. At least that's my experience.
My working theory is that I was banned due to using a VPN (Mullvad) with a location set to Europe, at a time when Europe users were not allowed to use the app.
I wasn’t actually in Europe, to be clear, but I live in a country where even WhatsApp voice calls are blocked and so I routinely have my VPN turned on.
The country I live in is officially supported by Anthropic, and so is Europe these days, so it’s quite frustrating that they won’t unban me.
I can’t use ChatGPT and Perplexity either when I have my VPN turned on, but at least they don’t ban my account.
Fortunately, Poe is VPN friendly.
I have an idea for a project that involves streaming 3 giabits of data per second from a USB 3.0 device out over a 10 gig Ethernet connection, and it was able to compare/contrast various levels of support for high-bandwidth USB 3 and Ethernet in multiple frameworks and languages.
And the whole conversation, with code examples, cost me 3 _cents_ of Anthropic credits.
My new fear is when people start asking AIs "Hey AI, here is my codebase, my org chart, and commit histories for all my employees - how can I reduce the number of humans I need to employ to get this project done?"
From what I can tell, most programmers are more ok with LLMs directly replacing them than artists are. I tend to agree that it is better to replace programmers, and protect artists.
It doesn’t yet do prompt to 100% finished artifact.
For concept artists and illustrators at least who have very generic or commercial styles Diffusion models can create a final artifact that passes as good enough to a lot of clients.
Think developers would feel a bit different if they were in the same situation today.
How do you know such information is correct though?
For example, I wasn't sure if there was good support for USB devices in Rust. Just not something I'd ever bothered to investigate. Claude knew that there was a libusb wrapper and a "rusb" crate, and claimed that support for USB in Rust was stable. I verified the existence of the crates, and I'll take Claude at its word for stability unless I discover otherwise.
Does what they say conflict with anything you already know to be true?
Is their argument internally inconsistent, does it contradict itself?
Can you corroborate parts of their argument from other independent sources?
It depends a lot on the domain of course, but I'd bet that frontier LLMs already exhibit superhuman capabilities in providing accurate answers in the vast majority of domains.
1. Good, detailed commits 2. Tests 3. Docstrings
In TFA they name OpenAI, Google DeepMind and Anthropic.
Its frustrating to work with the AI to implement something only to realise within a few interactions that it has forgotten or lost track of something I deemed to be a key requirement.
Surely the future of software has to start to include declarative statement prompts as part of the source code.
You could set a convention of having comments something like this:
# Requirement: function always returns an array of strings
And then have a system prompt which tells the model to always obey comments like that, and to add comments like that to record important requirements provided by the user.It may not be great for every workflow, but it certainly hits a sweet spot for intelligence x cost on most of my workflows.
If you want me to try your service, try using some flow with less friction than sandpaper, folks.
I did have some service refuse it, I want to say Twitter? But I'd definitely not consider it common. This is probably only the second or third time I've seen it, tbh.
Since Anthropic accounts come with free rate-limited access to their models they're trying to avoid freeloaders who sign up for hundreds of accounts in order to work around those caps.
Google Voice numbers are blocked because people can create multiple of those, which would allow them to circumvent those limits.
On one extreme usage of the service is totally open to everyone. On the other it is completely locked down and nobody can use it ever. Clearly neither of those are workable. So it’s a question of how far to go with validation.
My message was simply my opinion as a user, to anyone who may have a similar flow. No more, no less.
In other words: information with copyright on it that was (illegally) used/stolen from billions of people and companies will get copyright on it that will be resold as a set?
I don't know...
(Assuming it is really that good, which is something that I can never replicate.)
I'm legit elated that a smaller player is able to compete with large behemoths like OpenAI and Google. (I know they have Amazon backing them, but their team is much smaller. OpenAI is ~1000 employees now).
I'm building on top of their api. It's neat. I wish them the best.
It did seem slightly odd to me that openai made their supposedly best model free.
gpt-4o: "Yes, 400,000 is the same as 4 X 100 × 1000."
gpt-4:
Yes, 400,000 is the same as 4 X 100 × 100.
Here's the breakdown:
4 × 100 = 400
400 × 100 = 40,000
Oops, it looks like there was an error in my previous calculation.
So, 4 × 100 X 100 actually equals 40,000, not 400,000. To make it 400,000,
you would need to use 4 × 10 x 10,000 or adjust the numbers similarly.
claude-3.5: No, 400000 is not the same as 4 * 100 * 100.
Let's break it down:
4 * 100 * 100 = 4 * 10000 = 40000
400000 is ten times larger than 40000.
Everyone keeps saying gpt-4o beats benchmarks and stuff, but this is consistently my experience with it. The benchmarks fall far off my every day experience.Here is an example using http://github.com/robusta-dev/holmesgpt/ where ask it to troubleshoot a payment processing workload in a Kubernetes cluster vs the GPT-4o output underneath:
holmes ask --model="bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0" "what is wrong with the payment processor?"
AI: To investigate what's wrong with the payment processor, we'll need to gather more information about it. First, let's try to find the relevant Kubernetes resources related to the payment processor.
Running `kubectl get -A --show-labels pod | grep payment`
Thank you for your question about the payment processor. initial search, I've found two pods related to payment processing. Let's investigate the one that seems to be having issues:
1. In the default namespace, there's a pod named "payment-processing-worker-5bb5bd75cd-x4g99" which is in a CrashLoopBackOff state. This indicates that the pod is repeatedly crashing and restarting.
2. In the sock-shop namespace, there's a pod named "payment-ff86cd6f8-rs47b" which appears to be running normally.
Let's focus on the problematic pod in the default namespace. To get more information about what's wrong, we'll need to look at its logs and description.
Running `kubectl describe pod payment-processing-worker-5bb5bd75cd-x4g99 -n default`
AI: Based on the information gathered, it appears that there is an issue with the payment processor in the Kubernetes cluster. Let me explain what's wrong:
1 The payment processing worker pod is in a CrashLoopBackOff state:
• Pod name: payment-processing-worker-5bb5bd75cd-x4g99
• Namespace: default
• Status: CrashLoopBackOff
2 The pod is continuously restarting and failing:
• Restart Count: 242
• The container is in a "Waiting" state with the reason "CrashLoopBackOff"
3 The root cause of the issue appears to be related to an environment variable:
• The container is running a bash script that checks for the presence of the DEPLOY_ENV environment variable.
• If DEPLOY_ENV is not set, the container exits immediately after printing an error message.
• The command being run is:
if [[ -z "${DEPLOY_ENV}" ]]; then echo Environment variable DEPLOY_ENV is undefined ; else while true; do echo hello; sleep 10;done; fi
4 The environment section in the pod description shows:
Environment: <none>
This indicates that no environment variables are set for the container.
5 The container is exiting with a status of "Completed" and an exit code of 0, which suggests that it's running the script, finding that DEPLOY_E
To fix this issue:
1 The DEPLOY_ENV environment variable needs to be set for the payment-processing-worker deployment.
2 You should update the deployment configuration to include this environment variable.
3 After setting the environment variable, the pod should stop crashing and enter a running state where it prints "hello" every 10 seconds.
To resolve this, you'll need to update the deployment configuration for the payment-processing-worker. You can do this by editing the deployment and adding the DEPLOY_ENV environment variable with an appropriate value.
Would you like me to show you how to update the deployment with the required environment variable?
Note that it explicitly picks up on the ambiguity in the question and notes that there are 2 payment related workloads and one is healthy.GPT-4o:
$ holmes ask "what is wrong with the payment processor?"
Running `kubectl get -A --show-labels pod | grep payment-processor`
Running `kubectl get -A --show-labels deployment | grep payment-processor`
Running `kubectl get -A --show-labels service | grep payment-processor`
Running `kubectl get -A --show-labels pod | grep payment`
Running `kubectl describe pod payment-processing-worker-5bb5bd75cd-x4g99 -n default`
Running `kubectl logs payment-processing-worker-5bb5bd75cd-x4g99 -n default --previous`
AI: payment-processing-worker-5bb5bd75cd-x4g99 in namespace default crashes due to missing DEPLOY_ENV environment variable.
Fix: Define DEPLOY_ENV in the deployment spec.
Example patch:
env:
- name: DEPLOY_ENV
value: "production"> Given a body with momentum B and forques F, what is the differential of applying the forques to the momentum in PGA?
Claude gave a wrong answer, ChatGPT gave a correct one. I’m sticking with ChatGPT.
Testing on something widely known isn’t likely to stretch these systems.
This very thread is the fifth hit on Google for `"forques" geometric algebra`; the third and fourth hit are the same paper as each other; the second hit is https://bivector.net/PGAdyn.pdf which appears to have invented the term; and the first hit doesn't define it.
I (logic, computability, set and type theory) am in no position to know whether it's a standard term in geometric algebra, but I do strongly expect LLMs to do much worse on queries that don't appear much in their training set (for which I take Google search results as a proxy); even if they have the knowledge to answer, I expect them to answer better when the question uses common words. I do know that when I asked your question to ChatGPT, it silently re-spelt "forques" as "torques".