Somehow, OpenAI is playing catch with them rather than vice versa.
Somehow, OpenAI is playing catch with them rather than vice versa.
It really depends on the language and the prompt. Sometimes one shines and the other produces garbage and it's usually 50/50
I've found it to be on par with Stack Overflow / Google Search.
More convenient than cut/paste but more prone to inaccuracies and out of context answers.
But at no point did it remotely feel like a top tier programmer.
These models are good at generating template code and many straightforward things, but if you add anything complex, you start wasting your time.
A concrete example: I was doing shader programming with Sonnet 3.5 and ran into a visual bug. Sonnet asked me to add four debugging modes, cycle through each one, and describe what I saw for each one. With one more prompt, it resolved the issue. In my experience, GPT-4o has never bothered proposing debug modes and just produced more buggy code.
For non-trivial coding, Sonnet 3.5 was miles above anything else, and I didn't even have to try hard.
They might be able to debug it themselves, maybe they should be able to debug it themselves. But I feel like that is a completely different conversation.
If you're curious, I knew nothing about shader programming when I first played around. In that specific experiment, I wanted to see how far I could push Claude to implement shaders and how capable it is of correcting itself. In the end, I got a pretty nice dynamic lighting system with some cool features, such as cast shadows, culling, multiple shader passes, etc. Asking questions along the way taught me many things about computer graphics, which I later checked on different sources, it was like a tailored-made tutorial where I was "working" on exactly the kind of project I wanted.
i.e. there would be a lot of value if an AI could maintain a detailed understanding of say, the Linux kernel code base, when someone is writing a driver and actively prompt about possible misuses, bugs or implementation misunderstandings.
What its best for is creating components and functions that are labor intensive but fairly standardized
Like if you have a CRUD app and want to add a bunch of filters, complete with a solid UI, you can hand over this to Sonnet and it will do a fine job right out of the box
Umm... why?
Nobody else in the AI space wants to track my number.
I'm sure Anthropic has their "reasons". I just doubt it is one that I would like.
> Umm... why?
https://support.anthropic.com/en/articles/8287232-why-do-i-n...
My guess is, these models are incredibly expensive to run, Claude has a fairly generous free tier, and phone numbers are one of the easiest ways to significantly reduce the number of duplicate accounts.
> Nobody else in the AI space wants to track my number.
Given they're likely hoovering up all of the data you're sending to them, and they have your email address to identify you, this seems like an odd hill to die on.
Whether they are serious about it or use it as an excuse to collect more PII (or both/neither), collecting verified phone numbers presumably allows them to demonstrate compliance.
[0] https://cset.georgetown.edu/article/dont-forget-the-catch-al...
In the US, other locations may/may not have the same export controls. Base your AI business in one of the non-US countries and it'll be legal to not keep strict controls on who is using your service.