I would absolutely choose to use Claude as my model with ChatGPT if that happened (yes, I know it won't). ChatGPT as an app is just so far ahead: code interpreter, web search/fetch, fluid voice interaction, Custom GPTs, image generation, and memory. It isn't close. But Claude absolutely produces better code, only being beaten by ChatGPT because it can fetch data from the web to RAG enhance its knowledge of things like APIs.
Claude's implementation of artifacts is very good though, and I'm sure that is what lead OpenAI to push out their buggy canvas feature.
--
It also features one-click installation, OpenAI integration, a hub for downloading and running local models, a spec-compatible API server, global "quick answer" shortcut, and more. Really can't recommend it enough!
[0] https://msty.app
Its most unique feature is its "beam" facility, which allows you to send a query to multiple APIs simultaneously (if you want to cross-check) and even combine the answer.
Sonnet is better in the small, by a lot. It’s sharply up from idk, three months ago or something when it was still an attractive nuisance. It still tops out at “Best SO Answer”, but it hits that like 90%+. If it involves more than copy paste, sorry folks, it’s still just really fucking good copy paste.
But for sheer “doesn’t stutter every interaction at the worst moment”? You’ve got to hand it to the ops people: 4o can give you second best in industrial quantity on demand. I’m finding that if AI is good enough, then OpenAI is good enough.
Are you sure you're using Claude 3.5 Sonnet? In my experience it's absolutely capable of writing entire small applications based off a detailed spec I give it, which don't exist on GitHub or Stack Overflow. It makes some mistakes, especially for underspecified things, but generally it can fix them with further prompting.
And I’m not sure we disagree.
Vercel demo but Pets is copy paste.
* I'm iOS by experience; my main professional JS experience was something like a year before jQuery came out, so I kinda need an LLM to catch me up for anything HTML
Also, I wanted HTML rather than native for this.
Funny thing, TypingMind was ahead of them for over a year, implementing those features on top of the API, without trying to mix business model with engineering[0]. It's only recently that ChatGPT webapp got more polished and streamlined, but TypingMind's been giving you all those features for every LLM that can handle it. So, if you're looking for ChatGPT-level frontend to Anthropic models, this is it.
ChatGPT shines on mobile[1] and I still keep my subscription for that reason. On desktop, I stick to TypingMind and being able to run the same plugins on GPT-4o and Claude 3.5 Sonnet, and if I need a new tool, I can make myself one in five minutes with passing knowledge of JavaScript[2]; no need to subscribe to some Gee Pee Tee.
Now, I know I sound like a shill, I'm not. I'm just a satisfied user with no affiliation to the app or the guy that made it. It's just that TypingMind did the bloodingly stupid obvious thing to do with the API and tool support (even before the latter was released), and continues to do the obvious things with it, and I'm completely confused as to why others don't, or why people find "GPTs" novel. They're not. They're a simple idea, wrapped in tons of marketing bullshit that makes it less useful and delayed its release by half a year.
--
[0] - "GPTs", seriously. That's not a feature, that's just system prompt and model config, put in an opaque box and distributed on a marketplace for no good reason.
[1] - Voice story has been better for a while, but that's a matter of integration - OpenAI putting together their own LLM and (unreleased) voice model in a mobile app, in a manner hardly possible with the API their offered, vs. TypingMind being a webapp that uses third party TTS and STT models via "bring your own API key" approach.
[2] - I made https://docs.typingmind.com/plugins/plugins-examples#db32cc6... long before you could do that stuff with ChatGPT app. It's literally as easy as it can possibly be: https://git.sr.ht/~temporal/typingmind-plugins/tree. In particular, this one is more representative - https://git.sr.ht/~temporal/typingmind-plugins/tree/master/i... - PlantUML one is also less than 10 lines of code, but on top of 1.5k lines of DEFLATE implementation in JS I plain copy-pasted from the interwebz because I cannot into JS modules.
Which app are you talking about here?
Either way, Claude is great so this is a net win for everyone.
A commenter on another thread mentioned it but it’s very similar to how search felt in the early 2000s. I ask it a question and get my answer.
Sometimes it’s a little (or a lot) wrong or outdated, but at least I get something to tinker with.
Google used to prioritise big comprehensive articles on subjects for desktop users but mobile users just wanted quick answers, so that's what google prioritised as they became the biggest users.
But also, per your point, I think those smaller simpler less comprehensive posts are easier to fake/spam than the larger more compreshensible posts that came before.
Right now, I find each tool better at different things.
If I can only describe what I want but don't know key words, LLM are the only solution.
If I need citations, LLMs suck.
If this is the state of the art of coding LLMs, I really don't see why I should waste my time evaluating their confident sounding, but wrong, answers. It doesn't seem like much has improved in the past year or so, and at this point this seems like an inherent limitation of the architecture.
Codebases are exploding in size. Feature development has slowed down.
What might have been a carefully designed 100kloc codebase in 2018 is now a 500kloc ball of mud in 2024.
Companies need many more developers to complete a decent sized feature than they needed in 2018.
Once you start shifting it to micro services the business logic gets spread out and duplicated.
At the same time each micro-service now has its own code to handle rest, graphql, grpc endpoints.
And each downstream call needs error handling and retry logic.
And of course now you need distributed tracing.
And of course now your auth becomes much more complex.
And of course now each service might be called multiple times for the one request - better make them idempotent.
And each service will drift in terms of underlying libraries.
And so on.
Now we have been adding in LLM solutions so there is no consistency in any of the above services.
Each dev rather than look at the existing approaches instead asks Claude and it provides a slightly different way each time - often pulling in additional libraries we have to support.
These days I see so much bad code like a single microservice with 3 different approaches to making a http request.
Maybe I'll have better luck next time, or maybe I need to improve my prompting skills, or use a different model, etc. I was just expecting more from state of the art LLMs in 2024.
A lot of the devs in my office using Claude/gpt are convinced they are so much more productive but they aren't actually producing features or bug fixes any faster.
I think they are just excited about a novel new way to write code.
I ask it questions mostly about libraries I’m using (usually that have poor documentation) and how to integrate it with other libraries.
I found out about Yjs by asking about different operational transform patterns.
Got some context on the prosemirror plugin by pasting the entire provider class into Claude and asking questions.
It wasn’t always exactly correct, but it was correct enough that it made the process of learning prosemirror, yjs, and how they interact pretty nice.
The “complete” examples it kept spitting out were totally wrong, but the information it gave me was not.
I had a suspicion that what I was trying to do was simply not possible with that library, but since LLMs are incapable of saying "that's not possible" or "I don't know", they will rephrase your prompt and hallucinate whatever might plausibly make sense. They have no way to gauge whether what they're outputting is actually correct.
So I can imagine that you sometimes might get something useful from this, but if you want a specific answer about something, you will always have to double-check their work. In the specific case of programming, this could be improved with a simple engineering task: integrate the output with a real programming environment, and evaluate the result of actually running the code. I think there are coding assistant services that do this already, but frankly, I was expecting more from simple chat services.
Specific is the specific thing that statistical models are not good at :(
> how do I do X with library Y?
Recent research and anecdotal experience has shown that LLMs perform quite poorly with short prompts. Attention just has more data to work with when there are more tokens. Try extending that question like “I am using this programming language and am trying to do this task with this library. How do I do this thing with this other library”
I realize prompt engineering like this is fuzzy and “magic,” but short prompts have a consistent lower performance.
> In the specific case of programming, this could be improved with a simple engineering task: integrate the output with a real programming environment, and evaluate the result of actually running the code.
Not as simple as you’d think. You’re letting something run arbitrary code.
Tho you should give aider.chat a try if you want to test out that workflow. I found it very very slow.
I'm aware of that. The actual prompt was more elaborate. I was just mentioning the gist of it here.
Besides, you would think that after 30 minutes of prompting and corrections it would arrive at the correct answer. I'm aware that subsequent output is based on the session history, but I would also expect this to be less of an issue if the human response was negative. It just seems like sloppy engineering otherwise.
> Specific is the specific thing that statistical models are not good at
Some models are good at needle-in-a-haystack problems. If the information exists, they're able to find it. What I don't need is for it to hallucinate wrong answers if the information doesn't exist.
This is a core problem of this tech, but I also expected it to improve over time.
> Tho you should give aider.chat a try
Thanks, I'll do that eventually. If it's slow, it can get faster. I'd rather the tool be slow but give correct answers, than it slowing me down by wasting my time error correcting it.
Thankfully, these approaches can work for programming tasks. There is not much that can be done to verify the output of any other subject.
Anthropic's seems to have addressed the issue using pydantic but I haven't had a chance to test it yet.
I pretty much use Anthropic for everything else.
I agree, this was a tactical move designed to give them leverage over OpenAI.
Still i find very little use from LLMs in this front, but they do come in handy randomly.
Because.. yea, it is. However.. it keeps expanding, it keeps getting more useful. Yea people and especially companies are using it for things which it has no business being involved in.. and despite that it keeps growing, it keeps progressing.
I do find the "stochastic parrot" comments slowly dwindle in number and volume with each significant release, though.
Still, i find it weirdly interesting to see a bunch of people be both right and "wrong" at the same time. They're completely right, and yet it's like they're also being proven wrong in the ways that matter.
Very weird space we're living in.
If these systems showed understanding we would notice.
No one is denying that this form of intelligence is useful.
We just need more monkeys and it will be the same as a human brain.
More than once these tools fail at tasks a fifth grader could understand
You should re-read that very slowly and carefully and really think about it. Calling anyone that's skeptical a 'denier' is a red flag.
We have been through these AI cycles before. In every case, the tools were impressive for their time. Their limitations were always brushed aside and we would get a hype cycle. There was nothing wrong with the technology, but humans always like to try to extrapolate their capabilities and we usually get that wrong. When hype caught up to reality, investments dried up and nobody wanted to touch "AI" for a while.
Rinse, repeat.
LLMs are again impressive, for our time. When the dust settles, we'll get some useful tools but I'm pretty sure we will experience another – severe – AI winter.
If we had some optimistic but also realistic discussions on their limitations, I'd be less skeptical. As it is, we are talking about 'revolution', and developers being out of jobs, and superintelligence and whatnot. That's not the level the technology is at today and it is not clear we are going to do anything else other than get stuck in a local maxima.
There's the question, "is an LLM just autocomplete"? The answer to that question is obviously no, but the question is also a strawman - people who actually use LLM's regularly do recognize that there is more to their capabilities than randomized pattern matching.
Separately, there's the question of "will LLM's become AGI and/or become super intelligent." Most people recognize that LLM's are not currently super intelligent, and that there currently isn't a clear path toward making them so. Still, many people seem to feel that we're on the verge of progress here, and feel very strongly that anyone who disagrees is an AI "doomer".
Then there's the question of "are we in an AI bubble"? This is more a matter of debate. Some would argue that if LLM reasoning capabilities plateau, people will stop investing in the technology. I actually don't agree with that view - I think there is a lot of economic value still yet to be realized in AI advancements - I don't think we're on the verge of some sort of AI winter, even if LLM's never become super intelligent.
I think calling it intelligent is being extremely generous. Take a look at the following example which is a spelling and grammar checker that I wrote:
https://app.gitsense.com/?doc=f7419bfb27c89&temperature=0.50...
When the temperature is 0.5, both Claude 3.5 and GPT-4o can't properly recognize that GitHub is capitalized. You can see the responses by clicking in the sentence. Each model was asked to validate the sentence 5 times.
If the temperature is set to 0.0, most models will get it right (most of the time), but Claude 3.5 still can't see the sentence in front of it.
https://app.gitsense.com/?doc=f7419bfb27c89&temperature=0.00...
Right now, LLM is an insanely useful and powerful next word predictor, but I wouldn't call it intelligent.
Wouldn't this make chimpanzees and ravens and dolphins unintelligent too? You're asking it to do a task that's (mostly) easy for humans. It's not a human though. It's an alien intelligence which "thinks" in our language, but not in the same way we do.
If they could, specialized AI might think we're unintelligent based on how often we fail, even with advanced tools, pattern matching tasks that are trivial for them. Would you say they're right to feel that way?
Animals have the ability to learn and grow by themselves. LLMs are not intelligent and I don't see how they can be since they just follow the most likely path with randomness (temperature) sprinkled in.
Second, it's wrong. LLMs can learn within their context window. The main issue now is the limited size of their context window; animals have a lifetime of compressed context and LLMs only have approximately one conversation.
It honestly made no sense what you were saying so I didn't respond to that directly as I assumed it would be clear from my explanation as to why animals can be intelligent and LLM are not.
> LLMs can learn within their context window.
They don't learn from the context window as much as they use what is in the context window to define a probabilistic path. If you put something in the context window that it was never trained on, it would spit out BS or say it doesn't know.
If these tools boost tue productivity where is the output spike of all the companies, the spike in revenue and profits?
How often do we lose the benefit auto text generation to the loop of That’s wrong Oh yes of course, here is the correct version Nope, still wrong Prompt editing?
When it works, it's pretty good, and sometimes great. But when failure modes look like the above I'm very wary of accepting its output.
But it still does the tasks you asked for, so that's the part that really matters.
Phind is useful as you can switch between them -- but only get a handful of o1 and Opus a day which I burn through quick at moment on deeper things -- Phind-405b and 3.5 Sonnet are decent for general use