I think calling it intelligent is being extremely generous. Take a look at the following example which is a spelling and grammar checker that I wrote:
https://app.gitsense.com/?doc=f7419bfb27c89&temperature=0.50...
When the temperature is 0.5, both Claude 3.5 and GPT-4o can't properly recognize that GitHub is capitalized. You can see the responses by clicking in the sentence. Each model was asked to validate the sentence 5 times.
If the temperature is set to 0.0, most models will get it right (most of the time), but Claude 3.5 still can't see the sentence in front of it.
https://app.gitsense.com/?doc=f7419bfb27c89&temperature=0.00...
Right now, LLM is an insanely useful and powerful next word predictor, but I wouldn't call it intelligent.