It's great as a non offensive stack overflow replacement, but just look at aider benchmarks (amazing work by the way): most capable models really struggle to make basic changes to a real code base.
Does anybody actually using it in practice believes the hype? I thought the hype is just another theatre for investors
EDIT: just some points taken from: https://aider.chat/docs/leaderboards/#code-editing-leaderboa...
- The metric "Percent completed correctly" maxes out at 72.9% with gpt4o, while at the same time, giving out correctly formatted output only 96.2% of the time.
- Benchmark suite is based on https://github.com/exercism/python, which very likely is a part of the training material already! In real code bases, no LLM would have the advantage of seeing new or proprietary code