Currently it suffers from being very expensive (I blew through my 30€ monthly budget in a few hours of intensive use for coding).
Also it still is very confidently wrong in sometimes subtle but fundamental ways.
My recent example is this:
I had success with developing rust proc macros with it. I don't know much about them but as a developer I can read the generated code just fine.
Yesterday I wanted to code a macro that adds an attribute to fields in an existing struct. It's actually not possible to do that but gpt-4 send me down a wrong track by "fixing" it's bugs when asked to without getting anywhere.
Asking it whether this is even possible is unreliable because in such niche cases it'll flip flop between answers.
Copilot has become a valuable tool and I've learned a lot by using both versions of GPT
I'll give it a system prompt in the spirit of "you assist experts with developing software. Be brief and assume expertise".
I've found it to work well for smaller, contained problems or one off scripts where the alternative would have been to do it manually or not at all. Getting there 80% allows me to start them in the first place.
Another random example: Recently I needed a script to transfer some secrets from one k8s cluster to several others. It took about 3 minutes with GPT 4 and solved the problem within one iteration. There is probably a one liner in bash to do it but I don't know it of the top of my head;)
Copilot massively improved the quality of my logging and commenting
I wonder if more people have this suspicion or if it's just my imagination?
https://platform.openai.com/docs/model-index-for-researchers
You'll never hit a ratelimit as far as I can tell, and it's usage-based so it will probably come out cheaper than $20/mo for regular usage.
I didn't copy a reference, just been reading AI topics on HN lately.
However, very interestingly, the unicorn image got far worse when they trained the model to be safer by trying to correct discrimination against various demographics.
This isn't very intuitive to me why that may occur, and seems to conflict with what has been shown in ROME, etc. So I'm surprised it hasn't been commented upon more. It's certainly one of the best examples of how we don't understand what's going on with these models, and it causes very unexpected outcomes.
it's good, but the difference is not that big