So it's most useful to look at other capabilities and opportunities when evaluating LLM's with a different heritage.
Not to say we shouldn't evaluate this one for coding or report our evaluations, but we shouldn't be surprised that it's not leading the pack on that particular use case.
So they are going to be training on exactly the same data that is available to all.
The terms and conditions say as much https://docs.github.com/en/site-policy/github-terms/github-t...
One example is the data hosted in Google Cloud.
https://cloud.google.com/blog/topics/public-datasets/github-...
I added it to my custom instructions and it has helped a lot.
Today I have the same experience. The thing fills in placeholder comments to skip over more difficult regions of the code, and routinely forgets what we were doing.
Aside all the recent OpenAI drama, I've been displeased as a paying customer that their products routinely make their debut at a much higher level of performance than when they've been in production for a while.
One would expect the opposite unless they're doing a bad job planning capacity. I'm not diminishing the difficulty of what they're doing; nevertheless, from a product perspective this is being handled poorly.
Also, the only way for OpenAI to really know if a model is an improvement or not is to test it out on some human guinea pigs.
Are you prompting it with instructions about how it should behave at the start of a chat, or just using the defaults? You can get better results by starting a chat with "you are an expert X developer, with experience in xyz and write full and complete programs" and tweak as needed.
eg: Write clean {your_language} code. Include {whatever_you_use} conventions to make the code readable. Do not reply until you have thought out how to implement all of this from a code-writing perspective. Do not include `/..../` or any filler commentary implying that further functionality needs to be written. Be decisive and create code that can run, instead of writing placeholders. Don't be afraid to write hundreds of lines of code. Include file names. Do not reply unless it's a full-fledged production ready code file.
It's pretty funny that my second message is often "that doesn't look like any programming language I recognize. I tried running it in Python and got lots of errors".
"My apologies, that message was an explanation of how to solve your problem, not code. I'll provide a concrete example in Python."
> As of my last knowledge update in September 2021, the XY framework did not have a --abc or --bca option in its default project generator.
Huh...
Ideal output is when nobody elese is using the tool.
Sounds like a kinda expensive way of doing things, to me.
[1] https://www-files.anthropic.com/production/images/model_pric...
OpenAI? I use ChatGPT A LOT for coding as some mixture of pair programmer and boilerplate, works generally well for me. On the API side use it heavily for other work and its more directed and have a very high acceptance rate.
GPT4 massively sped up my ability to create this.
It is a tool and it takes a lot of time to master it. Took me around 3-6 months of every day use to actually figure out how. You need to go back and try to learn it properly, it's easily 3-5x my work output.