On the other hand, if ChatGPT provides a competitive edge in CodeForce problems then it would be worrying indeed.
On the other hand, if ChatGPT provides a competitive edge in CodeForce problems then it would be worrying indeed.
When you feed one level of prompt output into the next using some functional tricks, the ability to layer high-level ideas into the engine become extremely powerful. Even if this isn't 'true' AGI, with a few prompt scaffolding libraries, it won't exactly matter. I'm convinced this is going to radically shake up the entire software field at a process level. Both exciting and scary times ahead.
What your thoughts on doing TDD with this, write the unit tests to write the code. But go one step further and have it generate tests from the spec. Repeat until 1=1.
With the pipeline as such: business need -> tasks/cards outlined -> tests written -> code written -> refactor as changes needed
There is still a need for someone to manage the business to task relationship. Let the business person prompt, and the engineer to confirm the prompt output tasks makes sense technically, edit for any errors, and do any modifications needed for specific architectural choices or desired abstractions. With well-enough defined tasks, you can start to write the tests that conform to those tasks. Again, engineer checks the prompt-output, ensuring the tests line up with the card, making any edits as required. With tests written, the code can be written such that it conforms to the tests for correctness, cards for general I/O, and business use-case for domain-specific variables and such.
It's the same domain splits that occur in our current day-to-day practice that are cause for pain. Business person and product person have a miscommunication, the wrong tasks/cards get outlined. The task writer and the test writer have a miscommunication, the tests get written poorly and problem the business is trying to solve gets murky. The tests are written poorly so the code is written poorly. The code is written poorly so the product must be refactored.
The understanding of each others intent must be had and communicated effectively or the downstream problems will mount quickly. It's in this area where GPT-3 still seems a bit lacking, and maybe GPT-4 will resolve the issue a bit. Another worry I have with this sort of model is the spaghetti and debuggability a misunderstanding may have. When code is written more slowly, as is the case today, one is able to mentor the junior and provide feedback over the process. This not only prevents a huge mess, but allows the team to resolve any initially unsaid misunderstandings.
The speed and scale of messes one can now create with this tool is completely massive!
In several previous years I have rescheduled things so I can stay up until the problems are released in my time zone to see how high on the leaderboard I could get.
It is sad that the AoC leaderboard will now just be filled with GPT entries.
For example, try giving ChatGPT the problem statement of https://adventofcode.com/2021/day/19. What comes out is basically gibberish. It seems to completely miss that the scanners report their coordinates in their own coordinate system.
But if nothing changes in the rules/community the leaderboards of early and more straight forward problems will probably be filled with GPT solutions.
But yeah, this probably mostly helps for easier problems, and is quite useless when either the problem needs a cleaver construction or some trick to make the time complexity good enough.
Deepmind has a project called AlphaCode, where they got in the middle of div2 rounds on codeforces in february 2022 (virtual): https://www.deepmind.com/blog/competitive-programming-with-a...
In the blogpost it actually comes up with a linear stack based solution (one of the intended solutions) to a problem where the naive solution is quadratic.