I think that the next step is getting an official "checked" mark by the SWE bench team
20 karma · joined November 17, 2022
I think that the next step is getting an official "checked" mark by the SWE bench team
We will see more of these frameworks for different use cases
Specifically for generating unit regression tests the Cover-Agent tool already works quite well in the wild for some projects, especially isolated projects (as opposed to complex enterprise-level code). You can see in the few (somewhat cherry-picked) examples we posted [0] that it generates working tests that increase coverage (they were cherry-picked in the sense that these are examples we like to work with often internally at CodiumAI).
I believe that it’s possible to generate additional meaningful tests including end-to-end tests by creating a more sophisticated flow that uses prompting techniques like reflection on the code and existing tests, and generates the tests iteratively, feeding errors and failures back to the LLM to let it fix them. Just as an example. This is somewhat similar to the approach we used with AlphaCodium [1] which hit 54% on the CodeContests benchmark (DeepMind’s AlphaCode 2 hit 43% [2] with the equivalent amount of LLM calls).
If like me you think tests are important but hate writing them, please consider contributing to the open source to help make it work better for more use cases. https://github.com/Codium-ai/cover-agent
[0] https://www.youtube.com/@Codium-AI/videos [1] https://github.com/Codium-ai/AlphaCodium [2] https://storage.googleapis.com/deepmind-media/AlphaCode2/Alp...
In the near future, the agent will suggest these for you, after first indexing your code base
To put it simple: Like copilot, it does have an auto-complete and also chat interface. Different than copilot, it focuses on generating a full code task plan, then having the auto completion work acrroding to your plan, and it check the of quality code.
Like agents, it tries to help you complete a full task, yet,does that in tandem with you, as you work inside tour favorite IDE writing the code with you. In addition, there is focus on code quality, testing, and fetching relevant context from your codebase.
We’ve come up with a bit of a different concept for what a coding agent should be. We believe it should work in tandem with a developer inside the IDE. Over time as the tech improves, it will get more and more autonomy.
We’ve been using our coding agent internally and see a 5-10x boost on some tasks.
The agent is available now to Codiumate VS Code users. We want to hear what kind of tasks it works well on and improve it over time to expand the task set. Would love to get feedback.
https://marketplace.visualstudio.com/items?itemName=Codium.c...
Paper: https://arxiv.org/abs/2401.08500 Blog: https://www.codium.ai/blog/alphacodium-state-of-the-art-code...
I wonder if you have tried tools that are dedicated to Code Integrity, e.g. generating tests?
I wrote a blog about it: https://www.codium.ai/blog/code-integrity-supercharges-code-...
disclaimer: I'm the co-maker of PR-Agent and CodiumAI
The benchmark - Codeforces programming contest.
GPT-4 Codeforces Rating is 392 points, improving GPT-3.5’s 260 points.
AlphaCode by DeepMind achieves 1,238 points!
Those who tried Bard or Codey will very likely agree that Google’s models and solutions are not better than OpenAI ones. So, what is going on here?
The secret sauce? AlphaCode is composed of two components, not one. A Code Generation component, and a Code Integrity component (that includes test generation, filtering and clustering according to tests runs).
Considering the latest news, like llama.cpp, new code generators, etc... maybe it is doable?
Can AI tools help in generating meaningful tests? In the linked post, Ankur Tyagi reviews CodiumAI tools and vision.
>>> CodiumAI’s vision is simple; an AI coding assistant/agent to assist developers in reaching zero bugs. >>> Codium AI automatically generates Happy Paths, Edge Cases, and Other test suites when you finish writing your code and saving it.
Can AI-empowered tools really understand develop intent and analyze the code to generate edge cases?
some selected quotes from the blog
1) "In SW 3.0, the business logic, I/O control, and the data pre/post-processing code are partially or even entirely created by the AI agent during the optimization process"
the `optimization process` for the entire chain of actions is still immature
2) "It is likely that any setting where the program is not obvious but one can repeatedly evaluate its performance (e.g. - programming competitions?) will be subject to this transition, because the optimization can probably find much better code than what an average human can write"
3) "It is likely that testing [automation] will become even more crucial in the SW 3.0 stack for a number of possible reasons. To name two: 1) programmers would be able to do more, including less experienced ones, 2) programs will be much more dynamic"
4) "In the coming years, we will see immense research efforts to improve the setup and learning schemes used to create AI agents that generate programs"
4) "Bottom-up: Functions, classes, and capabilities will be generated with newly introduced IDEs, code management, testing, code generation and analysis of SW 3.0 stack. Top-down: graphical and programable application interfaces will be generated with newly introduced no/low-code SW 3.0 based platforms"
did these age well?
do you think that the vision of AI Agents will be realized before 2025?
curious, would you consider `ChatGPT + Code interpreter` as an agent?