30 karma · joined December 13, 2018
We have been building agents for code review workflow for nearly two years. Our Code Review Knowledge base today serves more than 3 million repos. We pack more than 40 different points of information into the same LLM context, such as MCP Servers, Rule files, etc. We understand context poisoning/rot and how it creating problems.
We are now bringing the same learning/context engine to SDLC and Slack.
Please give it a try!
Here is the TL;DR
- GPT-5 outperformed Opus-4, Sonnet-4, and OpenAI’s O3 across a battery of 300 varying difficulty, error-diverse pull requests.
- GPT-5 scored highest on our comprehensive test and found 254 out of 300 bugs or 85% where other models found between 200 and 207 – 16% to 22% less.
- On our 25 hardest PRs from our evaluation dataset, GPT-5 achieved the highest ever overall pass rate (77.3%), representing a 190% improvement over Sonnet-4, 132% over Opus-4, and 76% over O3.
This reminds me of "big companies moves slow.." line.
They should have restricted the Marketplace several years ago, however, they are doing it now.
With C++, they are part of MFC's, they are the legal owners, not like Google vs Oracle in case of Java.
Lastly, with AI Code IDEs I think yes, there is a case, the need for IDE might be very less. Like a steering on a self driving car.
Suddenly after reasoning models, it looks like OSS models have lost their charm
We are hiring in BLR and SF.
I agree. I don't perse hate linters, too. I used a combination of them and love getting nits done.
At some places, we even had a pre-commit hook configured to have linters completed, so that was fun.
Lint rules should evolve over a period of time for teams to get better or as the project gets better, in my opinion. And I think there is no such process at many teams today.
End of they day, they are a business!
1. When you say backends, do you plan to integrate like a client with some "vector" stores. 2. Also any benchmarks? 3. Lastly, why python?
OR set the S3 bucket to public :D
It looks like we can get full context-aware config reviews with Gen AI.
> Microsoft Research and LinkedIn researchers have open-sourced CodeReviewer, a pre-trained transformer model that can automatically assess code changes, generate review comments, and suggest fixes. Trained on 7.9M pull requests across 9 programming languages, it achieves a 71.5% F1 score in identifying problematic code changes and can generate relevant review comments with 3.6/5.0 informativeness rating from human evaluators.
> Unlike existing code models, CodeReviewer is specifically trained on code diffs and real-world review comments from high-quality GitHub repositories. The model outperforms previous approaches by learning to "think" like a code reviewer rather than just understanding source code.
Technical details and model available at: https://github.com/microsoft/CodeBERT/tree/master/CodeReview...
I'm thinking GitHub is gonna use some of this learning in Copilot?
I'm afraid that some issues could arise at runtime, such as CORS problems, which even an experienced developer might overlook.
I don't believe this is simply a list of static analysis issues that SonarQube can identify. One of the advantages of using AI (despite its tendency to hallucinate and be overly picky about minor details) is its ability to generate a fix or several variations of fixes that we can test.
P.S. This gave me a good idea to check if I can run the community edition for testing purposes.
We at CodeRabbit are building a product that helps augment and cut existing code review hubris. We don't think it is just a review problem, through the automated review, we help run code quality workflow on PRs. So it is like improving code quality PR-by-PR.
We also focus on supressing the noise from linters and bringing important issues to the top.
CodeRabbit is used by many OSS projects today. With nearly 4 million PRs reviewed, nearly a million repo's, 15k+ interactions. We are just getting started. Please give it a shot and let us know your feedback. Here or our discord (https://discord.gg/coderabbit)
What I love with these tools, I'm getting productive. Although, I'd still like to do the review myself.
Besides, I think X aka Twitter is not supported as of now?
Lastly, what do you think of Typefully? It is amazing too I think.
We cant say, one is need and one is not!
They tried to analyse on:
- Mathematical Accuracy - Data Analysis - Actionability - Summarisation
They have a DORA metrics product which show: Lead Time for Changes, Deployment Frequency, Mean Time to Recover (MTTR), Change Failure Rate (CFR) for engineering teams. (https://github.com/middleware/middleware)
I think still it is not the best or efficient way to compare. What do you think? Any suggestions on comparisions for specific use cases like these.