99 karma · joined November 13, 2022
I put those in another box, seal it completely, and write a message on the outside along the lines of "presumed useless - Sep 2016".
If I come across that box in, say, September 2026, I feel entirely comfortable tossing the whole thing.
The result is http://getcaliper.dev.
It has a number of mechanisms that help substantially:
1. It can extract deterministic quality checks from your CLAUDE.md text; these checks then get executed after every agent turn.
2. It performs a lightweight ai-powered review at every commit; feedback goes directly to the agent, which can then make corrections.
3. It performs a more 'traditional' deep AI review at merge, or on-demand.
Free to use, just bring your own API key. Any and all feedback is welcome!
At that rate, even if you just look at extending the A/C/E from Jamaica to JFK, you're talking about 15B or so USD. And compared to today's [subway|LIRR] -> airtrain system, you probably only cut about 25% of the travel time (from 60 minutes down to 45 minutes)
Compare that to, for example, the Gateway Tunnel, estimated to cost about 16B USD and double the daily commuter capacity from NJ to NYC (including traffic to and from EWR!), and it's hard to justify new infrastructure to make it easier to get to the airport.
1. In NYC Subway, a Case Study in Runaway Transit Construction Costs - Bloomberg https://share.google/SPcN8iRDZG7lNiwt9
Model output volumes mean that code review only as a final check before merge is way too late, and far too burdensome. Using AI to review AI-generated code is a band-aid, but not a cure.
That's why I built Caliper (http://getcaliper.dev). It's a system that institutes multiple layers of code quality checks throughout the dev cycle. The lightest-weight checks get executed after every agent turn, and then increasingly more complex checks get run pre-commit and pre-merge.
Early users love it, and the data demonstrates the need - 40% of agent turns produce code that violates a project's own conventions (as defined in CLAUDE.md). Caliper catches those violations immediately and gets the model to make corrections before small issues become costly to unwind.
Still very early, and all feedback is welcome! http://getcaliper.dev
I've tried to tackle a similar problem with a couple different approaches.
One is a command I call "/retro" which basically goes through all recent history on a project - commits, prs, pr comments, etc, and analyzes the existing documentation to identify how to improve it in ways that would prevent any observed issues from happening again. This is less about adding structure to the docs (as AgentLint does) and more about identifying coverage gaps.
The other is a set of tooling I've built to introduce multiple layers of checks on the outputs of agentic code. The initial observation was that many directives in CLAUDE.md files can actually be implemented as deterministic checks (ex: "never use 'as any'" --> "grep 'as any'"; by creating those deterministic checks and running them after every agent turn, I'm able to effectively force the agent to retain appropriate context for directions.
The results are pretty astounding - among early users, 40% of agent turns produce code that doesn't comply with a project's own conventions.
The system then layers on a sequence of increasingly AI-driven reviews at further checkpoints.
I'd love feedback: http://getcaliper.dev
I have found that this no longer needs to be an issue with agentic coding tools.
Once I am happy with the end state of a branch, I tell Claude to rebuild the change from scratch as a set of atomic incremental commits. It adds about 2 minutes to the dev process, but creates a pr that is infinitely easier to review.
The overall thrust of the article is great, though. The tooling around prs needs a ton of attention.
I've been coming at this from a slightly different angle, building a system that encodes checks that run against code authored by Claude, with escalating levels of enforcement.
It's still really early, but, among early users, we see that about 40% of agent turns produce code that violates the guidance in the project's own CLAUDE.md. We catch the violations at the end of the turn and direct Claude to correct them.
Still looking for early users and feedback. http://getcaliper.dev
In general I think there has been (a) too little utilization of the hooks provided by agents for deterministic checks and follow-ups to agentic coding, and (b) too little investment in creating re-usable org-wide monitoring controls and reporting on convention enforcement.
Oculi looks like a strong advance on both fronts.
However, git doesn't work in my team as a document collaboration mechanism - the commenting and discussion on a file has so much more friction that Google Docs or Notion.
What I've ended up on is a workflow where I iterate on the doc until I am satisfied, then export it to notion (via MCP) and publish that for feedback. Once there is consensus on the notion doc, that becomes the artifact of record.
It's a suboptimal solution, though - claude makes a number of poor design choices when exporting .md -> notion.
I strongly agree with you that the solution likely involves pushing the correction mechanism much closer to the point of code generation. You want to put the AI back on track as soon as it starts to stray, you can't let it build a lot on top of a mistake.
My own attempt at resolving this involves running a set of deterministic checks on agent-generated code at the end of every agent turn, along with a lightweight AI-powered review on every commit, and deep AI review on PRs before merge. I am pretty happy with the results so far.
We're not concerned about hiring for the 'skill' of using these things, but more as a culture check - we are a very AI-forward company, and we are looking for people who are excited to incorporate AI into their workflow. The best evidence for such excitement is when they have already adopted these tools.
Among the team, the expectation is that most code is being produced with AI, but there is no micromanager checking how much everyone is using the AI coding tools.
We built AI code generation tools, and suddenly the bottleneck became code review. People built AI code reviewers, but none of the ones I've tried are all that useful - usually, by the time the code hits a PR, the issues are so large that an AI reviewer is too late.
I think the solution is to push review closer to the point of code generation, catch any issues early, and course-correct appropriately, rather than waiting until an entire change has been vibe-coded.
IIUC, the core of your approach is that decisions get locked as immutable constraints and then served to agents via MCP when they query for context — is that right? Is there anything you can do to ensure that the MCP actually gets called?
I've been experimenting with an approach that triggers a set of checks on the Claude Code stop hook for lightweight, deterministic checks, then does a quick AI-powered review as a pre-commit hook, with a mechanism in place to ensure that appropriate elements of context from repo docs get loaded into context during the review.
One thing I've had to put a good bit of work into is orchestrating what happens when the existing rules don't cover some particular case. How do you think about that?
For the messier stuff — logic bugs, security issues, spec drift against your project's actual policy — there's an optional AI review layer that evaluates changes before commit. It learns from feedback and runs entirely in your environment. You bring your own API key, and select the model to use, so you control the costs.
The two feel complementary to me: Scryer gives you the visual understanding of what changed architecturally, Caliper enforces that it stays compliant. If you're exploring this workflow, we're actively looking for alpha testers.
https://getcaliper.dev has docs and the install is just `npm install --save-dev @caliperai/caliper && npx caliper init`.
I ran into a related problem at scale though: deterministic linters catch syntax and obvious anti-patterns, but they miss the harder stuff. Logic bugs that pass all the rules. Spec drift. Security issues hiding in otherwise "clean" code.
So I built Caliper to layer AI review on top of deterministic checks. It reads your coding conventions (CLAUDE.md, .cursor/rules, whatever you use) and compiles them into checks that run after each AI turn—free and instant. Then optionally an AI layer evaluates changes against your project's actual policy. Catches what linters structurally can't.
Very different approach from Vibecheck but complementary. Actively looking for alpha testers if you want to try it — https://getcaliper.dev
Two points of anecdata from that experience:
- The students believe that the path to a role in big tech has evaporated. They do not see Google, Meta, Amazon, etc, recruiting on campus. Jane Street and Two Sigma are sucking up all the talent.
- The professors do not know how to adapt their capstone / project-level courses. Core CS is obviously still the same, but for courses where the goal is to build a 'complex system', no one knows what qualifies as 'complex' anymore. The professors use AI themselves and expect their students to use it, but do not have a gauge for what kinds of problems make for an appropriately difficult assignment in the modern era. The capabilities are also advancing so quickly that any answer they arrive at today could be stale in a month.
FWIW.
There are some good tools out there for automating pr review; IMO, they don't catch enough, and they catch it too late.
I've been experimenting with some ideas about a very opinionated AI code reviewer, one that makes an ideal tradeoff between cost and immediacy (eg, how soon after composition does the code get feedback).
Currently in an invite-only alpha, but check out the landing page and lmk if you'd like to be a trial user!
They had to run it for a few years before they realized CS kids who did poorly in the class dropped the major - the implicit signal being "you don't know how to think like a computer scientist".
I've been obsessed with making it easier to handle tab overload in the browser without requiring any sort of active "tab management".
I have a working extension that replaces the "new tab" page with a clean view of all open tabs, along with simple ways to search and select which tab to switch to, including search over bookmarks and history. There are also some simple tools to allow for creating and reorganizing tab groups.
For a small group of people, it revolutionizes the browser experience. I'm still trying to decide if there is a widely-useful product there, or if it's just a niche use case.
Any and all feedback welcome!
As a software engineer, my favorite book on the topic is "Unix: A History and a Memoir" by Brian Kernighan. It tells many of the same stories that are told elsewhere about the modern history of the Labs, but through the first-person perspective of someone who was there, for whom the characters are not just historical figures, but friends.
With a well written objective / key result (ex: "grow DAU by 30%"), you can abandon your entire roadmap two weeks into the quarter and still hit your OKRs. They enable you to respond to new information and lessons learned, rather than locking the whole team into a rigid plan for the entire quarter.
Ctrl-T opens a new tab page, <tab> highlights the search bar, and then I get instantaneous search over open tabs, bookmarks, and history. Everything stays 100% local.
Therefore, any offer should be backed by logic of the form "we want to hire this candidate for $Y or less. We will make an initial of $X (X < Y), so that we have room to negotiate upwards when they ask."
In my experience, usually Y==(X * 1.1).