10,427 karma · joined March 19, 2009
Daybreak Blue is the not the same thing as Daybreak Red, which has a more significant hurdle. I don't know anyone who has gotten access to Red.
Tail risks have always existed in software development. The tail risk of a bug introduced by some dev who quit five years ago is similar to the tail risk of a bug introduced by Claude six months ago. Deal with it by building better visibility into how your systems work. Demand that your agents write good documentation to accompany their code-writing.
If you’re doing it right these days, it means you’re thinking of a much bigger picture and containing downside risks as boldly as you’re expanding the frontier of upside opportunities.
I don't know why the models were generally trained to be so brief, but it's definitely not the way anyone I know actually writes. A second pass is always a good idea to clean this stuff up. And, thankfully, the models are all pretty good at that.
Nice work. It’s awesome to have these capabilities so close at hand and so trivially easy to integrate with.
sometimes with residual connections, but we can ignore that for sake of simplicity.
Build the systems around the code and let the agents do their work.
But reading code? What does that accomplish, other than to slow your dev process down enormously? Serious question.
Canadians are right to be skeptical that the next presidency will be significantly better than the current one. A serious amount of goodwill has been burned and I think the average view up north is that it's time to build a very solid backup plan.
Do you have an architectural explainer yet for Jev or are you holding that close to your chest and letting the magic rip for now?
If you work at a SaaS of any kind, I think it's worthwhile considering what things will look like when scale is the only thing that is really defensible anymore.
OpenAI and Anthropic have recently hired legions of "forward-deployed engineers" specifically to help companies do this automation work. It's a solid move. And, if you look at some recent product announcements, they are also hard at work building the necessary plumbing. For instance, the OpenAI Agents API lets you, "Build and run cloud agents with the Codex harness, fully managed by OpenAI."
This kind of enterprise-ready, cloud-hosted stuff really accelerates implementation of AI workflows within large organizations. Not every company is in the tech space (not by a long shot). Slop isn't the primary concern. Accuracy and reliability is the primary concern, and beyond that, just the capacity to actually make the changes happen.
But here's the thing: As we use agents to automate previously manual processes, we are elevating the humans to do work that is less amenable to automation. And the surface area of that work keeps expanding because the competitive market we exist in demands it of us.
To stretch an analogy, businesses are like organisms in a pond. A new nutrient (agentic LLMs) was recently added to the pond that makes business organisms more efficient and able to eat new kinds of food and explore new areas. As a result, those organisms that do the extra exploring and consuming grow much faster than their peers who do not. At the end of the day, the new nutrient will just be part of the pond and the old kind of organism will be a fossil.
It also had access to our internal docs (Confluence), JIRAs, Slack conversation history… All via MCPs. So it could dig around to its heart’s content as would a human developer trying to figure out the same problem.
I did also add a Vectorize database (the whole thing is Cloudflare hosted behind zero trust OAuth) as a second step and that can be helpful in surfacing concepts via the atlas’ MCP interface.
I strongly recommend trying this approach out yourself. The recipe is not rocket science. Get your coding agent to take a first cut at building the atlas itself, and then manually correct it. Once you’re happy that it got things right, put an MCP on it or a CLI or whatever. And your LLMs will know what to do from there.