89 karma · joined August 9, 2023
Propelcode looks awesome! So cool to see different people and perspectives shaping this space.
Historically, plan mode served two different roles:
1. making the agent’s instructions precise enough to execute 2. helping the human understand what was about to happen
I think #1 is less necessary as agents get better. #2 is going the other direction, it becomes more important as the model is able to do more on its own because larger chunks of work are happening with increasing complexity.
Where I’ve changed my mind is the interface for #2. I increasingly think an interactive, iterative workflow is closer to how people actually build understanding than being handed a long generated document, especially one they didn’t author themselves.
The human-understanding problem is very real though
1. giving the agent sufficiently precise instructions 2. helping the human understand what’s being built
I think #1 is increasingly going away as models get better. #2 is a separate problem, and I actually think it matters more as agents get more capable. I’m not arguing that human understanding should disappear with plan mode.
I’m also not pushing this on behalf of a model lab, I don’t represent one. These are lessons I learned from my own mistakes building Nuanced (https://www.nuanced.dev). We built around a very explicit plan-oriented workflow because that’s how I used to work. I spent years at GitHub writing ADRs, RFCs, and design docs, and I still really like writing because it’s how I clarify my own thinking.
What changed for me is that AI-generated plans don’t give me that same effect. The agent reads between the lines, generates a lot of detailed prose, and now I’m parsing decisions and assumptions I didn’t actually make.
So I still think human understanding deserves a first-class primitive. I’m just increasingly unconvinced that a big blob of generated text is the right one, especially as we move toward many agents working in parallel.
but I think I overfit the interface to how I used to work without agents. when I was at GitHub, I wrote a lot of ADRs, RFCs, and design docs, and I liked that because writing is how I clarify my own thinking. with agents though, I’m often not doing the writing myself. I give the model a rough intent and it fills in a bunch of gaps, and then I get back a long, polished plan containing decisions I didn’t explicitly make.
That’s the part that feels broken to me. The plan can be detailed and technically correct, but still be hard to review because the important bits are buried and feel distant from my own thinking. All the assumptions and tradeoffs and questions may or may not be legitimate, but it’s hard for me to get into flow state and carefully check them.
So I still want the collaboration step. I’m just less convinced that generated prose à la plan mode is the right interface for it.
In the interim, this is a test I did with Sonnet 3.5 + Cursor, showing how it impacted explanations (not solutions): https://github.com/nuanced-dev/nuanced/issues/31
In the interim, this is a test I did with Sonnet 3.5 + Cursor, showing how it impacted explanations: https://github.com/nuanced-dev/nuanced/issues/31
With context windows limited to 200K tokens, cramming in random files isn't just inefficient, it's impossible for large codebases. If you’re debugging a failing test, you only need to understand the relevant files in the call chain. It's not about more context, it's about relevant context. That's what Nuanced provides through static analysis and machine learning.
As AI writes more code, we need better tools to trust it and technologies that ensure our human understanding keeps pace with this rapid development.
While everyone else races to ship new features with AI, we're focused on addressing gaps in AI coding tools and ensuring those features are reliable and maintainable rather than code that works today but becomes a liability tomorrow.
We're starting with an AI-powered Python language server that makes AI-generated code more reliable by understanding your entire system—using a deeper semantic understanding of code than LLMs have today, but also artifacts outside of code such as commit histories, configs, and team patterns.
We're a team of ex-GitHub engineers and researchers who've scaled some of the world's largest developer platforms. I'm Ayman (https://www.aymannadeem.com/about/), and before founding Nuanced, I spent seven years at GitHub where I helped build Semantic(https://github.com/github/semantic), an open-source library for parsing and analyzing code across languages—and scaled security systems to detect anomalous code patterns across millions of repositories. Our team’s deep experience in static analysis and large-scale system design shapes our approach to the AI reliability challenge today.
We've all been on-call at 2 AM, untangling complex service dependencies, and more recently, we've seen firsthand how AI accelerates development—both the wins and the wounds.
If you're building an AI coding tool and any of this sounds interesting to you—we should talk!
Read more at https://nuanced.dev/blog/the-reliability-gap