Engineering for Bounded Cognition
shapeofthesystem.com
shapeofthesystem.com
For me personally, that means setting up 'attention getters' for the important things in life - 'totems' that force a context switch. For AI agents, it means well-designed CLI tools that help the agent orient itself in a task and pull exactly the 'context-for-the-job' it needs right then.
This is exactly what makes building modern GenAI decision-support systems so difficult. It's no longer just about finding the right software abstractions. You now have to account for the unknown cognitive construct of a completely different intelligence.
This is exactly what I am trying to solve and I have what I call smart repositories that demonstrates this at
https://github.com/gitsense/smart-ripgrep
https://github.com/gitsense/smart-codex
The issue I am finding is, getting the agent to pull what it needs, even when the data is there is still challenging since LLMs are trained on blind discovery where the pattern is:
grep -> read -> grep -> read ...
What is working for me now is thanks to Pi (pi.dev). I am working on a pi-brains extension that makes it dead simple to control the lifecycle for an agent so if I detect that it uses `rg` without `gsc rg`, I can block the agent and inject a steering message that says always search with context.
I can also see if they try to "read" without first looking at the files metadata and so forth.
I'm finalizing things right now, but I think pi with my brains extension should allow domain experts to better guide agents so they can find what they need, when they need it.
But your LLM training corpus covered that, right? /i
Tons of self-help books could be summarized in those 2 lines.
> to fit there "evidence"
> useing
> unoticed
> it's own right
I guess the connection between reading/listening, comprehension and retention ability on the one hand, and language generation ability on the other, isn't as strong as I'd been assuming till now.
But at some level context engineering is very similar to what this article talks about.
I was also confused by the tag menu - I thought they are sections in the article - I completely failed here :(
By the way I have a page with some more sources for the context degradation phenomenon: https://zby.github.io/commonplace/notes/agent-context-is-con...
That's funny, isn't it the same for dogs?
What I've heard is human short-term memory can hold seven things at once. Fortunately the mind is much more.
https://news.ycombinator.com/item?id=48706307
Even if it were written by hand, it’s a very poor and frankly stupid essay about an interesting topic. “The model's attention is a fixed quantity, and it has to add up to one, so the more things you make it look at, the less of that attention any single earlier thing can keep.” This is borderline gibberish and it outright rejects the interesting question about LLMs and attention, namely that they have very different capacities from us. LLMs can read an entire OpenAPI schema in seconds and immediately construct valid requests from it. The article first points this out, and then switches to arguing that LLMs have similar limits to us. It’s completely incoherent.
> But an unbounded queue isn't a safety margin, it's a debt that keeps compounding [1]
before getting a headache.
I wish the whole thing was written better, because the idea of designing a codebase for humans and our limitations sounds fascinating. It's why I personally love type systems, you can keep less things in your head and let the type checker alert you of any possible errors.
[1]: https://shapeofthesystem.com/posts/2026/05/10/the-queue-that...
I would please urge you to read further into the manifesto itself but would also recommend you start at the foreword so you can understand the reason for the use of AI assistance in my writing.
The actual criticisms you have about the content however, I'd like to challenge:
The "adding up to one" is just a simplified gloss over softmax. It's very possible it reads poorly, and thats on me - not LLM gibberish.
As for the incoherence - I have to totally disagree. You have merged the 2 things the post keeps apart - capacity and attention over it. That a model can swallow a schema and write code is a competence humans share. We have been doing it for decades. Besides, the claim was never about us sharing capacity (other than stating it is always bounded) - it was about our attention failing in eerily similar ways.
So, AI slop, no. AI assisted, absolutely. It's sad that some judge the "who" more important than the "what" - especially for this kind of writing. But it's fair feedback nonetheless. I'll see what else I can do with assisting my delivery.
Maybe you just write like an LLM.
Please read https://marcusolang.substack.com/p/im-kenyan-i-dont-write-li...
To me this is optimization problem. How can we solve a problem if we don't understand it? Understanding takes a lot of effort exactly because our minds are wandering through useless context all the time, and not to mention interruption.
I formulate this problem as:
> Optimize for understanding
I know how to approach solving this exact problem. In fact, I've been doing exactly this since March 2026. We need to figure out how to isolate problems and context around them. And so my best bet right now is using graphs. Links can be easily added or removed between two nodes. And context is simply a group of links and/or linked nodes.
Now. What exactly is "understanding"?
To me, this is process when we look at some unpredictable, chaotic system and then creating structure from it. The chaotic system is an entangled, spaghetti-like graph. The ordered one is one we [hopefully] have in our brains, which allows us to act on it. I don't want to repeat entire article I wrote about this so if you want you can find it on my recent project (it's not ready for HN prime-time yet but I'll post demo soon).
But tl;dr we have "chaotic/unknown graph" and "structured/understood graph", and the bottleneck is moving nodes and links between these two.
The faster we understand the world around us, the better we understand why problems appear, and how to fix them. And once I realized this, to me it became clear where we need to move forward: to connect everything together in a way that makes understanding quick.
And fun fact, I already did this: I connected my article to yours.
Basically, when I was using my first version of Coherence to write down what system does on lowest level I found myself imagining in my head database records. And this is exhausting when you have more than 2 tables to think about.
And so I don't want to imagine them, I want DB records represented as entities I can simply render onto screen and look at them while I think how they should interact with each other.
Now, this is topology + domain modeling. Topology/semantics here is describing what happens with those domain models over time, what states they, how they transition, what attributes models might have, and most important — which code symbol describes all of that.
So for example, I am making Job Tracking System. I need to persist Job record. I visualize it as a row in table and I can model it without attaching to code. And then we can say:
- "this entity (job) --(has_attribute)--> status", or
- "job -(described_by)-> class JobRecord",
- "job -(observed_in)-> trace",
- "job -(verified_by)-> job_test.rb".
This gives us useful semantics / description / context / domain knowledge + links it to code + test that confirms model shape, or system state transition, or process, or product behavior.
And you can describe it at any level you want.
If you don't know the implementation details yet, a greenfield project? Well, just start with Product level specs, high level description what users will see.
Is it brown-field legacy monolith 500k LoC? Well, just reverse-engineer intent from code (I even have guide how to do this and I did it on 10k LoC successfully), then align people on that "intent graph", and then connect it into CI/CD so that it becomes default review artifact on EVERY single pull request.
And the key here is to isolate context around domain model we're interested in. Make it minimal but useful to make decisions. And make it easy to pull more [relevant] context, or remove [unnecessary] context.
I still don’t fully know how to avoid tunnel vision once you have only one slice of the graph and models in front of you. E.g. failure mode might be somewhere not even described by a link or something exists entirely outside of the graph.
This would be the next thing to solve.
And meanwhile I'd like to thank you for showing me this problem.