it's designing the environment and invariants so whole categories of failures can not happen at all
the agent ui nagging the user for approval is a ux anti-pattern, we already know how well this works for operating system permission dialogues
188,895 karma · joined May 4, 2010
https://findableapp.com (SEO toolkit for Google & ChatGPT)
https://kiwilang.com (k-like language and implementation in Zig with support for GPU via Apple MLX)
https://chesscatsapp.com (a fun way to play chess)
https://moreepisodes.com (tv show episodes generated by GPT-4)
https://jamshelf.com (open source Clubhouse)
https://magic.do (early stage fund)
https://lemmings.io (sci-fi themed hackathons (think "zombie apocalypse", "aliens"))
https://applesilicongames.com (game compatibility and game performance on Apple Silicon Macs)
https://devmonthly.com (curated news & input & jobs for software engineers)
https://blossom.io (project tracking for distributed teams)
more from me around the web:
https://twitter.com/__tosh
https://github.com/tosh
https://angel.co/tosh
https://medium.com/@__tosh
https://lobste.rs/u/tosh
https://dribbble.com/tosh
https://linkedin.com/in/tschranz
https://instagram.com/thomas.schranz
https://facebook.com/thomas.schranz
{UX Service Game} Design, Typography, Clojure, Lisp, Python, Tea, Minimalism, Dao, …
[ my public key: https://keybase.io/tosh; my proof: https://keybase.io/tosh/sigs/gBG56O339IbBPh1iaudGL3dgJuX2yONIATYv62xHXdg ]
it's designing the environment and invariants so whole categories of failures can not happen at all
the agent ui nagging the user for approval is a ux anti-pattern, we already know how well this works for operating system permission dialogues
all with the same zero 3rd party dependencies approach
also want to do clojure, unfortunately it looks like java does not come with json support out of the box
- 21 lines of Go
- no 3rd party dependencies
https://github.com/smol-env/smoleasier to add and customize stuff when you start from a small base
think of it as your starter dough
But pulling orchestration off is very very tricky.
Even if it is just a small, simple orchestrator.
Ideas like planner, memory, log, subagents, graphs (each on their own) sound great and very promising.
So promising that one would think they must work, how could they not?
I've been there as well!
The challenge is that all these parts of the orchestrator are intertwined with each other
and they are all causing overhead in the main context window in some form or at least overall complexity that is difficult to grasp and predict/engineer for
(even though the idea is to help exactly with the fact that the context window is limited)
To save context window there is also more communication that has 'stille post' ('chinese whispers') like dynamics
Turns out it is very difficult to find out the right context to bubble up and down.
It's very similar to human org communication challenges (think large org stucture vs small teams vs one person that can keep it all in their head)
Yeah, what do you do if one person can't keep it all in their head?
But how great is it when it's possible?
Companies must have figured out how that works right? Maybe we can adopt and implement these ideas?
And yet … easy it is not, especially when you're not dealing with run-of-the-mill well-defined tasks.
But more like with open-ended software development?
I'm not saying it's not possible or that it should not be tried.
On the contrary, I think this is worth pursuing and a bit like the search for the holy grail.
But I also think the other direction of the search space is under-explored.
The holy grail is glamorous.
With 'smol' I'm spelunking on this other extreme (non-orchestration?)
(welcome, join us, we have cookies, and context windows with a lot of room for work items!)
smol is a minimalist agent harness that protects the context window
- no system prompt
- no tool spamming (just 1 tool: sh)
- no agents.md
- no mcp
- no planning, todos, graphs, beads, …
and figuring out how that looks like and performsit is a worthwhile thread to pull I think
at least from the dozens of benches I'm looking at I see that less stuff in the context window does help a lot
- cheaper per task
- finishing faster
- better tool composition (sh and pipes are great!)
but also for more complicated longer-term tasks the model gets less confused when the context window is not getting spammedthe context window is precious
I also see codex do it that way quite often
and at the same time Opus struggles with using the edit tool in Claude Code even though model and harness are by the same company
no need to start from smol (even though I think it makes a decent starting point)
I'm confident you can understand what it does (and does not do) and why in an afternoon
(probably in 20-30 minutes actually, even if you are not familiar with Go)
the same is way more difficult with larger agent implementations even if you only want to understand the direct implementation ignoring all the 3rd party dependencies that come with them
here is a typical bench run with 9 runs and traces for opencode, pi and smol
you can look at every step and which tools are used and how
https://smolenv.com/t/nested-template-includes-60636/
sh is pretty versatile and composes well
pi also only has 4 tools (good!)
diy ftw
but also highly opinionated (no mcps, no agents.md, no system prompt, …)
so not sure it checks all of your boxes
that said: because smol is so smol you can adapt it easily and agents (including smol) can work well with it because the whole implementation fits comfortabliy into the context window
- which env do you provide?
- which model(s)?
- subagents?
- system prompt (default or custom?)?
- agents.md
etc etcalso you kinda have to look at many runs and study their traces, if you look at too few runs the outcome variability you're drawing from is too high
Also happy with how much love codex gets from OpenAI.
That said: I was looking at existing agents to find one to build upon and to me they were all too complex and were leaning too heavily into 3rd party dependencies.
Nothing I could understand comfortably in an afternoon (that's also on me I guess). Pi was closest to what I was looking for but still too big and too modular.
(It's hard to come up with good abstractions that work well across all major models + keep up with new concepts that come and go all the time with new releases.)
The more complex agents err on the side of supporting many models 'ok' instead of focusing on taking advantage of a specific model.
With a tiny implementation it is easier to adapt it.
Adding new stuff, removing stuff again, changing it from working well specifically with GPT 5.6 Sol to working with the exact model I want.
a few months ago I tried to implement an MCP server from scratch in python (instead of using the existing reference implementation) and I could not get it to work reliably across clients
1) system prompt in pi is quite small (way smaller than the one from OpenCode)
2) when your agents.md file changes pi does not re-spam it (preserves cache, good trade-off!)
3) only 4 tools, every tool comes with a description for how to use it and causes reasoning overhead (fewer tools is good)
all of these things add up
here are pi, opencode and smol working on the same tasks in 9 fresh runs
https://smolenv.com/t/nested-template-includes-60636/
you can step through the traces and see how the system prompt + tools steer the agent in a certain way
with GPT 5.6 Sol you can even get away without a system prompt (see smol) and only 1 tool (sh)
If I have agents.md or other context I want it to read I mention it at the beginning of the session
re MCP: I am not using an MCP with smol
but there are ways to convert MCPs into CLI tools or typed js
I imagine that would work well/more token efficient with smol (or most harnesses actually)
The best thing I found so far re smol is that it fits into the context window with plenty of room to spare
So it is easy to adapt (and add stuff to it, even stuff you only need specifically for just 1 project)
Whereas adapting a more complex harness is more error prone
at least with GPT 5.6 Sol fwiw
https://smolenv.com/t/nested-template-includes-60636/
sh is all you need
I agree, next it makes sense to try more open ended tasks + have humans (and/or multiple models) grade the runs and their results
- same task result (passed)
- finished faster
- fewer tokens, less cost
- fewer requests for inference
- fewer tool calls
- less peak RAMI found that removing
- system prompt
- skills
- agents.md
- mcps
+ reducing tools to just 1 (sh)gives better results than having 'more' of them
(e.g. look at these traces to see more vs less in action:)
https://smolenv.com/t/nested-template-includes-60636/
not saying the right context does not help
(it definitely does!, but it's not trivial to provide the right context)
even with a robust, well thought through setup every model behaves in its own way, some adhere more to a system prompt, another model has more recent cut-off time and knows about new parts in the stdlib
some oversteer, some understeer …
if you want the best performance unfortunately there is not really a way other than to constantly adapt the harness/clutches/context to the model du jour
(it has way less context spam, fewer tools and smaller system prompt than the usual suspects)
databricks also looked at this: https://earendil.com/posts/pi-autoresearch-and-databricks/
actually the opposite: they matter a lot because they do steer the agent
with great power comes great responsibility
try a fresh coding session without skills, agents.md, system prompt and additional tools
I think you will be positively surprised how good current models like GPT 5.6 Sol are when they are not oversteered and context spammed
Here is a task (python templating) with 9 runs with OpenCode, Pi and smol
https://smolenv.com/t/nested-template-includes-60636/
you can read each run step by step and see what the agents are doing and how the system prompt and available tools are steering their behaviour to take longer and higher cost
(disclaimer: I'm working on smol)
install or build stuff that they can then use
it changes the environment instead of the agent/harness but in a sense how separate is the agent from its environment and why do we apply this distinction re self-improvement?
animals and humans do the same thing and are great at it, without 'self-improvement' with emphasis on the 'self'
I think it is worth adding support for it though