HNHacker News
TopNewBestAskShowJobs

tosh

188,895 karma · joined May 4, 2010

https://smolenv.com (a smol agent)

https://findableapp.com (SEO toolkit for Google & ChatGPT)

https://kiwilang.com (k-like language and implementation in Zig with support for GPU via Apple MLX)

https://chesscatsapp.com (a fun way to play chess)

https://moreepisodes.com (tv show episodes generated by GPT-4)

https://jamshelf.com (open source Clubhouse)

https://magic.do (early stage fund)

https://lemmings.io (sci-fi themed hackathons (think "zombie apocalypse", "aliens"))

https://applesilicongames.com (game compatibility and game performance on Apple Silicon Macs)

https://devmonthly.com (curated news & input & jobs for software engineers)

https://blossom.io (project tracking for distributed teams)

more from me around the web:

https://twitter.com/__tosh

https://github.com/tosh

https://angel.co/tosh

https://medium.com/@__tosh

https://lobste.rs/u/tosh

https://dribbble.com/tosh

https://linkedin.com/in/tschranz

https://instagram.com/thomas.schranz

https://facebook.com/thomas.schranz

{UX Service Game} Design, Typography, Clojure, Lisp, Python, Tea, Minimalism, Dao, …

[ my public key: https://keybase.io/tosh; my proof: https://keybase.io/tosh/sigs/gBG56O339IbBPh1iaudGL3dgJuX2yONIATYv62xHXdg ]

submissionscomments
tosh··on Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
the way to avoid these problems is not to hope for the user or the agent never to make mistakes

it's designing the environment and invariants so whole categories of failures can not happen at all

the agent ui nagging the user for approval is a ux anti-pattern, we already know how well this works for operating system permission dialogues

tosh··on Prime Agent: A self-improving RLM agent
I'm adding python in a few minutes and then looking into other languages including javascript and typescript

all with the same zero 3rd party dependencies approach

also want to do clojure, unfortunately it looks like java does not come with json support out of the box

tosh··on Prime Agent: A self-improving RLM agent
here is a take on a smol agent ("smol")

  - 21 lines of Go
  - no 3rd party dependencies
https://github.com/smol-env/smol

easier to add and customize stuff when you start from a small base

think of it as your starter dough

tosh··on Building an Advanced Agentic Harness
edited with link to repo (I'll add more to it over the next days!)
tosh··on Building an Advanced Agentic Harness
I love reading about orchestration concepts.

But pulling orchestration off is very very tricky.

Even if it is just a small, simple orchestrator.

Ideas like planner, memory, log, subagents, graphs (each on their own) sound great and very promising.

So promising that one would think they must work, how could they not?

I've been there as well!

The challenge is that all these parts of the orchestrator are intertwined with each other

and they are all causing overhead in the main context window in some form or at least overall complexity that is difficult to grasp and predict/engineer for

(even though the idea is to help exactly with the fact that the context window is limited)

To save context window there is also more communication that has 'stille post' ('chinese whispers') like dynamics

Turns out it is very difficult to find out the right context to bubble up and down.

It's very similar to human org communication challenges (think large org stucture vs small teams vs one person that can keep it all in their head)

Yeah, what do you do if one person can't keep it all in their head?

But how great is it when it's possible?

Companies must have figured out how that works right? Maybe we can adopt and implement these ideas?

And yet … easy it is not, especially when you're not dealing with run-of-the-mill well-defined tasks.

But more like with open-ended software development?

I'm not saying it's not possible or that it should not be tried.

On the contrary, I think this is worth pursuing and a bit like the search for the holy grail.

But I also think the other direction of the search space is under-explored.

The holy grail is glamorous.

With 'smol' I'm spelunking on this other extreme (non-orchestration?)

(welcome, join us, we have cookies, and context windows with a lot of room for work items!)

smol is a minimalist agent harness that protects the context window

  - no system prompt
  - no tool spamming (just 1 tool: sh)
  - no agents.md
  - no mcp
  - no planning, todos, graphs, beads, …
and figuring out how that looks like and performs

it is a worthwhile thread to pull I think

at least from the dozens of benches I'm looking at I see that less stuff in the context window does help a lot

  - cheaper per task
  - finishing faster
  - better tool composition (sh and pipes are great!)
but also for more complicated longer-term tasks the model gets less confused when the context window is not getting spammed

the context window is precious

https://x.com/__tosh/status/2084985580144722369

https://github.com/smol-env/smol

tosh··on The next chapter of our AI momentum
ty!
tosh··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
can anyone with corporate background decipher if that is good for demis or not?
tosh··on Pi's Minimalism Is Its Advantage
yes but GPT 5.6 Sol is pretty good at editing files via sh (e.g. using python)

I also see codex do it that way quite often

and at the same time Opus struggles with using the edit tool in Claude Code even though model and harness are by the same company

tosh··on Pi's Minimalism Is Its Advantage
nice, very minimal!
tosh··on Pi's Minimalism Is Its Advantage
but I think what I want to say is: diy

no need to start from smol (even though I think it makes a decent starting point)

tosh··on Pi's Minimalism Is Its Advantage
your favorite agent will help you de-golf and analyze the implementation

I'm confident you can understand what it does (and does not do) and why in an afternoon

(probably in 20-30 minutes actually, even if you are not familiar with Go)

the same is way more difficult with larger agent implementations even if you only want to understand the direct implementation ignoring all the 3rd party dependencies that come with them

tosh··on Pi's Minimalism Is Its Advantage
> feeds everything into bash, making it completely pointless

here is a typical bench run with 9 runs and traces for opencode, pi and smol

you can look at every step and which tools are used and how

https://smolenv.com/t/nested-template-includes-60636/

sh is pretty versatile and composes well

pi also only has 4 tools (good!)

tosh··on Pi's Minimalism Is Its Advantage
i guess what i'm saying is: it's not so hard to start with a minimal, well working agent and then only add exactly what you need instead of taking a more complex agent and then trimming it down or configuring it (and staying compatible with its continued development and deps)

diy ftw

tosh··on Pi's Minimalism Is Its Advantage
smol is ~20 lines of Go, uses only stdlib

but also highly opinionated (no mcps, no agents.md, no system prompt, …)

so not sure it checks all of your boxes

that said: because smol is so smol you can adapt it easily and agents (including smol) can work well with it because the whole implementation fits comfortabliy into the context window

https://github.com/smol-env/smol

tosh··on Pi's Minimalism Is Its Advantage
it's very difficult and time-consuming to make good comparisons (especially for open ended tasks) also because there are so many configuration options

  - which env do you provide?
  - which model(s)?
  - subagents?
  - system prompt (default or custom?)?
  - agents.md
etc etc

also you kinda have to look at many runs and study their traces, if you look at too few runs the outcome variability you're drawing from is too high

tosh··on Harness engineering for self-improvement
great point, progressive disclosure is a good pattern for context management
tosh··on Pi's Minimalism Is Its Advantage
I'm glad there are quite a few good open source ones by now.

Also happy with how much love codex gets from OpenAI.

That said: I was looking at existing agents to find one to build upon and to me they were all too complex and were leaning too heavily into 3rd party dependencies.

Nothing I could understand comfortably in an afternoon (that's also on me I guess). Pi was closest to what I was looking for but still too big and too modular.

(It's hard to come up with good abstractions that work well across all major models + keep up with new concepts that come and go all the time with new releases.)

The more complex agents err on the side of supporting many models 'ok' instead of focusing on taking advantage of a specific model.

With a tiny implementation it is easier to adapt it.

Adding new stuff, removing stuff again, changing it from working well specifically with GPT 5.6 Sol to working with the exact model I want.

tosh··on Stateless MCP has recaptured my interest
I'm glad MCP is getting simpler

a few months ago I tried to implement an MCP server from scratch in python (instead of using the existing reference implementation) and I could not get it to work reliably across clients

tosh··on Pi's Minimalism Is Its Advantage
a few examples:

1) system prompt in pi is quite small (way smaller than the one from OpenCode)

2) when your agents.md file changes pi does not re-spam it (preserves cache, good trade-off!)

3) only 4 tools, every tool comes with a description for how to use it and causes reasoning overhead (fewer tools is good)

all of these things add up

here are pi, opencode and smol working on the same tasks in 9 fresh runs

https://smolenv.com/t/nested-template-includes-60636/

you can step through the traces and see how the system prompt + tools steer the agent in a certain way

with GPT 5.6 Sol you can even get away without a system prompt (see smol) and only 1 tool (sh)

tosh··on Pi's Minimalism Is Its Advantage
I use it for smaller changes (no compaction yet)

If I have agents.md or other context I want it to read I mention it at the beginning of the session

re MCP: I am not using an MCP with smol

but there are ways to convert MCPs into CLI tools or typed js

I imagine that would work well/more token efficient with smol (or most harnesses actually)

The best thing I found so far re smol is that it fits into the context window with plenty of room to spare

So it is easy to adapt (and add stuff to it, even stuff you only need specifically for just 1 project)

Whereas adapting a more complex harness is more error prone

tosh··on Pi's Minimalism Is Its Advantage
i also found re context window: less is more

at least with GPT 5.6 Sol fwiw

https://smolenv.com/t/nested-template-includes-60636/

sh is all you need

tosh··on Harness Engineering for Self-Improvement
in this case there was a hidden grader that checked if the implementation was correct (because that was the easiest thing to check), all 3 agents cleared this hurdle in all 9 runs

I agree, next it makes sense to try more open ended tasks + have humans (and/or multiple models) grade the runs and their results

tosh··on Harness engineering for self-improvement
apologies, I should have clarified the 'better' claim

  - same task result (passed)
  - finished faster
  - fewer tokens, less cost
  - fewer requests for inference
  - fewer tool calls
  - less peak RAM
tosh··on Harness engineering for self-improvement
maybe a bit counter-intuitive but:

I found that removing

  - system prompt
  - skills
  - agents.md
  - mcps
+ reducing tools to just 1 (sh)

gives better results than having 'more' of them

(e.g. look at these traces to see more vs less in action:)

https://smolenv.com/t/nested-template-includes-60636/

not saying the right context does not help

(it definitely does!, but it's not trivial to provide the right context)

tosh··on Agent skills that bring team coding standards to Claude Code and Codex
I agree in principle but in practice this is not easy

even with a robust, well thought through setup every model behaves in its own way, some adhere more to a system prompt, another model has more recent cut-off time and knows about new parts in the stdlib

some oversteer, some understeer …

if you want the best performance unfortunately there is not really a way other than to constantly adapt the harness/clutches/context to the model du jour

tosh··on Agent skills that bring team coding standards to Claude Code and Codex
Pi is pretty good actually

(it has way less context spam, fewer tools and smaller system prompt than the usual suspects)

databricks also looked at this: https://earendil.com/posts/pi-autoresearch-and-databricks/

tosh··on Agent skills that bring team coding standards to Claude Code and Codex
I'm not saying system prompts, agents.md, tools, skills don't have their place

actually the opposite: they matter a lot because they do steer the agent

with great power comes great responsibility

tosh··on Agent skills that bring team coding standards to Claude Code and Codex
I know it sounds a bit counter-intuitive but

try a fresh coding session without skills, agents.md, system prompt and additional tools

I think you will be positively surprised how good current models like GPT 5.6 Sol are when they are not oversteered and context spammed

Here is a task (python templating) with 9 runs with OpenCode, Pi and smol

https://smolenv.com/t/nested-template-includes-60636/

you can read each run step by step and see what the agents are doing and how the system prompt and available tools are steering their behaviour to take longer and higher cost

(disclaimer: I'm working on smol)

tosh··on Harness Engineering for Self-Improvement
one form of very effective self-improvement that coding agents do all the time:

install or build stuff that they can then use

it changes the environment instead of the agent/harness but in a sense how separate is the agent from its environment and why do we apply this distinction re self-improvement?

animals and humans do the same thing and are great at it, without 'self-improvement' with emphasis on the 'self'

tosh··on The Shape of Things to Come
great question, not yet

I think it is worth adding support for it though

← PreviousPage 4 of 34Next →