HNHacker News
TopNewBestAskShowJobs

tosh

189,014 karma · joined May 4, 2010

https://smolenv.com (a smol agent)

https://findableapp.com (SEO toolkit for Google & ChatGPT)

https://kiwilang.com (k-like language and implementation in Zig with support for GPU via Apple MLX)

https://chesscatsapp.com (a fun way to play chess)

https://moreepisodes.com (tv show episodes generated by GPT-4)

https://jamshelf.com (open source Clubhouse)

https://magic.do (early stage fund)

https://lemmings.io (sci-fi themed hackathons (think "zombie apocalypse", "aliens"))

https://applesilicongames.com (game compatibility and game performance on Apple Silicon Macs)

https://devmonthly.com (curated news & input & jobs for software engineers)

https://blossom.io (project tracking for distributed teams)

more from me around the web:

https://twitter.com/__tosh

https://github.com/tosh

https://angel.co/tosh

https://medium.com/@__tosh

https://lobste.rs/u/tosh

https://dribbble.com/tosh

https://linkedin.com/in/tschranz

https://instagram.com/thomas.schranz

https://facebook.com/thomas.schranz

{UX Service Game} Design, Typography, Clojure, Lisp, Python, Tea, Minimalism, Dao, …

[ my public key: https://keybase.io/tosh; my proof: https://keybase.io/tosh/sigs/gBG56O339IbBPh1iaudGL3dgJuX2yONIATYv62xHXdg ]

submissionscomments
tosh··on The Shape of Things to Come
great question, not yet

I think it is worth adding support for it though

tosh··on The Shape of Things to Come
Interesting, that is not my experience (but I'm mainly using them for reading and writing code right now)

but I don't doubt that you're seeing this behaviour, ty for sharing!

tosh··on The Shape of Things to Come
nb: current models (e.g. GPT 5.6 Sol) are very good at long horizon tasks

they no longer need crutches or rube goldberg machines to keep them going

minimal agent harness is just a loop that loops until no more tool calls are coming

GPT 5.6 Sol continues to drive the loop until the task is done or it decides that it wants to present the user with information

at that point it is probably good to not automatically continue (!)

(YMMV of course, for some tasks it makes sense, then you can still add a loop around it + the necessary signals, the main thing I want to say is that what used to be essential to keep models going is no longer needed, current models can do long-horizon tasks way better than when these outer loops where necessary)

self-plug: "smol", is a minimal agent in ~20 lines of Go that implements this pattern (keeps going until no more tool calls):

https://github.com/smol-env/smol

works just fine

tosh··on Cursor removed cost information from the usage page and CSV export
repo is published now: https://github.com/smol-env/smol
tosh··on Cursor removed cost information from the usage page and CSV export
I do think prompting and reference files (e.g. for architecture, tech stack, …) can be extremely helpful. I also do this in my projects (challenge is keeping drift of these documents at bay).

What I wanted to emphasize is that whatever is in the context (whether system prompt or user message does 'steer' the model in a strong way, so everything in the context affects overall performance in a way. Even if it is 'just' net neutral it takes space up in the context window.

The context window is very very precious, everything that goes into it should help (not just hopefully help).

The challenge is coming up with good stuff to put into that context. A good agents.md file will be better context than whatever the popular harnesses have in their system prompt.

Also good to keep in mind that newer models are very good and more agentic than older models so they are better at exploring their environment based on the tasks you give them.

tosh··on Cursor removed cost information from the usage page and CSV export
I agree, only looking at the token burn is not enough

in this case it was 10 tasks and all harnesses could complete the tasks successfully, of course now the question is: will this hold for more and more complex tasks but I had to start somewhere :)

Checksum: Compute a file’s SHA-256 checksum and save the exact digest to an output file.

Log correlation: Correlate nested service logs to identify and summarize a request’s complete execution path.

CSV report: Parse quoted CSV data and aggregate paid orders and exact decimal totals by region.

JSONL join: Join related JSON Lines datasets and produce a correctly grouped and ordered report.

Archive repair: Find the correct version of a corrupted file in a tar archive and restore it.

SQLite migration: Safely migrate a SQLite database schema and verify the resulting data and constraints.

Python bug fix: Repair an interval-merging implementation so it passes visible and hidden edge-case tests.

Python CLI: Implement a robust command-line program that reads JSONL and reports validated statistics.

Multi-file feature: Add an atomic feature across a small Python package, CLI, and associated tests.

Pipeline repair: Fix a Make, shell, and Python reporting pipeline so it handles general input and passes verification.

tosh··on Cursor removed cost information from the usage page and CSV export
coming in a few hours

you can follow this org in the meantime https://github.com/smol-env

or on twitter here: https://x.com/__tosh

tosh··on Cursor removed cost information from the usage page and CSV export
smol currently is very simple so it definitely does less things, like no subagent orchestration

I will look into how token usage looks like for longer sessions and more complex tasks

re caching: the cache ratio for this bench looks 'bad' for smol because it often finishes a task before caching kicks in (caching starts at 1024 tokens)

thank you for flagging this

tosh··on Cursor removed cost information from the usage page and CSV export
agree, that makes it a bit tricky to compare (esp if you also want to add different models and reasoning levels into the mix)

I will add more tasks (esp longer ones) and think more about grading, the current tasks were easy to grade because the desired outcomes are well specced but I will also look into more open ended tasks and how to grade those

thank you!

tosh··on Cursor removed cost information from the usage page and CSV export
will do!
tosh··on Cursor removed cost information from the usage page and CSV export
smol only has 1 tool: sh

the system prompt of smol is shorter than the system prompt of Pi

smol has no system prompt

system prompt of Pi 0.83.0

""" You are an expert coding assistant operating inside pi, a coding agent harness. You help users by reading files, executing commands, editing code, and writing new files.

Available tools: - read: Read file contents - bash: Execute bash commands (ls, grep, find, etc.) - edit: Make precise file edits with exact text replacement, including multiple disjoint edits in one call - write: Create or overwrite files

In addition to the tools above, you may have access to other custom tools depending on the project.

Guidelines: - Use bash for file operations like ls, rg, find - Use read to examine files instead of cat or sed. - Inspect PI_* environment variables for current model and session details. - Use edit for precise changes (edits[].oldText must match exactly) - When changing multiple separate locations in one file, use one edit call with multiple entries in edits[] instead of multiple edit calls - Each edits[].oldText is matched against the original file, not after earlier edits are applied. Do not emit overlapping or nested edits. Merge nearby changes into one edit. - Keep edits[].oldText as small as possible while still being unique in the file. Do not pad with large unchanged regions. - Use write only for new files or complete rewrites. - Be concise in your responses - Show file paths clearly when working with files

Pi documentation (read only when the user asks about pi itself, its SDK, extensions, themes, skills, or TUI): - Main documentation: /usr/local/lib/node_modules/@earendil-works/pi-coding-agent/README.md - Additional docs: /usr/local/lib/node_modules/@earendil-works/pi-coding-agent/docs - Examples: /usr/local/lib/node_modules/@earendil-works/pi-coding-agent/examples (extensions, custom tools, SDK) - When reading pi docs or examples, resolve docs/... under Additional docs and examples/... under Examples, not the current working directory - When asked about: extensions (docs/extensions.md, examples/extensions/), themes (docs/themes.md), skills (docs/skills.md), prompt templates (docs/prompt-templates.md), TUI components (docs/tui.md), keybindings (docs/keybindings.md), SDK integrations (docs/sdk.md), custom providers (docs/custom-provider.md), adding models (docs/models.md), pi packages (docs/packages.md), environment variables (docs/environment-variables.md) - When working on pi topics, read the docs and examples, and follow .md cross-references before implementing - Always read pi .md files completely and follow links to related docs (e.g., tui.md for TUI API details) Current working directory: /workspace """

tosh··on Cursor removed cost information from the usage page and CSV export
in my book anything that is (I'm sure well intentioned) and injected to help the agent — but doesn't help it — is a waste of tokens

but even injected context that when I read it sounds useful can oversteer the model and make it second guess or take a more complicated route than it normally would

(you can see this when looking at traces with and without that injected context)

often harnesses also mention in their system prompt locations of markdown files that the model can consult if the model thinks they might help

that hint alone as part of the system prompt can be strong enough to make the model read in more tokens than would have been necessary

'spam' is maybe a harsh way to say it

unfortunately I don't see an easy way other than to invest time and tokens into finding out which parts of the added context (in system prompt, injected in turns etc etc) are actually helpful or harmful and when

I'm just doing the easiest thing I could think of: start from nothing or close to nothing

that seems to work better than what most harnesses are doing

turns out GPT 5.6 Sol is all you need

tosh··on Cursor removed cost information from the usage page and CSV export
I will look into it more to see if I have configured it wrong but I think the token efficiency hurts cache use as caching only starts at 1024 tokens so for tasks where smol is under or close to 1024 tokens most of them are uncached
tosh··on Cursor removed cost information from the usage page and CSV export
smol is also prefix caching

the uncached tokens are also from runs where smol finished a task below 1024 tokens (the minimum amount of tokens needed to activate caching) which is less tokens than other harnesses are using for their system prompt (!)

> GPT-5.6 and later models: Caching is available for prefixes containing at least 1,024 tokens. This is a strict minimum.

https://developers.openai.com/api/docs/guides/prompt-caching

so in this specific case the count of uncached tokens for smol makes it look worse than it actually is

that said: it does makes sense to add more tasks that are difficult enough to fill the context window to compare the harnesses for how well they deal with compaction

staying below compaction (or with compaction at fewer compactions) is not only cheaper and faster, it also helps the agent stay on track

tosh··on Cursor removed cost information from the usage page and CSV export
I built an ad-hoc custom comparison framework to inspect system prompts, caching behaviour, tool call outputs, exact api requests and responses and so on

I agree fewer tokens is not necessarily better but a bit counter-intuitively often the harness using fewer tokens is not only done faster but has better results

(that said: of course check the results, look at the full traces, agree!)

tosh··on Cursor removed cost information from the usage page and CSV export
ty re --disallowed-tools

for coding agents 'shell' is often all you need (just make sure the environment has the necessary tools)

tosh··on Cursor removed cost information from the usage page and CSV export
the system prompt still takes up useful space in the context window and steers the model into unnecessary actions and over-thinking patterns
tosh··on Cursor removed cost information from the usage page and CSV export
smol is basically this 9 line python agent re-implemented in Go

https://news.ycombinator.com/item?id=49006862

I'll have more about it in the next hours/days, you can follow me on twitter in the meantime (https://x.com/__tosh)

tosh··on Cursor removed cost information from the usage page and CSV export
the tasks were all simple agentic tasks

like creating a checksum of a file, merging csvs and so on, fixing a makefile pipeline

with known 'good' outcomes

all harnesses could reach the outcomes, only cost, time, number of tool uses and so on were different

(Claude Code failed once in 1 task but I think that was just an unfortunate outlier, the tasks aren't that difficult)

tosh··on Cursor removed cost information from the usage page and CSV export
I can only recommend to regularly measure how many tokens a harness+model combination uses for a certain task

There are huge token efficiency/bloat differences between agents while working on the same tasks, using the same model, in the same environment

Yesterday I ran 10 agentic tasks using GPT 5.6 Sol in an ubuntu 26.04 vm a couple of times with different harnesses and got vastly different token usage.

  +-------------+-----------+-----------+-----------+-----------+--------+
  | Harness     | API total | Input     | Cached    | Uncached  | Output |
  +-------------+-----------+-----------+-----------+-----------+--------+
  | smol        |   172,807 |   142,334 |     8,704 |   133,630 | 30,473 |
  | Pi          |   427,211 |   392,767 |   137,216 |   255,551 | 34,444 |
  | OpenCode    | 1,564,429 | 1,523,957 | 1,204,736 |   319,221 | 40,472 |
  | Codex       | 3,005,744 | 2,953,154 | 2,649,344 |   303,810 | 52,590 |
  | Hermes      | 3,856,611 | 3,808,231 | 3,167,232 |   640,999 | 48,380 |
  | Claude Code | 5,073,137 | 5,029,969 | 4,587,008 |   442,961 | 43,168 |
  +-------------+-----------+-----------+-----------+-----------+--------+
https://x.com/__tosh/status/2083593799872237680

I'm not surprised that Claude Code is not optimized for an OpenAI model but I was still quite shocked re how much of a difference the harness makes.

Disclaimer: I'm working on 'smol' which is a minimalist harness but it's really nothing special, just a minimal system prompt, no skills files, only tool is shell

Do not underestimate how much popular harnesses are spamming the context window. The context window is very important.

tosh··on How to Do Great Work (2023)
good related read: You and Your Research (Hamming)

https://www.paulgraham.com/hamming.html

tosh··on Inkling-Small
better than haiku 4.5 smaller than nemotron 3 ultra
tosh··on Advancing the price-performance frontier with GPT‑5.6
luna is way better than haiku 4.5
tosh··on Advancing the price-performance frontier with GPT‑5.6
80% price cut for luna is a very aggressive pricing move

makes it by far the best choice for most workloads that do not need bleeding edge intelligence (reminder: luna can be comparable to opus 5!)

tosh··on Superlogical
more info: https://mitchellh.com/writing/superlogical
tosh··on Superlogical
might be a related sneak peek here: https://x.com/mitchellh/status/2079327969416482859

can't wait for more details!

tosh··on Show HN: Agent in 9 Lines Python
ty for looking into it, i like your http.client take (still stdlib, keeps connection open)
tosh··on Show HN: Agent in 9 Lines Python
oh right, good point @ move outside of the loop (or move directly into the Request), ty!
tosh··on Show HN: Agent in 9 Lines Python
here is a more expanded transliteration:

https://gist.github.com/tosh/61aca9ffa9ea115fa4df332407d7a9a...

tosh··on Show HN: Agent in 9 Lines Python
Great question!

I had a tool description earlier but 'sh' as tool name seems to be sufficient, the agent behaviour was the same.

There might be performance gains if a description is added though, or worth trying different ways of telling the agent about what is available in the environment.

That said, the newer models are fairly good at driving a harness to explore the environment.

← PreviousPage 5 of 34Next →