if the benches hold it did catch up
188,728 karma · joined May 4, 2010
https://findableapp.com (SEO toolkit for Google & ChatGPT)
https://kiwilang.com (k-like language and implementation in Zig with support for GPU via Apple MLX)
https://chesscatsapp.com (a fun way to play chess)
https://moreepisodes.com (tv show episodes generated by GPT-4)
https://jamshelf.com (open source Clubhouse)
https://magic.do (early stage fund)
https://lemmings.io (sci-fi themed hackathons (think "zombie apocalypse", "aliens"))
https://applesilicongames.com (game compatibility and game performance on Apple Silicon Macs)
https://devmonthly.com (curated news & input & jobs for software engineers)
https://blossom.io (project tracking for distributed teams)
more from me around the web:
https://twitter.com/__tosh
https://github.com/tosh
https://angel.co/tosh
https://medium.com/@__tosh
https://lobste.rs/u/tosh
https://dribbble.com/tosh
https://linkedin.com/in/tschranz
https://instagram.com/thomas.schranz
https://facebook.com/thomas.schranz
{UX Service Game} Design, Typography, Clojure, Lisp, Python, Tea, Minimalism, Dao, …
[ my public key: https://keybase.io/tosh; my proof: https://keybase.io/tosh/sigs/gBG56O339IbBPh1iaudGL3dgJuX2yONIATYv62xHXdg ]
if the benches hold it did catch up
I wouldn't call this 'abandoning .com' though
it's basically still .com dominated with a good chunk .ai
and I guess the ones with good .com are keeping the .com
being able to talk to each of the agents via dm (but also in group chats) sounds interesting
does that mean that you have 1 chat per domain specific agent? can you also start multiple sessions/threads or is that not part of the way you interact with them currently?
(not just JVM and JavaScript runtimes)
(any OpenAI Responses API compatible endpoint works, if your endpoint does not support 'custom' tools you can have your agent change the smol implementation to use 'function' tool implementation instead)
that said: be aware that smol does not come with any system prompt and does not load agents.md files by default
some older not so strong models benefit from a system prompt and guidance in agents.md that complements them
that said 2: system prompt and or loading agents.md automatically is easy to add though if you want it
https://github.com/smol-env/smol
here are traces from an agentic task around using duckduckdb
comparing CPU and RAM usage of the whole container over time w/ OpenCode, hermes, pi, codex, smol
https://x.com/__tosh/status/2086882367126286466
https://x.com/__tosh/status/2086882204060160350
smol is very minimal only using stdlib (in this case it is the go version but you can also take a look at implementations in python, clojure, php)
it sounds like this might benefit from a ui that helps to edit/re-write the history
I also think this would make sense for programming but there it is a bit harder to justify the effort
when you are working on something that really matters this can make the difference though
ty for sharing!
prompt engineering 2026: i believe in you
intentionally moving the security boundary to where it should be (the environment)
i will also provide ungolfed versions
ty for sharing this + your books
also has minimal implementations in clojure, go, php:
https://github.com/smol-env/smol
cheaper, faster, less peak RAM per task than OpenCode, Hermes, Codex, pi:
https://smolenv.com/t/nested-template-includes-60636/
https://smolenv.com/t/duckdb-sessionization-23811/
exploring what the minimal setup is for an agent to work well
I think Jolt is done mainly by yogthos and a few contributors
(and mid-to-long-term, often also short-term end up cheaper than weaker models)
this might change soon if we are reaching a certain capability threshold
but right now that's still the case
unless you are working on throw-away trivial stuff where iteration speed and trying many speculative things might give you an edge
agree, this works, undervalued!
look at minimal agents that protect the context window:
- pi (https://github.com/earendil-works/pi)
- smol (https://github.com/smol-env/smol)
some thoughts on the other tips (for coding):1) stronger models are more token efficient for open ended tasks because at the limit …
- stronger models can solve tasks that the weaker models can not solve
- stronger models make fewer mistakes, compose things better (cli, abstractions, …)
- navigate the code base better
- are better at removing and simplifying the code base again
that of course is difficult to benchmark, so most attention goes to simple benchmarks that show cheaper models can get similar results on 'closed' tasks with easy to 'eval' results2) dynamic request and task routing sounds great/obvious but is very very hard
- to benefit from caching you don't want to switch model or inference endpoint
- to _know_ a certain request can be routed to a weaker/cheaper model needs good context and a strong model to get right and often is still unknowable because the active coding session can go many ways and turn from trivial to challenging in a few turns, always in motion is the future, if you get it wrong you are back in the problem space of #1
using cheaper models and auto-routing do work well for 'closed' tasks where you have something repeatable and can evaluate whether a certain quality threshold is reached that you are comfortable withfor open ended coding sessions it is not so easy
that said: cheaper does not have to mean weaker, you want to look at the pareto frontier and stay up to date on new good models
there are many models like deepseek v4 flash and luna that are both cheaper and way better than most other models
promising!
but the runtimes are everywhere
and a smol agent /w no 3rd party dependencies is probably a good option to have in a pinch
luna is very good
I like the babashka angle!
it's designing the environment and invariants so whole categories of failures can not happen at all
the agent ui nagging the user for approval is a ux anti-pattern, we already know how well this works for operating system permission dialogues