HNHacker News
TopNewBestAskShowJobs

tosh

188,728 karma · joined May 4, 2010

https://smolenv.com (a smol agent)

https://findableapp.com (SEO toolkit for Google & ChatGPT)

https://kiwilang.com (k-like language and implementation in Zig with support for GPU via Apple MLX)

https://chesscatsapp.com (a fun way to play chess)

https://moreepisodes.com (tv show episodes generated by GPT-4)

https://jamshelf.com (open source Clubhouse)

https://magic.do (early stage fund)

https://lemmings.io (sci-fi themed hackathons (think "zombie apocalypse", "aliens"))

https://applesilicongames.com (game compatibility and game performance on Apple Silicon Macs)

https://devmonthly.com (curated news & input & jobs for software engineers)

https://blossom.io (project tracking for distributed teams)

more from me around the web:

https://twitter.com/__tosh

https://github.com/tosh

https://angel.co/tosh

https://medium.com/@__tosh

https://lobste.rs/u/tosh

https://dribbble.com/tosh

https://linkedin.com/in/tschranz

https://instagram.com/thomas.schranz

https://facebook.com/thomas.schranz

{UX Service Game} Design, Typography, Clojure, Lisp, Python, Tea, Minimalism, Dao, …

[ my public key: https://keybase.io/tosh; my proof: https://keybase.io/tosh/sigs/gBG56O339IbBPh1iaudGL3dgJuX2yONIATYv62xHXdg ]

submissionscomments
tosh··on Grok 4.6
> where Grok finally catches up

if the benches hold it did catch up

tosh··on Grok 4.6
gpt 5.6 sol and fable 5 level if the benches hold
tosh··on YC startups are abandoning .com
.com is evergreen but it's more and more difficult to get good .com names

I wouldn't call this 'abandoning .com' though

it's basically still .com dominated with a good chunk .ai

and I guess the ones with good .com are keeping the .com

tosh··on My Agent Setup
makes sense, ty
tosh··on My Agent Setup
ty for the writeup!

being able to talk to each of the agents via dm (but also in group chats) sounds interesting

does that mean that you have 1 chat per domain specific agent? can you also start multiple sessions/threads or is that not part of the way you interact with them currently?

tosh··on llama.cpp
ty for digging this up!
tosh··on llama.cpp
I was a bit suspicious of the url but it is also listed on llama.cpp github

https://github.com/ggml-org/llama.cpp

tosh··on Jolt: Clojure compiler implemented with Chez Scheme
it's good to have even more (and more slim!) ways to run Clojure

(not just JVM and JavaScript runtimes)

tosh··on Claude Code is leaking real email address as a User-Agent string in curl command
is this confirmed? this is a github issue with little context and only one comment
tosh··on Show HN: Ante, a coding agent in a single binary that runs offline
also no need to trust these bench runs, you can just run your own

(any OpenAI Responses API compatible endpoint works, if your endpoint does not support 'custom' tools you can have your agent change the smol implementation to use 'function' tool implementation instead)

that said: be aware that smol does not come with any system prompt and does not load agents.md files by default

some older not so strong models benefit from a system prompt and guidance in agents.md that complements them

that said 2: system prompt and or loading agents.md automatically is easy to add though if you want it

tosh··on Show HN: Ante, a coding agent in a single binary that runs offline
if you're interested in an open source agent that uses a minimal amount of CPU and RAM and has source that is easy to audit:

https://github.com/smol-env/smol

here are traces from an agentic task around using duckduckdb

comparing CPU and RAM usage of the whole container over time w/ OpenCode, hermes, pi, codex, smol

https://x.com/__tosh/status/2086882367126286466

https://x.com/__tosh/status/2086882204060160350

smol is very minimal only using stdlib (in this case it is the go version but you can also take a look at implementations in python, clojure, php)

tosh··on Learning more about Claude's mathematical capabilities
interesting!

it sounds like this might benefit from a ui that helps to edit/re-write the history

I also think this would make sense for programming but there it is a bit harder to justify the effort

when you are working on something that really matters this can make the difference though

ty for sharing!

tosh··on Learning more about Claude's mathematical capabilities
prompt engineering 2025: you are an expert programmer, use industry best practices, test driven development and use modularity and abstraction to anticipate future features, …

prompt engineering 2026: i believe in you

tosh··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
good to see new open weights releases from meta
tosh··on Ask HN: What are you working on? (August 2026)
that is by design

intentionally moving the security boundary to where it should be (the environment)

tosh··on Ask HN: What are you working on? (August 2026)
ty for the feedback!

i will also provide ungolfed versions

tosh··on Ask HN: What are you working on? (August 2026)
I like the racket+deepseek combo

ty for sharing this + your books

tosh··on Ask HN: What are you working on? (August 2026)
smol, a smol agent harness in 9 lines python

https://smolenv.com

also has minimal implementations in clojure, go, php:

https://github.com/smol-env/smol

cheaper, faster, less peak RAM per task than OpenCode, Hermes, Codex, pi:

https://smolenv.com/t/nested-template-includes-60636/

https://smolenv.com/t/duckdb-sessionization-23811/

exploring what the minimal setup is for an agent to work well

tosh··on Show HN: A Project Oberon System version running on RISC-V instead of RISC-5
https://en.wikipedia.org/wiki/RISC5
tosh··on Show HN: A Project Oberon System version running on RISC-V instead of RISC-5
what does RISC-V instead of RISC-5 mean?
tosh··on Jolt: A Clojure compiler implemented on top of Chez Scheme
I just found it!

I think Jolt is done mainly by yogthos and a few contributors

tosh··on Managing AI Coding Costs at Scale
for complex open ended coding tasks better models are better

(and mid-to-long-term, often also short-term end up cheaper than weaker models)

this might change soon if we are reaching a certain capability threshold

but right now that's still the case

unless you are working on throw-away trivial stuff where iteration speed and trying many speculative things might give you an edge

tosh··on Managing AI Coding Costs at Scale
> Using harnesses that are “less chatty” (more token efficient), or tuning existing harnesses to generate less token overhead.

agree, this works, undervalued!

look at minimal agents that protect the context window:

  - pi (https://github.com/earendil-works/pi)
  - smol (https://github.com/smol-env/smol)
some thoughts on the other tips (for coding):

1) stronger models are more token efficient for open ended tasks because at the limit …

  - stronger models can solve tasks that the weaker models can not solve
  - stronger models make fewer mistakes, compose things better (cli, abstractions, …)
  - navigate the code base better
  - are better at removing and simplifying the code base again
that of course is difficult to benchmark, so most attention goes to simple benchmarks that show cheaper models can get similar results on 'closed' tasks with easy to 'eval' results

2) dynamic request and task routing sounds great/obvious but is very very hard

  - to benefit from caching you don't want to switch model or inference endpoint
  - to _know_ a certain request can be routed to a weaker/cheaper model needs good context and a strong model to get right and often is still unknowable because the active coding session can go many ways and turn from trivial to challenging in a few turns, always in motion is the future, if you get it wrong you are back in the problem space of #1
using cheaper models and auto-routing do work well for 'closed' tasks where you have something repeatable and can evaluate whether a certain quality threshold is reached that you are comfortable with

for open ended coding sessions it is not so easy

that said: cheaper does not have to mean weaker, you want to look at the pareto frontier and stay up to date on new good models

there are many models like deepseek v4 flash and luna that are both cheaper and way better than most other models

tosh··on DeepSeek V4 Flash 0731
results comparable to gpt 5.6 luna but cheaper

promising!

tosh··on AI psychosis is the new leadership blind spot
does this work better than putting a bag of popcorn into the microwave?
tosh··on Prime Agent: A self-improving RLM agent
i'm also not a fan of javascript and typescript (let alone the ecosystem they come with)

but the runtimes are everywhere

and a smol agent /w no 3rd party dependencies is probably a good option to have in a pinch

tosh··on Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users
free unlimited luna is a pretty badass move

luna is very good

tosh··on Kitesurf: The agent-first browser that runs in V8 isolates on Cloudflare Workers
> we opted for native Rust whenever possible and to compile directly to WebAssembly
tosh··on Prime Agent: A self-improving RLM agent
ty for the pointers

I like the babashka angle!

tosh··on Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
the way to avoid these problems is not to hope for the user or the agent never to make mistakes

it's designing the environment and invariants so whole categories of failures can not happen at all

the agent ui nagging the user for approval is a ux anti-pattern, we already know how well this works for operating system permission dialogues

← PreviousPage 3 of 34Next →