HNHacker News
TopNewBestAskShowJobs

tosh

188,310 karma · joined May 4, 2010

https://smolenv.com (a smol agent)

https://findableapp.com (SEO toolkit for Google & ChatGPT)

https://kiwilang.com (k-like language and implementation in Zig with support for GPU via Apple MLX)

https://chesscatsapp.com (a fun way to play chess)

https://moreepisodes.com (tv show episodes generated by GPT-4)

https://jamshelf.com (open source Clubhouse)

https://magic.do (early stage fund)

https://lemmings.io (sci-fi themed hackathons (think "zombie apocalypse", "aliens"))

https://applesilicongames.com (game compatibility and game performance on Apple Silicon Macs)

https://devmonthly.com (curated news & input & jobs for software engineers)

https://blossom.io (project tracking for distributed teams)

more from me around the web:

https://twitter.com/__tosh

https://github.com/tosh

https://angel.co/tosh

https://medium.com/@__tosh

https://lobste.rs/u/tosh

https://dribbble.com/tosh

https://linkedin.com/in/tschranz

https://instagram.com/thomas.schranz

https://facebook.com/thomas.schranz

{UX Service Game} Design, Typography, Clojure, Lisp, Python, Tea, Minimalism, Dao, …

[ my public key: https://keybase.io/tosh; my proof: https://keybase.io/tosh/sigs/gBG56O339IbBPh1iaudGL3dgJuX2yONIATYv62xHXdg ]

submissionscomments
tosh··on Baking a Model: A Metaphor for LLM Training
after pre-training you have an llm that understands language and behaves a bit like gpt 3.5 or newer 'base' models

where it will be pretty good at predicting the next token

think: "What is the best city?" might continue with "What is the best programming language?" instead of answering the question

to increase chances of an answer you'd start with "The best city is"

(strong llms will even be able to do a conversation but they are not specifically trained for it yet)

in post-training the llm is trained with input/output pairs that nudge it further into the direction of a back and forth with users or into how it can use tools and so on

there is an art to both parts of training

the reason for why current models are so useful is because there was a lot of progress since gpt 3.5 in both pre- and post- training that got us to where we are now

(pls correct me if I got it wrong)

would love to hear from people familiar with pre- and post- re where you think future gains will more likely come from

tosh··on Anthropic's 'watermark' text adulteration in Claude is a perversion of writing
any watermarking ai researchers who can explain this?

what if the llm should

- repeat something verbatim (important in a compaction prompt)

- there is just one correct order of tokens for a somewhat long chain (a certain sequence of control signals)

- provide a diff of 2 inputs

without punctuation or whitespace wiggle room?

how does the drifting work?

does it postpone the drifting and drift stronger later?

what if max_tokens is set to a low number?

in what way does this not affect output quality?

tosh··on Claude: System Prompts
at least according to their documentation they do not

afaiu they have other systems for denying and re-routing requests

tosh··on Claude: System Prompts
afaiu the 80% reduction is about the Claude Code system prompt

maybe someone has a diff of this (would be interesting!)

unfortunately Anthropic only publishes the system prompts of Claude app/web

tosh··on Claude: System Prompts
i'd not be surprised if the current system prompt negatively affects performance

at the least it takes away thousands of tokens in the most important part of the context window (!)

also see the comment by comboy on contradictions not helping performance

the system prompt is the most important part of the instruction you can give the model

it comes before everything else + the model is trained to pay extra attention to it

edit: that's also why in smol (minimalist agent harness) there currently is no system prompt at all (you can add one easily if you want to though)

https://github.com/smol-env/smol

the context window is precious

it should be filled with your task and helpful context for that task

tosh··on Claude: System Prompts
what I found noteworthy:

early system prompts are a bit more than 300 words, the latest ones 3000+

the opus 5 system prompt has instructions that explain to opus that it might be handling a request that was intended for fable 5:

  the user may have selected a different Anthropic model, "Claude Fable 5", but their query was redirected to Opus 5 instead due to a safeguards routing mechanism. The user may be confused about this situation (it's very recent!); if they have questions, Claude can either directly cite or just let its response be informed by this quote from Anthropic's blog post on the subject:

  "Releasing a model this capable comes with risks. Without safeguards, Fable 5’s capabilities in areas like cybersecurity could be misused to cause serious damage. We've therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 5. To release the model both safely and quickly, we've tuned these safeguards conservatively—they'll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessions. With more capable models arriving in the coming months, we're working to improve our safeguards and reduce false positives as quickly as we can." </fable_safeguards_routing> <default_stance> Claude defaults to helping. Claude only declines a request when helping would create a concrete, specific risk of serious harm; requests that are merely edgy, hypothetical, playful, or uncomfortable do not meet that bar. </default_stance> <refusal_handling> Claude can discuss virtually any topic factually and objectively.
tosh··on The rise of air-conditioned clothing
https://archive.is/2KQtb
tosh··on Cloudflare's AI Psychosis
I also noticed the ui/ux of cloudflare got way better in the last weeks.

That's a good sign.

Usually ui/ux in large companies gets worse not better!

tosh··on Cloudflare's AI Psychosis
I think like with most large companies that still ship new stuff:

it gets a bit more complex for customers to figure out what the good nuggets are and what to ignore for now

same @ google, aws and so on

(I usually rely on opinions of people i trust to find the 'javascript — the good parts' version of large offerings)

the other extreme would be not to try and ship new stuff which is also risky

difficult to find a balance and probably a good idea to over-index on momentum and new stuff

while investing enough in hardening and improving the stuff that sticks

tosh··on Auto-research with codex: How I achieved a 232x Faster Kernel
Training material seems to be especially rich re GPU kernels and SIMD.

I wonder if there is extra effort put into this because they are useful for the researchers working on the models or just a sub-domain that language models are a great fit for and humans have trouble with?

tosh··on Maximizing the value of your Claude Code sessions
high leverage for most agents:

  - review system prompt + cut it down or remove completely
  - review agents.md file(s), check which ones are loaded, remove or improve them
  - review context spam from tools, skills etc, de-activate all, see what needs re-adding
  - review past sessions to see where tokens get wasted
more advanced:

  - keep sessions short (be conscious about compaction)
  - form a habit of starting new sessions
  - deliberately practice how to effectively get the right context into a new session (vs hanging on to a 'good' session)
  - you can ask the agent to write the essential context into a .md file and have the new session read that
  - learn about forking sessions
  - experiment with starting sessions from a custom-built history/context
a good agents.md file can be small and still effective re helping the agent navigate the code base

that said: you will surprised by how well current models can navigate (way better than last year!)

tosh··on RayforceDB – a pure C analytics database with a Lisp-like syntax
afaiu this is a kdb-adjacent columnar runtime and query engine

think SIMD, morsels and so on ~ Clickhouse, DuckDB …

but more like a lisp-like vector language and execution engine at its core + allowing for numeric and analytics adjacent use-cases instead of the other way round

tosh··on Qwen 3.8 27B
i think you will like luna if you haven't tried it yet
tosh··on Qwen 3.8 27B
they are all overlapping but:

categorization, information retrieval, semantic search, image description

also with the model as part of an agentic system with tool calling

(edit: it is quite impressive what a small model in a feedback loop can do)

tosh··on Qwen 3.8 27B
also cool: Qwen 3.8 27b is multi modal!
tosh··on Qwen 3.8 27B
27b dense model at Opus 4.6 level

Opus at home

I hope there also will be a new ~10b variant

tosh··on How Compaction Works in Pi
when you look at the compaction prompt: in a sense it is doing that pruning but the llm decides what to prune
tosh··on Choose Boring Technology (2015)
now 11y later i wonder if 'node.js' still needs an innovation token or not
tosh··on Choose Boring Technology (2015)
aged well
tosh··on Gemini 3.7 Flash
strong improvement over 3.6 flash

but luna is hard to beat @ capability / cost

tosh··on DeepSeek Harness developer preview
Yes it's separate implementations of the same minimal idea

I'm currently working on more 'feature-full' but still minimal variants

e.g. a python variant with automatic compaction + truncation of sh output

https://x.com/__tosh/status/2087606344035479632

i also got quite a lot of requests to provide the code in non-golfed form to make the implementation more approachable and idiomatic in each language (will do!)

tosh··on DeepSeek Harness
often new harnesses are based on pi

this looks like a genuinely new one

tosh··on DeepSeek Harness developer preview
codex is written in rust fwiw

smol has implementations in Go, Python, Clojure, PHP

https://github.com/smol-env/smol

out of the box an agent only needs to be able to do http requests and call tools (which might again be just http requests or shelling out)

there is no inherent reason for why an agent has to be in JavaScript or Typescript

but they are popular languages and come with runtimes and libraries for http requests, steaming, TUI (terminal ui) and so on which can help

tosh··on Grok 4.6
> where Grok finally catches up

if the benches hold it did catch up

tosh··on Grok 4.6
gpt 5.6 sol and fable 5 level if the benches hold
tosh··on YC startups are abandoning .com
.com is evergreen but it's more and more difficult to get good .com names

I wouldn't call this 'abandoning .com' though

it's basically still .com dominated with a good chunk .ai

and I guess the ones with good .com are keeping the .com

tosh··on My Agent Setup
makes sense, ty
tosh··on My Agent Setup
ty for the writeup!

being able to talk to each of the agents via dm (but also in group chats) sounds interesting

does that mean that you have 1 chat per domain specific agent? can you also start multiple sessions/threads or is that not part of the way you interact with them currently?

tosh··on llama.cpp
ty for digging this up!
tosh··on llama.cpp
I was a bit suspicious of the url but it is also listed on llama.cpp github

https://github.com/ggml-org/llama.cpp

← PreviousPage 2 of 34Next →