HNHacker News
TopNewBestAskShowJobs

lsb

4,183 karma · joined August 22, 2007

classicist, computer scientist, founder of PoetaExMachina and NoDictionaries, biker around San Francisco

leebutterman@gmail.com

submissionscomments
lsb··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
The context window is limited to 64k or 128k. If you’re using this with a coding agent, there’s going to be a lot of compactions. I found that I had a subagent whose compacted context plus prompts and such was over the window and it errored out in opencode.
lsb··on Markdown Database Pattern
As we have more powerful AI plus agents that can run on your phone, keeping everything as a pile of markdown plus some compute on top makes a lot of sense
lsb··on Qwen3.8-Flash-Next
This is glitch art for text, I love it
lsb··on Three ways to smuggle SQLite into Nix
If this much JSON could fit in a microcontroller’s memory (7.5MB in text), and it’s a performance issue, maybe it’s worth upgrading the JSON parser?
lsb··on Ornith-1.5: From Self-Scaffolding to Self-Improvement
The page has comparisons with Qwen 3.6 27b and I’d love to see comparisons with Qwen 3.8 27b, the newer one is much more capable!
lsb··on Google is making private AI practical with homomorphic encryption
Google is making private AI practical with Gemma4 something that you can run without an Internet connection.

All of the proofs of privacy rely on us getting the math right. All of the privacy from unplugging your internet cable is there by default.

lsb··on Ancient Library – 1,060 Greek/Latin texts, click any word to parse it
Fixed!
lsb··on Ancient Library – 1,060 Greek/Latin texts, click any word to parse it
Neat! Similar to my graduate thesis, NoDictionaries: https://nodictionaries.com/cato/de-agri-cultura/156
lsb··on Who Is America's Homer?
Homer is the Homer of America, unless you believe in fencing off universal properties of mankind into exclusive tribal ownership.

(cf Saul Bellow asking “who is the Tolstoy of the Zulus”, and Ralph Wiley in the Atlantic saying that Tolstoy is the Tolstoy of the Zulus.)

lsb··on Show HN: Visualizing Tiny LLMs from OpenAI's Parameter Golf
Happy to answer any questions :)
lsb··on Healthchecks.io now uses self-hosted object storage
Self Hosted object storage looks neat!

For this project, where you have 120GB of customer data, and thirty requests a second for ~8k objects (0.25MB/s object reads), you’d seem to be able to 100x the throughput vertically scaling on one machine with a file system and an SSD and never thinking about object storage. Would love to see why the complexity

lsb··on Leanstral: Open-source agent for trustworthy coding and formal proof engineering
The real world success they report reminds me of Simon Willison’s Red Green TDD: https://simonwillison.net/guides/agentic-engineering-pattern...

> Instead of taking a stab in the dark, Leanstral rolled up its sleeves. It successfully built test code to recreate the failing environment and diagnosed the underlying issue with definitional equality. The model correctly identified that because def creates a rigid definition requiring explicit unfolding, it was actively blocking the rw tactic from seeing the underlying structure it needed to match.

lsb··on Lena by qntm (2021)
It’s named after the multi-decade data compression test image https://en.wikipedia.org/wiki/Lenna

Buy the book! https://qntm.org/vhitaos

lsb··on Ask HN: How are you doing RAG locally?
I'm using Sonnet with 1M Context Window at work, just stuffing everything in a window (it works fine for now), and I'm hoping to investigate Recursive Language Models with DSPy when I'm using local models with Ollama
lsb··on Tell HN: No continental US flights due to attack on Venezuela
The New York Times has said that the US president has reported capturing the president of Venezuela https://www.nytimes.com/live/2026/01/03/world/trump-united-s...

Source about aviation: primary (I am at an airport now) and also there are no flights going into or out of JFK right now https://www.jfkairport.com/flight-tracker?view=VIEW_DEPARTUR...

lsb··on Lite^3, a JSON-compatible zero-copy serialization format
This is super interesting!

Apache Arrow is trying to do something similar, using Flatbuffer to serialize with zero-copy and zero-parse semantics, and an index structure built on top of that.

Would love to see comparisons with Arrow

lsb··on Why are your models so big? (2023)
My threshold for “does not need to be smaller” is “can this run on a Raspberry Pi”. This is a helpful benchmark for maximum likely useful optimization.

A Pi has 4 cores and 16GB of memory these days, so, running Qwen3 4B on a pi is pretty comfortable: https://leebutterman.com/2025/11/01/prompt-optimization-on-a...

lsb··on Show HN: DSPy on a Pi: Cheap Prompt Optimization with GEPA and Qwen3
Happy to answer any questions you have :)
lsb··on Show HN: Apache Fory Rust – 10-20x faster serialization than JSON/Protobuf
Curious about comparisons with Apache Arrow, which uses flatbuffers to avoid memory copying during deserialization, which is well supported by the Pandas ecosystem, and which allows users to serialize arrays as lists of numbers that have hardware support from a GPU (int8-64, float)
lsb··on Solveit – A course and platform for solving problems with code
fast.ai (some of the authors of this) was transformative for me, and the community was super nice. Cannot recommend looking into this highly enough.
lsb··on Show HN: A store that generates products from anything you type in search
This is halfbakery! I love it!

(For example, a recent half baked idea there is a perpetually burning flag. https://www.halfbakery.com/idea/Perpetually_20Burning_20Flag... )

lsb··on LandChad, a site dedicated to turning internet peasants into Internet Landlords
How are you a landlord if you're paying property taxes?

Once you have everything else set up, you can migrate to a server hosted on your own internet connection. Running your own data center is one of the more tricky parts of the equation, compared to almost-free web hosting for a 10MB site.

You're also just renting a domain name.

lsb··on Show HN: I replaced vector databases with Git for AI memory (PoC)
Interesting! Text files in git can work for small sizes, like your 100MB.

That is what's known in FAISS as a "flat" index, just one thing after another. And obviously you can query by primary key to the key-value store that is git, and do atomic updates as you'd expect. In SQL land this is an unindexed column, you can do primary key lookups on the table, or you can look through every row in order to find what you want.

If you don't need fast query times, this could work great! You could also use SQL (maybe an AWS Aurora Postgres/MySQL table?) and stuff the fact and its embedding into a table, and get declarative relational queries (find me the closest 10 statements users A-J have made to embedding [0.1, 0.2, -0.1, ...] within the past day). Lots of SQL databases are getting embedding search (Postgres, sqlite, and more) so that will allow your embedding search to happen in a few milliseconds instead of a few seconds.

It could be worth sketching out how to use SQLite for your application, instead of using files on disk: SQLite was designed to be a better alternative to opening a file (what happens if power goes out while you are writing a file? what happens if you want to update two people's records, and not get caught mid-update by another web app process?) and is very well supported by many language ecosystems.

Then, to take full advantage of vector embedding engines: what happens if my embedding is 1024 dimensions and each one is a 32 bit floating point value? Do I need to save all of that precision? Is 16-bit okay? 8-bit floats? What about reducing the dimensionality? Is it good enough accuracy and recall if I represent each dimension with an index to a palette of the best 256 floats for that dimension? What about representing each pair of dimensions with an index to a palette of the best 256 pairs of floats for those two dimensions? What about, instead of looking through every embedding one by one, we know that people talk about one of three different topics, and we have three different indices for each of those major topics, and to find your nearest neighbors you want to first find your closest topic (or maybe closest two topics?) and then search in those lower indices? Each of these hypotheticals is literally a different “index string” in an embedding search called FAISS, and could easily be thousands of lines of code if you did it yourself.

It’s definitely a good learning experience to implement your own embedding database atop git! Especially if you run it in production! 100MB is small enough that anything reasonable is going to be fast.

lsb··on Gemma 3 270M re-implemented in pure PyTorch for local tinkering
That’s wild that with a KV cache and compilation on the Mac CPU you are faster than on an A100 GPU.
lsb··on Vendors that treat single sign-on as a luxury feature
Also: this SSO tax is deceptively framed. Many of these services allow one to sign in through, for example, Google, which can count as a single sign on, and many organizations have a mail account, but that isn’t taken into account.
lsb··on What's the strongest AI model you can train on a laptop in five minutes?
This is evocative of “cramming”, a paper from a few years ago, where the author tried to find the best model they could train for a day on a modern laptop: https://arxiv.org/abs/2212.14034
lsb··on Ask HN: With all the AI hype, how are software engineers feeling?
I used Claude Code to navigate a legacy codebase the other day, and having the ability to ask "how many of these files have helper methods that are duplicated or almost but not quite exactly duplicated?" was very much a superpower.
lsb··on Ask HN: Moving a not-for-profit web app off AWS
I’ve been running half-a-billion parameter models comfortably in a web browser, especially with WebGPU, and you can definitely run billion parameter LLMs in the browser. It becomes a heavyweight browser app, but if the main costs are running ML models you can pretty easily serve static files from a directory and let clients’ browsers do the heavy lifting. Feel free to reach out if you have questions, happy to help, I’ve been working on language web apps as well
lsb··on Ask HN: What's the most creative 'useless' program you've ever written?
I made an AI Art clock, using Stable Diffusion to render a 24-hour clock face into a landscape: https://leebutterman.com/diffusion-local-time/
lsb··on Fast B-Trees
Clojure, for example, uses Hash Array Mapped Tries as its associative data structure, and those work well
Page 1 of 33Next →