HNHacker News
TopNewBestAskShowJobs

brainless

2,589 karma · joined February 11, 2010

Hello, I am Sumit. I live in a little Himalayan village in India.

Software engineer for 17 years across multiple startups. Led teams in the US, Germany and India.

I focus on tiny/small LLMs, build my own products and work on a couple consulting gigs. I do not read or write code manually anymore.

- https://github.com/brainless

I have given up city life and hustle culture. I share my home as a co-living space, mainly for artists and digital nomads. I run Curry Hostel:

- https://www.instagram.com/curryhostel

Socials:

- https://meet.hn/city/in-Kolkata - https://linkedin.com/in/brainless

submissionscomments
brainless··on Coding is not solved
What I mean is that a lot of the popular languages have an engine written in C. Fairly common pattern. I get it that we want to abstract but sometimes it feels we just want to avoid dealing with a language that is one layer below.

Now, LLMs can.

brainless··on Coding is not solved
How many of the popular languages are running C underneath? Count them, I will wait.
brainless··on Coding is not solved
If a language is Turing complete, why build another language right on top? Most of the interpreted languages are running an engine written in C. Was C not enough? Or was it that we needed abstraction?

It is nice to be in our bubble but if what you said is true, why do we only have like 3 major instruction sets? Why not 40? Programming languages are just easy to create - this is a good thing and we created too many. But it absolutely fractured the industry.

Look at the average job post. There are requirements for system design - good. Algorightms - good. Then it goes into this whole language/framework land. Remembering syntax is literally there in many companies interview. These are preferences. If we understand how a computer executes binary/assembly, the layers above are not as relevant as we have made them.

brainless··on Coding is not solved
I have personally wanted to learn and actually use Lisp. I will get to it, even with all the LLM enabled programming. But yes, this is one of those things.

We have reinvented so many ideas many different ways in many different languages and frameworks simply because we want our flavored version.

brainless··on Coding is not solved
I have been programming for about 30 years (including school years). Professionally for 18 years.

Can anyone tell me why we have 40 or more programming languages, with about 10 popular ones? Then about 20 frameworks in each of them. And add another 200 popular libraries for each language? This matrix make no sense till you realize - it is preferences all the way down.

Most of us engineers have built our own mental model of programming. We are all right. But the users do not care. LLMs are here to produce code closer and closer to the metal as needed. They can sit and create a graph out of every spec, use an AST that they develop and run on the CPU if they have to. They will do it. No amount of us discussing will stop that.

Programming is going to be re-invented. I do not think the current ways to write software will even matter.

brainless··on Ember-1
There are lots of distilled models on huggingface that are much smaller than, say, Opus. They are distilled from Opus or Fable and show clear improvements. I do not have the budget to fine-tune a 30b or more parameter model but from my tiny model experiments, the results are quite clear. Again, I have only a couple small tests.

Have you actually fine-tuned yourself? Email categorization comes to mind and there are tons of non-LLM approaches even that will give fantastic results. How did spam filters work before LLM?

I think LLMs just made us think that is the only way. It is not.

brainless··on Ember-1
You can go quite far using a human language to Bash grammar based setup but at some point the input prompts are harder to translate. The OP has existing projects that work with AST quite deeply so I assume they know about that already.

I am building a natural language to CSV/Excel commands for a "wrangler" type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.

https://github.com/brainless/baho

brainless··on Ember-1
I have been trying a mix of fine-tuning and I am amazed that most people do not see this coming.

A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.

I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.

brainless··on What About Rails?
I think LLMs may actually help us get to CLI/API driven development. At least, that has been my experience.

Even though most of my projects have a UI, I build a CLI/API version so that the LLM can interact with it directly. I have been using this "CLI driven development" approach for more than a year now and have had fantastic results. The CLI arguments make it easy for LLM to interact with software it wrote.

I usually ask LLM to build a lib, then expose as a CLI and a RESTful API.

brainless··on Reverse-engineered Jev-like model
I came across this recently. I was scanning for tiny models from HF using their search API. The script was generated by an agent. When I ran it, Qwen 3.5 did not make it at the top. Turns out, models generally prefer older content (training) but that the scanner also did not give any importance to recency.
brainless··on Nvidia announces native GPU programming in Rust
I was learning Rust slowly when the LLM enabled coding became good enough. I switched from learning to full on building with Rust. I still learn high level concepts as needed but I will not be able to write Rust on my own at all.

And that sounds scary but the way I got over the fear is by realizing there are many things that I do very well but I do not know their internals very well. Driving is an example. I barely understand what the steering wheel, clutch or brake pedals do. I have driven over 130,000 Kms and I will perhaps drive more than double that in the next many years.

I have been building software since PHP/Drupal days. Got into AWS S3 as a beta user. Adopted Memcached (and MQ) in 2008 out of necessity. Then Python/Django for 10 years. Then Rust. And tons of JS/TS. I owe a lot to my curiosity. I believe we can keep learning what we need and still delegate most of programming to agents.

brainless··on Introducing System One Models and Jev
I am not an expert in this domain but as an engineer-turned-researcher, this looks a lot like GliNER with a fitting harness.

This is something I focus on in a bunch of my experiments - how to get immense value out of tiny models (<1b params). There are lots of different architectures out there and there is so much to optimize if you know what you are asking and have a grammar to constrain with.

Great to see this and I hope this is a lot on top of what is already openly available.

brainless··on Training a 3.8B LLM to 0.384 CORE for $998
Yes they are great for learning at own pace, trying out new things.

I have accepted two things that make me a happy engineer now: AGI is not here no matter what they say and LLMs are still very useful if one knows how to use them.

They are another layer of abstraction and like you said they do not tire. There is a lot of optimization needed so we can reduce wastage (running 1T+ LLMs for most work is wastage).

brainless··on Training a 3.8B LLM to 0.384 CORE for $998
More and more such experiments. I felt sad for a couple months when I realized that writing code will not be the same since. Now I am on the other side.

LLMs are interesting in their own ways but as an engineer, this is a way to unlock a new way of building software.

I recently build a Claude-assisted Excel/CSV parser for a US based property management system (tax compliance). Uses Haiku and has a lot of deterministic code to extract column/row combinations to check known formats and finally handing out the headers to Haiku to give us a translation plan to our support columns.

These would eventually become part of the software, in a tiny LLM. The gap between training (such tiny LLMs) and inference will shrink. We can consult Claude for edge cases, create sample dataset and train a the tiny LLM on demand so we go to Claude less.

The tooling that a project needs is really important. Something I have been feeling as well. Not just in LLM building projects, but regular software projects that are LLM generated.

brainless··on Training a 3.8B LLM to 0.384 CORE for $998
When we hyper focus on finding something, we find it all the time.

The world of tiny LLMs is so interesting. It is unlocking novel ways to encode information. Why focus on the style of writing instead of the subject matter?

brainless··on 27.5KB language-agnostic WebGPU syntax highlighter
This is a tiny LLM doing all the heavy-lifting. Any mention of the training process? I am obsessed with tiny LLMs and the do-one-thing-really-well approach that they seem to fit very well.
brainless··on My local model setup on an M4 Pro Mac Mini
I experiment a lot with local LLMs, particularly small ones like Qwen3.5 4B and 9B. I have build multiple experiments to make harnesses that use these models for code generation, planning, local search, etc.

These are really good models but the harness has to be built around them. I have a ton of generated system prompts for specific purposes. Even parts of a SolidJS stack, for example Route management, has its own prompt. These are experiments but the results are real. If we build harnesses around small models, we can build a locally running WYSIWYG editor which works on plain text prompts.

The performance, in simple tokens/second, is not the most important factor. For many private data points, like emails, I would rather have a local graph based search and LLM on top where the harness is specific to problems like calendar, contacts, finance, etc.

I run all experiments on an 16GB M4 Mac Mini but coding agents building the harness are a mix of Codex, Claude Code and opencode.

brainless··on Matrox: Graphics for Professionals
Somewhere around early 2000, I used to work part-time with a local video editing agency. Matrox video capture cards were the really high bar, other than Avid. At least that is what I remember. We had a couple of Matrox capture cards, expensive for our tiny budget. Other than that and a few 3D cards also around that time, I forgot about Matrox.
brainless··on Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
I do not think Cloudflare was a less-than-peers optimized product when they launched. This is one of their blog posts which describes taking one aspect even further.

I think Cloudflare became big only because they were so much more optimized than others that they offered some services for free that others were not offering. If running costs are high, you only burn (VC) cash and then you exit.

brainless··on The Harness Is the Thing
I use Claude Code, Codex and opencode pretty much interchangeably. I am currently using Claude more this month because (stupidly) I paid for Max ($100) since I have a large client project.

I generally use larger models to plan. All my generated Epics have similar structure. All my repos have similar structure (https://github.com/brainless/akar and https://github.com/brainless/daftprompt are recent examples).

I barely spend time or thought in making prompts. I have a simple text file with a few combinations. They refer all the common files (README, AGENTS, DEVELOP, etc.)

All reference software is cloned locally and the docs mention that. The prompt templates then boil down to research mode (write Epic) or worker mode (write software) or review mode (leave review notes in Epic). That's it.

Many of my harness experiments are about text manipulation, text search, graph on text. Because that is what LLMs are - text processing systems. Cut parts of prompts, cut parts of response, cut parts of user's intent. Join, break into epics/tasks, run with LLMs, repeat.

brainless··on Worst-case glacial lake flood scenarios in a transboundary Himalayan basin 2022
I can understand that this is not the exact location where the floods actually happened. But I am sad and I am frustrated and I am angry.

We had floods in Sikkim, which is practically a "stone throw" away from a world-level map point of view. I was there then. There was past research but everyone was so surprised. Activists talking about that event 10 years prior. Published reports.

Yet, we do nothing to move people away. To pause and think what the hell is going wrong. Just keep hustling and pretend shit is not going to go wrong. I am tired of our* culture.

our: all of us, I am not only talking about India, Nepal or China here.

brainless··on Nvidia agrees to acquire Hugging Face for $13B
I am on the same boat, I live in a small Himalayan village. I have 2x5G based devices and one Wireless bridge (Ubiquiti LiteBeam M5) for a local broadband. Generally I get 50 Mbps, sometimes up to 100 Mbps
brainless··on Nvidia agrees to acquire Hugging Face for $13B
HF is the default platform to find models. Not just ones published by big labs but also a lot of the distilled or fine-tuned versions, etc.

If Nvidia buying HF makes it tough for all the diverse models on HF, then what are some alternatives?

It seems models are the best things to be available on a Torrent platform? Of course HF is much more than just the files but perhaps the metadata can be separate and hosted on multiple community platforms.

brainless··on I were 17, I'd learn how to build LLMs from scratch
I already do this. I live in a small village where there isn't even a wired Internet connection (wireless only). I work full-time with LLMs, on own product ideas and client projects (all LLM led).

I started investing in farms, have 50 pigs and 100+ chickens now. We are planning to grow to 100 pigs and 2000 chickens in a year. We will start growing Shiitake mushrooms in a few months too.

brainless··on Models Are Getting Dumber on Purpose
This is how my experiments go. And I am sure there are popular agents that do this. How I am trying is to create "Rust Engineer", "Typescript Engineer" or even "Rust Diesel Engineer". I have not tried fine-tuning. I focus on a small model, usually Qwen3.5 9B. I take a bunch of open source repositories and build a KG on it. A small model should be able to enrich your prompt and add technical context. The final, enriched prompt goes to the capable model.
brainless··on Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
I am sorry I did not understand all of it. But, would this allow running large MoE LLMs on a local network with experts spread out over multiple cheaper GPUs (or even CPUs)? This would perhaps be more useful than over the Internet, within offices for example.
brainless··on Launch HN: Bullet (YC S26) – A Faster Coding Agent
Good to see more harnesses coming out. I think the initial set of "features" that made into harnesses like tool calling, multi-turn chat, MCP, skills and so on can all be optimized. And then much more can be done on top.

I am trying out a two-model approach where small model has access to tools, large model does not. Small model shapes prompts from the repo graph. And repo graph is the only tool that small model has when reading. The small model is already given a set of context from git log, codebase and Markdown/text files (generally design files) depending on the user's prompt.

I do not want to use use multi-turn chat. Small model would instead create fresh prompts for the larger model feeding context and reshaping the original ask every time.

Also, reference repositories can be added for small model to help ask right questions. Once a plan is made by large model, execution is mostly task-by-task, done by small model. Lots of deterministic code doing all this orchestration.

brainless··on Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
This is awesome. I will take some time to dig in. When I am not working for client(s), I focus entirely on tiny LLMs - I have specific approach to prompting, avoid multi-turn chat and build harness to fit the selected LLM as closely as possible.

My experiments are in https://github.com/brainless/

I will be happy to share what I learn.

brainless··on Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)
Any plans to support smaller models? I have a M4 Mac Mini with 16GB unified memory and an RTX 3060 (Laptop) with 6GB VRAM. My own product experiments all revolve around small models and harness around them. Happy to contribute.
brainless··on Position: LLMs Can't Jump
I have a weird thought experiment: If you give a GPT-2/3 level LLM tools to search the internet - any document, can it build bigger, better LLMs?

You may think this is not a good test because an older (or say a smaller) LLM can study from the knowledge on the Internet and build. But we are like that - we can access the Universe through our senses.

Can we ever produce anything that is beyond this Universe? I think an LLM that is lacking in knowledge can build more complex systems as long as it can access more data.

Page 1 of 20Next →