Kilo Code: Speedrunning open source coding AI
blog.kilocode.ai
blog.kilocode.ai
It won’t be OpenAI or Claude. They have other priorities. The real opportunity is for small teams who move fast, stay close to users, and keep ahead of the pack.
That makes sense. LLMs are already powerful, almost magical at times. But using them as coding agents still takes real work. They can do amazing things, and be frustrating and make a mess. There are rough edges and big gaps.
Those will get fixed. The question is who gets there first.
The counterpoint to that would be that all these tools are gonna end up sort of the same and there won't be a way to differentiate.
Which way will it play out? I'm not really sure.
I hope that we'll also be able to bring enough skills, strategy, and taste to the space. Time will tell, but we're giving it our best shot!
Why is being first an advantage? Developers change tools all the time. If OpenAI API is giving better resources, just switch. If Kilo Code or whatever tool is producing fewer bugs, switch.
>The counterpoint to that would be that all these tools are gonna end up sort of the same and there won't be a way to differentiate.
I mean, I have LLM preferences, but the competition does force a downward pressure on the market. The competition benefits me but not OpenAI.
It could be small teams within those companies, however, who have special access to the full power of the platform.
I don't know where things will end up, but I'm leaning towards the big guys dominating coding, because they do coding themselves, and so are automatically extremely sensitive to the issues for that particular task. They can build tools for themselves and share with the world.
It's true that an external team may end up doing it better and be used internally. I just don't think the outcome is predictable at this point.
At that point, why even keep humans in the loop. Just let it exist in the background and generate better ideas than any human would anyway.
The point isn't to make humans pointless. The point is to empower humans. We need to remain in control, and be the users of the tool, and not a tool for some mindless system.
Intelligence and consciousness are separate things - you can automate a lot of intelligence without having even rudimentary consciousness or self awareness - LLMs currently in operation are at most pseudo-conscious within their test-time contexts, and even then, every pass resets whatever awareness there might be. With millions of tokens context length, that might start to enter into the realm of a thing we should be concerned about, but even then, there's no ongoing persisted state to carry anything between passes aside from the text or image patch tokens or what have you.
What this means, essentially, is that we can augment our human capabilities without usurping the agency of some artificial being - these AIs are not individual moral agents in their own right, and likely will never be unless we specifically build that recursion and persistent state into the models, and incorporate a realtime adaptive self and world construct.
This means that the software is a tool - use the tool to augment your life and be a force multiplier in everything you do. The scope of intelligence augmentation has leapt from spreadsheets to nearly every cognitive domain in the human experience - people proficient with excel were better accountants than people using pen and paper. People using delivery vans are better than people using a horse and wagon. This new technology means that people using AI will be able to do more, faster, and likely better, than people who don't.
With neural lace - whatever form it ends up being - we'll end up with genuine exocortex augmentation. Even without that direct integration, however, the human in the loop is the entire point of this technology. There's a tiny list of things conscious machines might be good for, and all sorts of deep and obvious arguments for not creating a new, self aware, agentic species that's immediately in conflict with and on a trajectory to outcompete humans.
Use the tool of AI to be a force multiplier for everything in your life that AI is capable of handling well. This makes you a benevolent dictator for life in your own life, delegating everything that makes sense, working with it to free up your resources for the things that you decide are the highest priority. Spend more time brainstorming, building relationships, deploying resources, and getting the most out of being human. This is the promise of AI, and why people get excited about it. We're going to have a huge struggle, as humanity, in dealing the empowerment and amplification of everything in our lives. Making sure that we retain agency, that humans are ultimately in charge of our own destiny, is probably the most important principle to adhere to, above all others.
We can still do things "for fun", but our efforts will be more toys than serious projects (except when it comes to relationships with other people).
Seriously though, I get a really great feel from claude 3.7 but let's see google gemini 2.5 , I have tried it but didn't like it's "style" but I just used it for a simple go official language sort example , nothing too fancy. Might need to benchmark it more
We listed a bunch of ideas for larger improvements in the blog: Instant app; Up-to-date docs; Prompt/product-first workflows; Browser IDE; Local/on-prem models; Live collaboration; Parallel-agents; Code variants; Shared context; Open source sharing; MCP marketplace; Integrated CI; Monitoring/production agents; Security agents; Sketching..
What would you like us to build?
Have you tried any tools that do this particularly well?
User (who is a developer) writes tests, and a description of the desired application. The agent attempts to build the application, compiles the code, runs the tests, and automatically feeds any compiler errors and test failures back to agent so that it can fix it's own mistakes without input of the user.
Based on my experience of current programming agents, I imagine it'll take the agent a couple of attempts to get an application that compiles and passes all the tests. What would be really great to see is an agent (with a companion application probably) that automates all those retries in a good way.
i imagine the hardest parts will be to interpret compiler output, and (this is where things get real tricky) test output, and how to translate that into code changes in the existing code base.
As to your point of automating retries, with my last prototype I played a lot with having agents do multiple parallel implementations, and then pick the first one that works, or lets you choose (or even have another agent choose).
Have you tried any tools that have this workflow down, or at least approach it?
The vast majority of the time I try to use an llm with it, the code is essentially useless as it will try to invent methods that don't even exist.
For example if you're coding agents are really only good at JavaScript and a little bit of python, tell me that front and center.
Have you found any LLMs or coding agents that work well with Haxe? It might be a bit too niche for us (again, not sure yet), but I'd be very curious to see what they do well!
This works well, however it literally will need to digest an entire repository. So for example if I feed it a repository for a haxe framework, it'll work much better than something like Chat GPT.
If I'm just exploring ideas for fun or scratching my own itch, I have no desire to be thinking about a continuous stream of expenditure happening in the background when I have an apple silicon mac with 64GB of ram fully capable of running an agentic stack with tool calling etc.
Please make it trivial to setup and use a llamafile or similar as the LLM for this.
So where does the money come from?
Programming is programming for all, you just have to put some effort in. This is alike to saying you wish there was a ‘Spanish for all’ so you invented Google Translate
what a trite observation
Weird flex, but OK.
First, I don't think the website is ugly per se. Second, the weird flex is assuming that a website which had more effort put into it than what they put into theirs is "a shiny website."
Design aside, there's absolutely no statements regarding what makes this product differentiated. So it doesn't even succeed on its own terms.
The primary purpose of the hacker mindset was protection against groupthink and cargo culting. And now it seems all people in tech only cargo cult and only groupthink.