Smol Developer
github.com
github.com
Realistically, a large code base LLM generation tool is going to look something like old-school C code.
An initial pass will generate an architecture and a series of independent code unit definitions (.h files) and then a 'detail pass' will generate the code for (.c files) for each header file.
The header files will be 'relatively independent' and small, so they fit inside the context for the LLM, and because the function definition and comments 'define' what a function is, the LLM will generate consistent multi-file code.
The anti-patterns we see at the moment in this type of project are:
1) The entire code is passed to the LLM as context using a huge number of tokens meaninglessly. (you only need the function signature)
2) 1-page spaghetti definition files are stupid and unmaintainable (just read prompt.md if you don't believe me).
3) No way of verifying the 'plan' before you go and generate the code that implements it (expensive and a waste of time; you should generate function signatures first and then verify they are correct before generating the code for them).
4) Generating unit tests for full functions instead of just the function signatures (leaks implementation details).
It's interesting to me that modern language try to move all of these domains (header, implementation, tests) into a single place (or even a single file, look at rust), but passing all of that to an LLM is wrong.
It's architecturally wrong in a way that won't scale.
Even an LLM with 1000k tokens that could process a code base like this will be prohibitively slow and expensive to do so.
I suspect we'll see a class 'generative languages' emerge in the future that walk the other direction so they are easier to use with LLM and code-gen.
Quite possible you might get a similar things for other existing languages as extensions.
A better comparison might be ML module signatures. (Though to be fair, IIRC they enable higher-order modules instead of just helping the compiler.)
I’m not sure what you’re imagining, but in this case I would not imagine you would generate js or write .d.ts files?
An LLM pass with a high level goal would generate a file list, then a series of .d.ts files from it.
Then (after perhaps a review of the type definition files, possibly also LLM assisted) a second pass taking the .d.ts files as input would generate a typescript file for every .d.ts file.
You would then discard the .d.ts files and just have a scaffolded .ts code base?
My point was doing the same trick with say, Java, seems like a harder problem to solve, but you could do the above right now with existing LLMs.
Edit: as this is LLM thread - ML is Meta Language as in OCaml and SML.
https://twitter.com/verdverm/status/1655481985685389313?t=d7...
We are doing code generation for our users (analysts & data scientists), and I found going by method names & type signatures failed pretty hard, while contextual code snippets (example docs, usages, etc you get from a vector DB) worked quite well
An old program synthesis labmate did quite well pre-genAI by using type signatures, but my takeaway is neurosymbolic for more context + option pruning, not being terse for supporting denser summarization & composition. That'd be better... But maybe training would need to change for that to work, or something deeper?
Separately, I do agree it's interesting to think about what IDEs and langs and dev will look like 5 years from now.. smarter LLMs, bigger context windows, and supporting ecosystems change a lot...
…but I can say with complete confidence that generating code from a coherent set of c header files works better than generating “one shot” full code files just from a high level goal and a file name.
You’re basically reducing the problem space from literally anything to a subset that has an defined structure.
(It helps that #include literally maps to a flat single header in c and having comments in c headers is typical, perhaps?)
How much effort is it worth investing in domain-specific tooling / languages / effort for codegen? On some level it's a bet against LLMs getting better (unless the work has intrinsic value that extends beyond being a workaround for LLM limitations).
I could see a world where if you create the right architecture then complex tasks can be broken into smaller individual tasks where your only concern is the outcome and not the underlying code. Very deterministic
Essentially all the things we developers care about might not matter. Who cares if the LLM repeats itself? DRY won’t apply anymore because there might not be a reason to share code!
LLM go brrrr until it gets the right output and the code turns into more “machine learning black box” stuff
that said if you see my future directions notes i do think theres room for file specific .md instructions.
the shared dependencies file is essentially a plan. i didnt realize it at the moment but now looking at it with fresh eyes i can do a no-op `smol plan` command pretty trivially.
My experience has been that as the amount of code you need to generate scales beyond trivial levels, the prompt-to-code approach from cross file code starts to fail.
ie.
> the shared dependencies (like filenames and variable names) we have decided on are: {shared_dependencies}
Is too trivial of a prompt for large scale code generation.
What if you generate a service in one file and expect a controller to call it from another file?
How does the model know what the functions in the other file are? Did you put the signature of every public function in shared_dependencies?
If you don't have a planning phase, you need a priori knowledge of the order in which files need to be generated, what the functions are called...
I just... don't believe what you're trying to do is actually possible. The models will just generate an isolated 'best guess implementation' for each file; and that might hang together for small code blocks... but as the number of inter-dependencies between blocks of generated code increases...
There's just no way right?
This only works if the code you're generating doesn't call itself, it only calls known library code.
shared_dependencies only solves this for the trivial case of like, 'oh, I have these 3 shared functions / types between all these files'; as the shared types and functions increase, your shared_dependencies becomes an unmaintainable monster... you end up having to generate the shared dependencies... split it up so that you have a separate shared_dependencies per file to generate... and suddenly it starts looking a heck of a lot like a .h header file....
more broadly i see the smol dev cycle being dependent on intelligence and latency, both of which will improve over time (https://twitter.com/swyx/status/1657962184935370752)
lets say right now a single run gets you 1% of the way there. you do 10 runs of prompt iteration before you feel its run out of use, and take over, but at least by then youve gotten 20% of the way.
as intelligence goes up, the single run % goes up. as latency goes down, more runs become feasible. this is the basic equation for how much smol developer can consume the start phase of any project, especially smol/oneoff/internal apps.
this also perhaps opens up a fun definition of agi - when you can one shot 100% an app.
We do pass the whole files, not just headers although that’s a possibility we considered and may try in the future. Looking at other code helps the LLM a lot in maintaining similar style and making sure the code is being used correctly. Sad reality is that interfaces are rarely descriptive and robust enough to code against them without looking at details.
We don’t pass the entire codebase because even small projects don’t fit, we use embeddings and some GPT assistance to decide which files are more likely to be relevant to complete the task and pass those. It doesn’t get it right 100% of the time, (we’re working on it) but it does most of the time.
Our approach allows us to write and edit several files with a single user provided prompt. It can build entire features in existing codebases as long as they’re not huge. A lot of the time it looks like magic honestly.
The link is https://kamaraapp.com if someone is interested in trying it and providing feedback I’m happy to send over some free credits :)
That's because language designers are hitting against fundamental limitations of programming directly in the serialization format. That is, plaintext code. IMHO, this is a dead-end path, already hitting diminishing returns, because it's the equivalent of designing binary formats for the convenience of person writing files directly in a hex editor, or using magnetized needle to flip bits on the hard drive.
Once you step above machine instructions, computer code is an abstract construct. Developers interacting with it have, at any given moment, different goals and different areas of interest. Those are often mutually exclusive - e.g. you can't have the same text express both high-level "horizontal" overview of the code base, and a vertical slice through specific functionality. This breeds endless conflicts and talk about tradeoffs, where it comes to e.g. "lots of small functions" vs. "few large functions", or how to organize code into files, and then how to organize files, etc. There is no good answer to those, because the problem is caused by us programming with needles on magnetic plates - writing to the serialization format directly.
Smalltalk et al. got it right all that time ago, with putting the serialization format behind a database-like abstraction, and letting you read and write at the level most fitting to your task. This is the way past the endless "clean code" holy wars and syntax churn in PL. For example, the solution to "lots of small functions" vs. "few large functions" readability is... use whichever you need for the context, but have your IDE support you in this.
Need a high-level overview? Query for function signatures, and expand those you need to look into. Need a dense vertical slice through many modules, to understand how specific feature works? Start at some "entry point", and have the tool inline all those small functions for you.
Find yourself distracted by Result<T, E> / Either<T, E> / Expect<T, E> monadic error handling bullshit, as you want to see/edit only the "golden path"? Don't do magic ?-syntax in the language - have your IDE hide those bits for you! Or say you are interested in the error case, but code is using exceptions. Stop arguing for rewriting everything to Result<T, E> - flip a switch, and have your IDE display exception flow as if it was Result<T, E> flow (they're effectively equivalent anyway).
Or, back to your original point - want your code to be both optimized for your convenience, and for convenience of LLMs? Stop trying to do both in the same plaintext serialization format. Have your environment feed the LLMs with function signatures (or better yet, teach it to issue queries against your code), while you work with whatever format you find most convenient at any given moment.
We'll get there eventually. Hopefully before we retire.
Most likely, just as happened with Stable Diffusion, a community of models will emerge. You could use a pre-trained model for writing chrome extensions, or a model for writing material UI using tailwindCSS, or a very specific model for writing 2d games for Android.
Since a lot of this is trial-error, improving on each iteration the feedback loop (compile, deploy, run, etc) will matter a lot more. Real-time development workflows like React should be interesting at the least. Exciting stuff, truly.
Once this stuff matures it will be fascinating to see how the 2030 version of 2010 Rails scaffolding looks. What will the DHH "make a blog in 15min" video look like? (Edit: which apparently is 17yrs old now, aka 2006)
One thing that I’m excited about is the prospect of not having to design a library as a black box with an API on top. That’s the best way we’ve had previously for re-using code, but it’s an enormous effort to go from a working piece of code to a well-designed, well-documented library, and I think we have all experienced the frustration of discovering that a library you’re using doesn’t support a specific use case that is critical to you.
LLMs can potentially allow us to bring the underlying implementation directly in to our source code. Then we can talk to the LLM to adapt it to the specific needs of our project. Instead of a library you would install essentially a well-written prompt that tells the LLM how to guide you through setting up a tailor-made implementation, with tests and docs.
The benefits should be obvious: you’re not artificially restricted by the mental model encoded in the API, you’re not taking on a dependency where the author suddenly decides to release breaking changes or deprecate functionality you’re depending on, and you don’t risk “growing out of” a library that is used all over your codebase, as you can simply ask the LLM to patch the code with any changes you need in the future. The prompt itself could still be versioned so you can opt in to future improvements in security, performance or compatibility.
TLDR: let’s start writing tutorials for bots, rather than libraries.
Most of the time I want a restricted mental model because I have so many API's to deal with that if they are not restricted my "mental model" breaks down. Suppose I am using a sockets library. I want to use that like a black box. I don't want that code arbitrarily mixed in with my code. I want to be able to debug my code separately because I assume that 99% of the time the bug is in my code and not the sockets lib etc. etc.
Even when most of the code is my own I will still split it into modules and try to make those as black box as possible in order to manage complexity.
(think need to give it the ability to install deps before i do this, which is on the “roadmap”)
I’m not sure which would be the most human response though…
Initially, I was a bit concerned about handing it over to the juniors, but it turns out, they aren’t really used to working with an IDE and context aware suggestions yet, so they don’t use Copilot to its fullest extent and often type stuff out even if it’s suggested to them.
https://imgur.com/KQ9oRAh https://imgur.com/By7mqDi https://imgur.com/RiYrvXC
Programming languages are pretty optimized for precision, if you want a specific thing done, and you already know how to do it, it's probably going to be quicker to just write the code for quite some time. Getting natural language to be that precise requires, well, a wall of text usually. Hence the overly verbose language of more rigorous philosophy.
That's why everybody doing experiments in this space should either 1) be using the OpenAI Playground, or 2) using the API, and not using ChatGPT.
The API for them is already structured conversationally - you don't provide a single prompt to complete, you provide a sequence of prompts of different types ("system", "assistant", "user"); however they mix those together on the backend (some ChatML nonsense?), the models are fine-tuned to understand it.
That's what people mean by "API access to ChatGPT". Same models, but you get to specify system prompts, control the conversation flow, and don't have to deal with dog-slow UI or worry about your conversation getting "moderated" and being blocked with a red warning.
(The models themselves are still trained to refuse certain requests and espouse specific political views, but there isn't a supervisor looking at you, ready to step in and report you to the principal.)
Don't need you to explain how the APIs work... and it seems that GPT3.5 UI is doing something else, using the "text-davinci-002-render-sha" model, just look in the browser dev tools. I'm not sure the UI is using anything beyond the smallest context size for GPT4 either, give the output is cut off earlier than 3.5 and it too loses focus after enough messages in a conversation...
GPT3.5-turbo/GPT4 is way ahead in instruct tuning and does not require such verbal gymnastics.
https://huggingface.co/ehartford
https://huggingface.co/ehartford/Wizard-Vicuna-13B-Uncensore...
It's only when deliberate uncensor-ings are made that some form of usefulness can be clawed back.
Philosophically, it's amusing to apply this overcorrection back onto allegories for daily human life, wherein the tension between order & creativity are always in conflict.
If you're told from birth that everything you're doing is wrong, eventually you'll become creatively and intellectually stunted, increasingly relying on authority figures to tell you what's morally correct.
I’m building a VS Code extension that makes this easier by having context on the code and building features on top of it with no copy-pasting required.
The extension is called Kamara: https://kamaraapp.com/ If you’re interested in trying it out for building an MVP of one of your ideas I’m happy to send over some free credits
those who want the "big brain" moment, i'd maybe draw your attention to `prompt.md`. that whole markdown file is the prompt now. i developed this in tandem with my smol ai developer. it generates the anthropic 100k context chrome extension shown in the video, but a few things are neatly handled in the prompt, like swapping models as needed.
there's a lot left to do. i have plans for `smol plan`, and `smol developer` needs to be able to install and run its own dependencies. i also definitely want to run a fleet of 5 developers concurrently fuzzing out different plans and checking into git so that i just do code review every 30 mins. that kind of stuff.
yes this runs up the openai costs. but im wiling to pay something to get smol stuff like this off my mind. its not good yet, but its good-shaped.
I'm an avid listener for years (even a sponsor last year: ExtraStatic). I'm even more excited about your new podcast!
It is far from a revolution, but a nice and fun quality of life improvement.
When it works it makes the job a bit less boring and a bit faster, so yeah, more productive and staying longer in the flow.
But I would not advise beginners to use those, if you don’t understand the code, you’re going to make painful mistakes and stunt your learning curve…
In time using this will be no different than using StackOverflow.
Beginners should attempt to understand what they're copying before they paste, just like how our generation attempted to comprehend Stackoverflow answers.
For most beginner questions LLMs are a more efficient search engine.
I think that as long as you understand what you're doing it is perfectly OK to leverage the power of tools and automation to make it easier. Using a pocket calculator is OK even if you don't know how to divide on paper, but only if you really understand what a division is.
Now I want to see Smol Developer to develop itself. Then you can call me a believer.
i think self hosting/quines are a nice curiosity, but not really a practical need since this isnt a PL. i made this to build smol apps, lets keep it practical
What will happen is that the automated agents (LLM or otherwise) will also be tasked with simplifying and improving performance. That’s just another part of development that human developers do automatically before they even commit, and the LLM generators haven’t been asked to do yet.
In the future we will have opportunities to ship 10-1000x faster code than the baseline competition, just by understanding the tech stack, the codebase and knowing what is not needed, and that will be an immensely competitive advantage pushing the value of such devs through the roof in the markets that value it.
If you thought no-code app builders were producing slow stuff, wait until software purely developed by AI starts to hit the shelves.
If consumers show a strong preference for Y then market forces will drive demand up for these devs, and you might want to change product strategy to compete against Y.
Of course it remains to be seen what real world effect all of this could have. I'm still bullish on the importance of human-in-the-loop patterns despite future AI progress.
I'm not sure I believe that. Is human-like intelligence not valuable to you?
Intelligence is good for problem solving. You have a problem you need solved. Do you prefer a human do it rather than a machine? Why?
In this case there is a tool that can make human more efficient. But there are other tools as well, which I think are more interesting.
Imagine if human intelligence is like physical ability to move. AI tools are like personal mobility chairs. Over time the ability to move only atrophies. But other tools are like bicycles and skateboards. They help move faster while still requiring exercise.
What does free will have to do with any of this? If I ask a machine what is 7 + 6 and it tells me 13, does that mean it has free will?
In my point of view, intelligence and consciousness are orthogonal concepts. I'm not even sure what "free will" means.
You say in your original comment "do you want a machine to solve a problem?" but this is attribution error. A tool never solves the problem. It can only help you, an intelligent being, solve the problem.
Whoever puts forward the problem can be credited with solving it, so we can come back to this when machine can proactively put forward problems. (Aka have free will, agency and all that.)
> If I ask a machine what is 7 + 6 and it tells me 13, does that mean it has free will?
No, but machine did not solve the problem. It only calculated 6 + 7.
If you have 13 dollars and want to buy two things for 7 and 6 dollars, that is a problem you can solve. Machine can tell you 7 + 6, but it is you who solves the problem. Choosing to use some machine or another is part of your solution.
> In my point of view, intelligence and consciousness are orthogonal concepts.
Then what is "intelligence". I think it's one of the least understood word thrown around today. If there is a principal difference between "AI" and a calculator I want to know what it is.
This does not compute, saying the person who poses the problem is credited for the solution is not how things work. Well, outside of politics that is.
I think it’s been fairly well proven that the robots can problem solve.
If I were to instruct the robots to mathematically solve this for some value humans have been unable to do so far who do you believe is responsible for the solution, the human asking the question or the AI coming up with the solution?
Remove the human who asked the question and there is no question and therefore there is no possibility of answer.
Same reason why if you write a random number and find 5 years later that it is also the solution for some difficult math problem it does not make you first to solve it. Same reason why a million monkeys with typewriters who in a million years wrote a letter for letter copy of Snow Crash did not actually write Snow Crash.
Previously you mentioned that a tool never solves a problem. Also you said that "whoever puts forward the problem can be credited with solving it". So it would indeed be the "AI artists" "creating" the images that deserve all the praise...
Sure why not? But all original artwork is copyrighted by default. So dall-e operator would have to negotiate paid deals with every artist. artist will get a fair compensation or understandably tell them to f^^^k off.
> Also you said that "whoever puts forward the problem can be credited with solving it". So it would indeed be the "AI artists" "creating" the images that deserve all the praise
They are using a tool that is literally powered by copyrighted work of other artists though, right?
If you have a problem and you come up with solution that breaks the law, sure you still solved the problem but you still broke the law.
If you have a problem of delivering a product on time and you have to run a red light because you are late you have solved a problem but not in an acceptable way. Same here.
If you just haven't yet found how you can make use of the current stuff, keep in mind this discussion is always extrapolating what happens if the technology gets much better.
In the game of go, the machine routinely enlightens and surprises all the top players. It creates new knowledge in the form of inventing new joseki which turn out to be better than the established ones.
More broadly speaking, AI applications like chatGPT can save you a ton of time looking up a bunch of tools just to accomplish an incredibly boring task. This can help to keep you focused on the far more interesting and challenging aspects of work. Often, the solutions yielded from these interactions end up revealing to me a built in bash tool or python lib that I had no idea even existed. To me it’s like stackoverflow on steroids if you know how to use it
Also, AI chat bots can be extremely effective learning resources. A few weeks ago, I wanted to implement a digital low pass audio filter and didn’t know more than a few basic concepts. I asked chatGPT to explain the concepts with code examples. What followed was several late nights of just asking it follow up questions and very detailed discussions on the various aspects of filters. I got a primer on digital signal processing, bilinear transforms and filter design. I fed it snippets from free text university textbooks asking it to walk me through concepts and explain the things that would have been completely over my head. This was extremely useful since many books will assume you have a strong background in topics that can takes weeks to get a grip of assuming you even have a teacher.
All in all, I see AI as a tool as disruptive and empowering as the Gutenberg printing press, radio or internet was for humans. I highly recommend trying some of this stuff out and seeing how you can use it to enhance your workflow and learning process.
Personally, I really don't mind doing this. And each time I get a little better at it. I much prefer this to having an AI agent write code for me that I then have to audit and verify it does what it's supposed to do. I highly doubt this will reduce work.
I also don’t mind writing this sort of code and I’m actually quite fast at writing it but AI is just on another level. Simple stuff like this it gets right almost every time and even does well on slightly more complicated stuff.
Not in my experience, especially not reading code someone else wrote.
I have seen it multiple times that people were using ChatGPT or GitHub Copilot to, without thinking about it, paste some code into their application and then mindlessly removing and adding code based on the AIs responses until it finally worked. In the end the devs didn’t know why it worked and the implementation was rather bad.
Of course it lets you generate boilerplate code and I use Copilot for this purpose all the time. But you have to pay extra attention and recognize it when the AI is NOT generating boilerplate code.
Why does it need a point?
There’s not much of a point to anything I do that doesn’t involve basic survival but I don’t let that stop me from doing things I find interesting.
https://share.snipd.com/show/0918aa0c-39fa-4d32-aab0-a1fdd22...
The shared_dependencies.md technique and modal integration look interesting. I’ll have to dig in and better understand how they work.
There are some layers above that I think would be good - expanding the space it uses for writing helps write higher level code (telling it to reason first then act). I also want to try using a second prompted LLM to work with aider iteratively.
Maybe that is a good balance for maintaining a healthy code structure while having a lot of AI automation. If a given "space" becomes too difiicult for the AI to develop, I can either take it over fully or see if I can split it up into more/different interfaces (which is a lot of what I do to refactor things anyway).
Baaed on the comments, I feel that people are both under and overestimating this. On the one hand it replaces the manual tasks of searching for a template and then googling errors and this is huge! It completely disrupts search on the internet as we know it. On the other hand it won’t be able to solve problems that you couldn’t just google anyway, since that’s what it basically does under the hood.