102 karma · joined June 11, 2015
I will add more model support soon - any models you particularly want to see?
I’m also hoping to test out RL on the tools to get a fine-tuned model specifically for browser automation eventually.
Also there's a matter of taste, as commented above, the best way to use these is going to be running multiple runs at once (that's going to be super expensive right now so we'll need inference improvements on today's SOTA models to make this something we can reasonably do on every task). Then somebody needs to pick which run made the best code, and even then you're going to want code review probably from a human if it's written by machine.
Trusting the machine and just vibe coding stuff is fine for small projects or maybe even smaller features, but for a codebase that's going to be around for a while I expect we're going to want a lot of human involvement in the architecture. AI can help us explore different paths faster, but humans need to be driving it still for quite some time - whether that's by encoding their taste into other models or by manually reviewing stuff, either way it's going to take maintenance work.
In the near-term, I expect engineering teams to start looking for how to leverage background agents more. New engineering flows need to be built around these and I am bearish on the current status quo of just outsource everything to the beefiest models and hope they can one-shot it. Reviewing a bunch of AI code is also terrible and we have to find a better way of doing that.
I expect since we're going to be stuck on figuring out background agents for a while that teams will start to get in the weeds and view these agents as critical infra that needs to be designed and maintained in-house. For most companies, foundation labs will just be an API call, not hosting the agents themselves. There's a lot that can be done with agents that hasn't been explored much at all yet, we're still super early here and that's going to be where a lot of new engineering infra work comes from in the next 3-5 years.
o3 in codex is pretty close sometimes. I prefer to use it for planning/review but it far exceeds my expectations (and sometimes my own abilities) quite regularly.
The rest is glorified boilerplate that I find usually saps me of my energy, not gives me energy. I'm a fan of anything that can help me skip over that and get to the more enjoyable work.
Just looking at all of the amazing tools and workflows that people have made with ComfyUI and stuff makes me wonder what we could do with diffusion LMs. It seems diffusion models are much more easily hackable than LLMs.
- 2 frontend React/Tailwind codegen projects (1 an agent, and 1 a really cool website builder)
- 1 Node.js codegen AI agent
- 1 ecommerce product placement image generator
I'm not working on these anymore (and the code and dependencies are probably far out of date now unfortunately - haven't looked at some of these for a year at this point).
But I thought I would just go ahead and make the code public and share it out there in case it can help anyone or inspire some new ideas.
Besides just sharing the code, if you click through to the detailed pages for each project you can see a demo video of each one showing how it works and read some of my notes about each one.
About the projects:
I really love the website builder, it's probably my favorite project I've done - so many cool details.
Other projects have some cool agent things going on, maybe some novel approaches to codegen in there - not sure.
The image generation one has a bunch of ComfyUI workflows that I spent a bunch of time on.
Happy to answer any questions here - although I'll probably have to review the code if anything is too in the weeds as I've forgotten a lot already.
Also the code is not very well-organized and there's some dead code and stuff floating around and things commented out or half-implemented at places. I didn't bother to clean it up too much besides making sure I'm not leaking any env secrets or API keys (if you spot one, let me know please!)
And last - these are mostly written in Node.js, React, and a bit of Python. I don't claim to be an expert in these languages and was pretty unfamiliar with using them before I built these - so don't use this as a reference on how to write great code, rather I hope you enjoy the concepts behind the features.
Hope it can be of interest and potentially of help to someone out there!
I'm not working on these anymore (and the code and dependencies are probably far out of date now unfortunately - haven't looked at some of these for a year at this point).
But I thought I would just go ahead and make the code public and share it out there in case it can help anyone or inspire some new ideas.
Besides just sharing the code, if you click through to the detailed pages for each project you can see a demo video of each one showing how it works and read some of my notes about each one.
About the projects:
I really love the website builder, it's probably my favorite project I've done - so many cool details.
Other projects have some cool agent things going on, maybe some novel approaches to codegen in there - not sure.
The image generation one has a bunch of ComfyUI workflows that I spent a bunch of time on.
Happy to answer any questions here - although I'll probably have to review the code if anything is too in the weeds as I've forgotten a lot already.
Also the code is not very well-organized and there's some dead code and stuff floating around and things commented out or half-implemented at places. I didn't bother to clean it up too much besides making sure I'm not leaking any env secrets or API keys (if you spot one, let me know please!)
And last - these are mostly written in Node.js, React, and a bit of Python. I don't claim to be an expert in these languages and was pretty unfamiliar with using them before I built these - so don't use this as a reference on how to write great code, rather I hope you enjoy the concepts behind the features.
Hope it can be of interest and potentially of help to someone out there!
I made a bunch of GenAI projects in 2024: - 2 frontend React/Tailwind codegen projects (1 an agent, and 1 a really cool website builder) - 1 Node.js codegen AI agent - 1 ecommerce product placement image generator
I'm not working on these anymore (and the code and dependencies are probably far out of date now unfortunately - haven't looked at some of these for a year at this point).
But I thought I would just go ahead and make the code public and share it out there in case it can help anyone or inspire some new ideas.
Besides just sharing the code, if you click through to the detailed pages for each project you can see a demo video of each one showing how it works and read some of my notes about each one.
About the projects:
I really love the website builder, it's probably my favorite project I've done - so many cool details.
Other projects have some cool agent things going on, maybe some novel approaches to codegen in there - not sure.
The image generation one has a bunch of ComfyUI workflows that I spent a bunch of time on.
Happy to answer any questions here - although I'll probably have to review the code if anything is too in the weeds as I've forgotten a lot already.
Also the code is not very well-organized and there's some dead code and stuff floating around and things commented out or half-implemented at places. I didn't bother to clean it up too much besides making sure I'm not leaking any env secrets or API keys (if you spot one, let me know please!)
And last - these are mostly written in Node.js, React, and a bit of Python. I don't claim to be an expert in these languages and was pretty unfamiliar with using them before I built these - so don't use this as a reference on how to write great code, rather I hope you enjoy the concepts behind the features.
Hope it can be of interest and potentially of help to someone out there!
Deep Research works reliably but I wish the agent would do some due diligence on its link selection and also today I uncovered where it misquoted a website which was very misleading.
Doing batch merges with a merge queue can speed up things if you have a ton of longer running end to end and integration tests. But then if a test fails you need to identify which commit out of the batch is causing it so you don’t reject the entire batch.
I wrote about that here: https://qckfx.com/blog/ai-powered-stagehand-git-bisect-findi...
For the most part, code reviews should be optional - if you want to get a review from someone, tag them on your PR and ask. If someone you didn't tag spots something and your PR landed, you can always figure it out and still make a fix.
I will give an exception to maybe super fragile parts of the code but ideally you can refactor/build tests/do something else that doesn't require blocking code review to land changes.
I saw this first hand when I worked on operational excellence at Meta. Large companies can afford to hire teams to manually reproduce bugs before passing off to developers, but that option is not available to most companies.
I can also see benefit in using this to develop an AI that is potentially more data/power efficient than something like GPT. In the same line of thinking, there is the potential for novel algorithm discoveries since evolution is based on chance mutations vs being constrained to human creativity.
Has there been any other similar research into this from an AI perspective? If not, why?
It sounds like the focus is on migrations, refactors, hiding stuff behind feature flags, etc. That's useful but it's less than what it sounded like from codegen's marketing materials which makes sense, they are selling the vision and not the current product.
If it's limited in handling complexity right now then it sounds like a nice feature for Linear. That also sounds like a good way to enter developer's workflows if you have some button on linear tasks to just automate it. That's not L5 engineer though and I bet it's expensive on the order of thousands of dollars per month for the OpenAI use.
Congrats to the team on the round, excited to see how it develops.