Show HN: We are building an open-source IDE powered by AI
github.com
github.com
When I tried to “talk” to gpts, it was always hard to formulate my requests. Basically, I want to code. To tell something structured, rigorous, not vague and blah-blah-ey. To check and understand every step*. Idk why I feel like this, maybe for me the next ten years will be rougher than planned. Maybe because I’m not a manager type at all? Trust issues? I wonder if someone else feels the same.
Otoh I understand the appeal partially. It roots in the fact that our editors and “IDEs” suck at helping you to navigate the knowledge quickly. Some do it to a great extent, but it’s not universal. Maybe a new editor (plugin) will appear that will look at what you’re doing in code and show the contextual knowledge, relevant locations, article summaries, etc in another pane. Like a buddy who sits nearby with a beer, sees what you do and is seasoned enough to drop a comment or a story about what’s on your screen any time.
Programming in a natural language and trusting the results I just can’t stand. Hell, in practice we often can’t negotiate or align ourselves in basic things even.
* and I’m not from a slow category of developers
This is just gpt-3. With chatGPT i finally managed to get unit test backbones for the most annoying methods in our code base that stuck our coverage at 92%. Once we get full copilot gpt-4 integration we will likely get to 98% coverage in a days time. That’s not nothing.
With GPT-4, I thought: "fuck. It's all over. How soon will I get laid off?"
Only half-joking, but GPT-4 is truly good at complex programming and problem solving, and noticeably better than GPT-3.5.
Something about the iterative context aware approach of chat helps me think through the problem. The only think I don't like about copilot is how you cant tell it to augment the output. Sometimes it's close but not perfect.
But you can actually think of this as a code quality measure. You want to avoid this. You want to organize the project in a way that writing or modifying the code is easy. That means, make it straightforward how functions are called, how things are done. Minimize interdependencies, side effects.
Then Copilot will work quite well.
A specification can be rigorous, structured, precise and formal: a declarative description of what is to be done - not how. It is a very different perspective from designing algorithms - even one that proves the thing to be done, is done.
But I think trusting the results is misplaced. It's more like being a partner in a law firm, where the articled clerks do all the work - but you are responsible. So you'd better check the work. And doing that efficiently is an additional skill set.
Some things that helped me to become a lot more efficient:
- Vim. Since you mostly edit code, not write it, it is reasonable to be in the editing mode by default, with a full keyboard of shortcuts you can use, not only ctrl/alt combinations. I use IdeaVim[1] plugin for Vim emulation.
- Never use your mouse. Keyboard actions are faster and more consistent once you get used to it. IDEs and editors are highly customizable and should allow to bind almost any action.
- Get rid of tabs, keep 1-3 active buffers in your vicinity. And then use ctrl+tab or ctrl+e to switch between them.
- Learn tricks for code navigation: jump to definition/usage/implementation, search file/class/symbol.
- Learn tricks for text navigation: pretty much any decent tutorial on Vim navigation will do.
However, this is the state of the art today. In the future, the training set will be based on the prompt-result output and refinement process, leading me to believe that next generation tools will be much better at prompting the user to provide the details. I've already seen this enhancement in gpt4 recently, I think this is a common and interesting use case.
Overall, these tools will become more and more advanced and useful for developers. Now is a great time to become proficient and start learning how the future will require you to adapt your workflow in order to best leverage chatGPT for development.
This response was written by a human.
It's still mental work. You still have to read the code over and edit stuff.
It helps me save lots of time so I can do quality things like comment on hn during work..
GPT seems really smart and you can give it very high-level prompts, which generate good quality code. But it's limited by the chat interface, and if it doesn't get it right the first time, it's tricky to get a tweaked version. For example, I asked it to generate some code for a login endpoint, and gave it all the details about language, libraries, database etc. It produced some pretty good code. But then I asked it to do a version didn't store passwords in plaintext. It did that, but also made other changes including using a different library than one I had specified. So then I had to try to fix that, which led to more weirdness. I ended up just using the code it generated as inspiration and writing it myself.
Copilot, with its integration into VSCode is quite different. It's *awesome*. It makes suggestions in inline, as you type. It seems to have a lot of context, and generates very good code, apparently based on information contained in other files in the same project. It matches style and naming conventions, and even when it's not quite right, I find it easy to just accept the suggestion, and tweak it. It's a huge help.
So my current practice is to rely on Copilot to generate code and use GPT as sort of a pairing partner that I can talk to about the code. It's pretty good at suggesting alternatives, developing ideas, explaining error messages and so on. This is a very, very early attempt at coding with the help of AI, and I haven't actually seen a productivity improvement, because I don't know how to use these tools well. But it's fun and exciting!
I.e. for a backend dev that hardly knows react and needs to write some, who usually gets by doing small modifications and copy-paste-driven development there, ChatGPT / Copilot is basically a 10x productivity multiplier.
I agree that I've not been feeling that much value yet when working in areas I'm proficient in.
I, too, was skeptical until...
I've making the switch to VS Code (not by choice...my previous daily driver, Atom, has been EOLed since December). Accordingly, I needed to port an Atom extension I wrote a long time ago to VS Code. What I didn't want to do was spend hours or days getting up to speed on how to write VS Code extensions, especially since this is the only one I'm ever likely to write (it's the only one I ever wrote for Atom :-)).
I used GPT-4 and it worked like a charm -- almost like magic. It wrote almost all the boilerplate and bullshittery for me. Was it perfect? No, the regex it came up with for recognizing the code pattern to which the extension applies (Scheme atoms or pairs) was crap. That's okay -- I know how to write regexes (in fact, I pretty much just dropped in my state machine parser from the old extension). But not having to spend hours and hours groveling through Microsoft docs and Googling to learn stuff I'm not likely to ever need again? Man, that was great.
> Programming in a natural language and trusting the results I just can’t stand.
This would worry me, too. But in this case, I was already well-familiar with Javascript and Typescript, so it was no problem for me to read the resulting code and give it a sanity check, test it, etc. I just didn't have to write it.
If you haven't already, look into trying out Copilot (I don't think there's a free version anymore?) or Codeium (free and has extensions for a lot of editors, just be careful with what code you're okay sending to a third party). Using AI prompts as a fancy auto complete is what has been giving me the most benefit. Something simple like a comment "# Test the empty_bank_account method" in an existing file using pytest, I go to the next line and then hit tab to auto-complete a 10-15 line unit test. It's not always right, but it definitely helps speed things up compared to either typing it all out or copy/pasting.
My biggest annoyance so far, at least with Codeium, is that sometimes it guesses wrong. With LSP-based auto-complete, it only gives me options for real existing functions/variables/etc. But Codeium/copilot like to "guess" what exists and can get it wrong.
> Like a buddy who sits nearby with a beer, sees what you do and is seasoned enough to drop a comment or a story about what’s on your screen any time.
I agree with you, this is probably where it would be most useful (to me). I don't need someone to write the code for me, but I'd love an automated pair-programming buddy that doesn't require human interaction.
My take is that data-centric programming requires too much context for GPT, and we're going to see a move back to doing things in a more OO way, hopefully with better languages than exist now. The ability to reason locally about objects helps both human and AI developers, so we can build larger systems. Data-centric/functional programs are akin to hand-crafted, artisanal goods that were slowly overtaken by standardized parts/division of labor production methods in the industrial revolution, and we are in the middle of one for software [0]. By the ended of it, software engineering may no longer be an oxymoron.
0. https://www.researchgate.net/profile/Brad-Cox-2/publication/...
The reason I mentioned data pipelines is I've adopted a compiler type view of programming, where data is mutated through series and or parallel steps until a satisfactory side-effect or output is reached. I find its a lot easier to define a problem when I think about it this way.
You will be if you don't embrace AI ;)
Our editor isn't a regular coding editor. You don't actually write code with e2b. You write technical specs and then collaborate with an AI agent. Imagine it more like having a virtual developer at your disposal 24/7. It's built on a completely new paradigms enabled by LLMs. This enables to unlock a lot of new use cases and features but at the same time there's a lot that needs to be built.
You could tell me, in the most painstaking detail, what you want me to paint, and I still couldn't paint it. You can take any random person on the street and tell them exactly what to type and they'd be able to "program".
And over time, people will discover some basic phrases and keywords they can use to get certain results from the AI, and find out what sentence structures work best for getting a desired outcome. These will become standard practices and get passed along to other prompt engineers. And eventually they will give the standard a name, so they don’t have to explain it every time. Maybe even version schemes, so humans can collaborate amongst themselves more effectively. And then some entrepreneurs will sell you courses and boot camps on how to speak to AI effectively. And jobs will start asking for applicants to have knowledge of this skill set, with years of experience.
Although you may know about LLMs, you might specialize in speaking to specific models and know how to get optimal results based on their nuances.
This just sounds like a language that is hard to learn, undocumented, and hard to debug
Because funky brackets are confusing.
If I want to change 1 character of a generated source file, can I just go do that or will I have to figure out how to prompt the change in natural language?
It's in the frustration valley between WYSIWYG and just writing code. The worst of both worlds.
Great question. I would love to hear the devs thoughts here. This is one of those questions where my intuition tells me there may be a really great "first principles" type of answer, but I don't know it myself.
For example, the last ai company installer I just clicked "decline" to (a few minutes ago) says that you give it permission to download malware, including viruses and trojans, onto your computer and that you agree to pay the company if you infect other people and tarnish their reputation because of it. Literally. It was a very popular service too. I didn't even get to the IP section
edit: those terms aren't on their website, so I can't link to them. They are hidden in that tiny, impossible to read box during setup for the desktop installer
per the end of the readme:
- one step to go from a spec to writing code, run code, debugging itself, install packages, and deploying the code
- creates "ephemeral UI" on demand
Jira?
Only slightly joking. It really sounds like we're moving in the direction of engineers being a more precise technical version of a PM, but then engineers could just learn to speak business and we don't need PMs.
It's pretty nascent and a lot of work needs to be done so bear with us please. I'm traveling in 30 minutes but I'm happy to answer your questions.
A little about my co-founder and myself: We've been building devtools for years now. We're really passionate about the field. The most recent thing we built is https://usedevbook.com. The goal was to build a simple framework/UI for companies to demo their APIs. Something like https://gradio.app but for API companies. It's been used for example by Prisma - we helped them build their playground with it - https://playground.prisma.io/
Our new project - e2b - is using some of the technology we built for Devbook in the past. Specifically the secure sandbox environments where the AI agents run are our custom Firecracker VMs that we run on Nomad.
If you want to follow the progress you can do follow the repo on GH or you can follow my co-founder, me, and the project on Twitter:
- https://twitter.com/t_valenta
And we have a community Discord server - https://discord.gg/U7KEcGErtQ
[1]: In the spirit of https://wiki.mozilla.org/Areweyet
You're welcome to use it if you want to get a link page started, and I'd be glad to help - you can also add comment sections on the page to get user input/contributions so if anyone else has some links they can comment them there. I eventually want to more fully formalize user contributions to pages so that they can be used as crowdsourced freeform sites, if theres enough interest out there.
Do I understand correctly that the GPT4All provides a delta on top of some LLAMA model variant? If so, does one need to first obtain the original LLAMA model weights to be able to run all the subsequent derivations? Is there a _legal_ way to obtain it without being a published AI researcher? If not, I'm not sure that Gpt4All is viable when looking for legal solutions.
I think the crucial part is indeed not being able to deterministically go from NL to code but to take an existing state of the codebase and spec and "continue the work".
So my question would be, what is the use case?
I guess it’s more like planing software and not implementing it.
You can pretty well plan your software with ChatGPT. But it will just help you not really doing the job.
> mlejva: Our editor isn't a regular coding editor. You don't actually write code with e2b.
then what licensing problems arise from its use? In theory, if you only prompt the AI to write the software, is the software even your intellectual property?
It seems like this is a public domain software printing machine if you really aren’t meant to edit the output.
I think everyone had this idea and is building something similar. I know I am.
I’m not really planning on turning it into a product. It sounds like this guy is a lot farther along than me if you’re looking for a competitor - I think you’re going to have plenty. https://mobile.twitter.com/codewandai
My gut feeling is we're still a few LLMs generations away from this being really usable but I'd love to hear how the authors are thinking about this.
All of that (issues list, version, tone, etc) is then formulated into a GPT prompt. The prompt is structured such that it returns written release notes. That note is then stored and the user can edit it using a rich text editor.
Once the first note is created the system can help the user write future notes by predicting release version, etc.
This isn’t that complex imo, but I’m curious to see if this is what people consider complex.
ChatGPT wrote 90% of the code for this.
> Line of business (AKA LOB) is a term that describes a business’s product or service, the resources used, and the process for delivering value to a market segment. It could be the primary or one of the main processes that bring revenue.
> For example, manufacturing dry-erase markers is a line of business. Everything that happens from concept, developing the markers, marketing, selling, to fulfillment, and staying competitive makes up the business line. So, a LOB could also describe a product line.
How long did it take?
I suspect it has to do with the equivalent of prompt engineering: it's too difficult to cross the cultural and linguistic barriers, as well as the barrier of space that could have mitigated the other two. By the time you've directed somebody to do the work with sufficient precision, you could have just done it yourself.
And it's part of the reason we keep looking for that 10x superdeveloper. Not just that they produce more code per dollar spent, but that there is less communication overhead. The number of meetings scale with the square of the number of people involved, so getting 5 people for the same price as me doesn't actually save you time or money.
I have no idea what that means for AI coding. Thus far it looks a lot like that overseas developer who really knows their stuff and speaks perfect English, but doesn't actually know the domain and can't learn it quickly. (Not because they aren't smart, but because of the various human factors I mentioned.)
I'd be thrilled to be completely wrong about that -- in part because I've been mentally prepared for it for so long. I hope that younger developers get a chance to spin that into a whole new career, hopefully of a kind I can't even imagine.
And it’s much slower because “do it” includes trial-error-decision cycle, which is fast when you’re alone and weeks if you are directing and [mis]communicating. Also wondering where it goes and how big of a bubble it is/will be.
A lot of UX, UI, DX work related to LLMs is completely new. We ourselves have a lot of new realizations.
> While you mention that you can bring your own model, prompt, etc, the current main use case seems to be integrating with OpenAI
You're right. This is because we started with OpenAI and to be fair it's easiest to use from the DX point of view. We want to make it more open very soon. We probably need more user feedback to learn what would be the best way how this integration should look like.
> How, if at all, do you plan to address the current shortcoming that the code generated by it often doesn't work at all without numerous revisions?
The AI agents work in a self-repairing loop. For example they write code, get back type errors or a stack trace and then try to fix the code. This works pretty well and often a bigger problem is short context window.
We don't think this will replace developers, rather we need to figure out the right feedback loop between the developer and these AI agents. So we expect that developers most likely will edit some code.
It doesn’t so for many/most humans either. I hope they have a revisions prompting, but I did not try it yet.
I noticed adding in a feedback/review loop often fixes it, but the you still need someone saying ‘this is it’ as it doesn’t know the cut off point.
Hahaha. I’ve been coding for over 20 years and this is definitely not the case.
> Current AI models often generate code that doesn't even compile.
Most of the code ChatGPT has given me, has run/compiled on the first try. And it’s been a lot longer and complex than what I would have written on a first pass.
Let’s just learn to use these tools instead of trying to justify human superiority.
Aha. Maybe you know super clever people or people who learned in the 60-80s when cycles (and reboots etc) mattered or were costly; this is incredibly far from the norm now.
The feedback/review loop is spot on - a lot of the problems can be fixed automatically in a few steps but you actually need the outputs/errors.
Sure, if you're doing greenfield, just ask it to write a new file. If you only have a simple script, you can ask it to rewrite the whole file. The tricky bit however, is figuring out how to edit relevant parts of the codebase, keeping in mind that the context window of LLMs is very limited.
You basically want it to navigate to the right part of the codebase, then scope it to just part of a file that is relevant, and then let it rewrite just that small bit. I'm not saying it's impossible (maybe with proper indexing with embeddings, and splitting files up by i.e. functions you can make it work), but I think it's very non-trivial.
Anyway, good luck! I hope you'll share your learnings once you figure this out. I think the idea of putting LLMs into Firecracker VMs to contain them is a very cool approach.
But context windows are far from being large enough to fit entire repos, nor even entire files (if they're big). I'm not sure how hard just scaling up the context window is, from the current early access of OpenAI GPT-4.
> in a way that the LLM can retrieve and focus on the relevant files
This I think is something we haven't really figured out yet, esp. if a feature requires working on multiple files. I wouldn't be surprised if approaches based on the semantic level (actually understanding the code and the relationships between it's parts; not the textual representation of it) won't be needed in the end here.
This also implies that the first codebases to really benefit from LLM collaboration will be those written in strongly typed languages which are already amenable to static analysis.
And in terms of context windows, it's not like humans keep the entire codebase in their head at all times either. As a developer, when I'm focused on a single task, I'm only ever switching between a handful of files. And that's by design; we build our codebases using abstractions that are understandable to humans, given our limited context window, with its well-known limit of about seven simultaneous registers. So if anything, perhaps the risk of introducing an LLM to a codebase is that it could create abstractions that are more complicated than what a human would prefer to read and maintain.
It was just a PoC so it wasn't bulletproof and I only ever made it work on a small Next.js codebase (with every file being intentionally small so it fits the context window) but it did work correctly.
What happens if you use the README.md and associated documentation as a prompt to re-implement this whole thing?
Don’t most devs tend to be extremely sticky to their preferred dev env (last big migration was ST -> VS Code back in 2015-2017)
Read all the comments in here. Still not getting why this isn’t a VS Code plugin.
Distribution almost always beats product.
Also it's a somewhat recent update, but VSCode asks you if you trust the code author now when you open a project https://code.visualstudio.com/docs/editor/workspace-trust
For some time I've been giving serious thought about an automated web service generator. Given a data model and information about the data (relationships, intents, groupings, etc.) output a fully deployable service. From unit tests through container definitions, and everything I can think of in-between (docs, OpenAPI spec, log forwarder, etc.)
So far, while my investment hasn't been very large, I have to ask myself: "Is it worth it?"
Watching this AI code generation stuff closely, I've been telling myself the story that the AI-generated code is not "provable". A deterministic system (like I've been imagining) would be "provable". Bugs or other unintended consequences would be directly traceable to the code generator itself. With AI code generation, there's no real way to know for sure (currently).
Some leading questions (for me) come down to:
1. Are the sources used by the AI's learning phase trustworthy? (e.g. When will models be sophisticated enough to be trained to avoid some potentially problematic solutions?)
2. How would an AI-generated solution be maintained over time? (e.g. When can AI prompt + context be saved and re-used later?)
3. How is my (potentially proprietary) solution protected? (e.g. When can my company host a viable trained model in a proprietary environment?)"
I want to say that my idea is worth it because the answers to these questions are (currently) not great (IMO) for the AI-generated world.
But, the world is not static. At some point, AI code generators will be 10x or 100x more powerful. I'm confident that, at some point, these code generators will easily surpass my 20+ years of experience. And, company-hosted, trained AI models will most likely happen. And context storage and re-use will (by demand) find a solution. And trust will eventually be accomplished by "proof is in the pudding" logic.
Basically, barring laws governing AI, my project doesn't stand a cold chance in hell. I knew this would happen at some point, but I was thinking more like a 5-10 year timeframe. Now, I realize, it could be 5-10 months.
>1. Are the sources used by the AI's learning phase trustworthy? (e.g. When will models be sophisticated enough to be trained to avoid some potentially problematic solutions?)
Probably not, but for most domains reviewing the code should be faster than writing it.
>2. How would an AI-generated solution be maintained over time?
I would imagine you don't save the original prompts. Rather, when you want to make changes you just give the AI the current project and a list of changes to make. Copilot can do this to some extent already. You'd have to do some creative prompting to get around context size limitations, maybe giving it a skeleton of the entire project and then giving actual code only on demand.
> When can my company host a viable trained model in a proprietary environment?
Hopefully soon. Finetuned LLaMA would not be far off GPT-3.5, but nowhere close to GPT-4. And even then there are licencing concerns.
1> Relying on code reviews has concerns, IMO. For example, how many engineers actually review the code in their dependencies? (But, I guess it wouldn't take that much to develop an adversarial "code review" AI?)
2> Yes, agreed, that would work. Provided the original solution had viable tests, the 2nd (or additional) rounds would have something to keep the changes grounded. In fact, perhaps the existing tests are enough? Making the next AI version of the solution truly "agile"?
3> So, at my age (yes, getting older) I'm led to a single, tongue-in-cheek / greedy question: How to invest in these AI-trained data sets?
AWS roughly has one of these in Amplify. The data mapping parts are pretty great, though lots of the rest of it suck. The question I'd ask is if those parts suck by nature of the setup, or is it just amplify that has weird assumptions
Copilot already exists and copilot X already packs the features this package promises AND much more, why use this application over Copilot?.
Disagree with your fears about “give them data”.
Here’s their data policy: https://openai.com/policies/api-data-usage-policies
Literally states it won’t use your data. Ofc there’s non trivial risk that this policy will change over time. Still, I don’t feel like there’s any huge lock-in risk with OpenAI right now, so advantage of using it outweighs the risk for most.
How are you folks looking beyond something that is so simple and should be correct always to be useful? How does it save you time if any small detail may be a potential problem?
You think improving their produce for them while submitting yourself to the authority of their plugin market is a good idea?