Add an AI Code Copilot to your product using GPT-4
windmill.dev
windmill.dev
https://twitter.com/transitive_bs/status/1646778220052897792
https://www.youtube.com/watch?v=rd-J3hmycQs
Features like their "AI Fix" feel much more like the direction things should go. Providing context-based actions without requiring user input.
E.g. use the database schema, existing code, etc as context to suggest actions without a user having to type it out.
Personally copilot is a magical typing speedup and that got me enthusiastic about next step - but it looks like next step is going to require another breakthrough considering GPT4 HW requirements and actual performance.
Chat interface as a search bot is the only use case I've found it useful outside of copilot - regurgitating relevant part of internal documentation as search on steroids, even 3.5 is relatively decent at this.
You also need to provide adequate context. I don't think source code alone is good enough context most of the time (E.g. need an error message as well), unless you're pulling in multiple chunks of code.
1. Write a prompt, get a full script in response (good! as I expected)
2. Realize something missing in the prompt, or want to improve a specific function.
3. Prompt like "For function foo_bar(), give me just a updated version of that function that adds exception handling for missing files".
4. Chat GPT ignores the (admittedly implied in my language above) "just that function" part, and rewrites the entire script.
While usually it only modifies the part of the script I asked it to, it's annoying because (1) it's slower to give me the whole thing back (2) need to do additional checks in case the function I care about depends on things outside of it that might have been changed.
I've had this a few times, so I'm in the habit of not trying to iterate with chatGPT when asking it for scripts, but instead just doing the rest myself.
(Disclaimers: I haven't tried too hard to find a solution to this; there might be one; I also don't generally combine ChatGPT with Copilot; one or the other on different projects; and haven't tried other forms of using GPT-4...)
This is a small recent example: https://chat.openai.com/share/2aa979ca-c796-4dfa-ae13-b21d39...
And another example where I ask ChatGPT to remove some error handling: https://chat.openai.com/share/90ce6336-f35c-40fc-8fce-baefc5...
GitHub Copilot seems a bit more temperamental. If it doesn't give me something decent first time I don't bother trying to follow-up.
I'm also not doing too much, volume-wise; maybe about a script/week. I'll keep this in mind in future work!
"you will be given a section of code wrapped in ``` your task is analysis this code for any possible issues. You should list any issues you find and why you believe it's an issue. Then explain how to correct it"
I started working on something for this purpose called j-dev [0]. It started as a fork off smol-dev [1] which basically gets GPT to write your entire project from scratch. And then you would have to iterate the prompt to nuke everything and re-write everything, filling in increasingly complicated statements like "oh except in this function make sure you return a promise"
j-dev is a CLI where it gives a prompt similar to the one in the parent article [2]. You start with a prompt and the CLI fills in the directory contents (excluding gitignore). Then it requests access to the files it thinks it wants. And then it can edit, delete or add files or ask for followup based on your response. It does stuff like make sure the file exists, show you a git diff of changes, respond if the LLM is not following the system prompt, etc.
It also addresses the problem that a lot of these tools eat up way too many tokens so a single prompt to something like smol-dev would eat up a few dollars on every iterations.
It's still very much a work in progress and i'll prob do a show hn next week but I would love some feedback
[0] https://github.com/breeko/j-dev
[1] https://github.com/smol-ai/developer
[2] https://github.com/breeko/j-dev/blob/master/src/prompts/syst...
For you.
Long before YouTube versus WikiHow we've had "visual learners", "practical learners", and those who prefer learning from clear writing.
So it's probably not just some sort of generational or way-of-thinking divide, it's probably a cohort of personas who far prefer text to rapidly understand what things do, and a cohort who prefer someone to "show and tell"[^1].
That said, the balance does seem to have shifted in recent decades, perhaps for Americans around the same time (correlation not causation) as free play outside and big three over-the-air TV networks gave way to helicopter parenting of fully programmed days and/or the first 150 channel MTV generation.
[^1]: "show and tell" (1954?) - https://en.wikipedia.org/wiki/Show_and_tell
https://www.loom.com/share/969562d3c0fd49518d0f64aecbddccd6?...
The CI integration sounds like the most interesting part to me, since I usually let things fail in CI then go back and fix them.
It's kind of in an interesting spot because it's not instant feedback like a linter/type checker, but only running it in CI feels like a waste of potential.
I hope it becomes a successful product!
Some things we've done with windmill so far:
* Slack bots / alerts from scheduled jobs -> All sourced in one location
* One source of truth for connectors (google drive / sheets, slack, postgres database aggs)
* Complex ATS flows (We load our applicants through a windmill task that runs LLMs to parse against a rubric we made, scores then get pushed back into the ATS system as an AI review, then that job reports to slack which helps us prioritize)
* Dashboards (Though I'll admit a bit of sharp-edges here on the UI builder) running complex python + pandas tasks that pull from postgres and run dashboards for us with pretty VEGAlite plots
* Job orchestration -- (though this is partially on hold) we have prototyped orchestrating some of our large workloads, data pipelines, etc. in a few different ways. (we run spark jobs and ray jobs on distributed clusters, and have been trying to use windmill to orchestrate and schedule these, with alerts)
Additionally, windmill made running and tracking these things (like the ATS system) accessible to my "low"-technical co-founder, who regularly will hop in and run the job or browse through error logs of previous runs.Last, I found some bugs, reported in their discord and I heard back from rubenf very quickly and fixes went in rapidly too! Huge shoutout to the great work the windmill team is doing.
An interesting question is, what are you providing above copying and pasting a load of my data into ChatGPT in a way that I’m not sure I gave you permission to?
This response is powered by HackerGPT4.
For a given task, ChatGPT seems to do way better if it given a precise prompt with detailed instructions. This includes describing the context of the problem, maybe what kind of data it will be operating on, and especially the format in which to output its answer. Once the answer is produced in the expected format, the application can then integrate it into its own data and make it immediately available.
An embedded AI assistant in a product can provide a lot of this context, whereas a simple user of ChatGPT on the web might not end up with an answer that they can use immediately. I have no doubt that many people spend more way time preparing data for and extracting data out of ChatGPT than how long it took to answer.
You can see an example of this in this demo of a tool called Tana: https://youtu.be/FlqpK8ucf8s?t=310 – the user is asking for airline home bases, and has to add "do not mention the country" for the data to have the right format. Once the prompt is correct though, the app is able to ingest this generated data and merge it with the user's.
I do this with code, design documents, and task planning.
I was blown away by how effective it is when I tried it. Unfortunately there are cases where it doesn't work well, and it can still hallucinate solutions when it has no idea what the problem is, but generally speaking it's excellent.
My goal was to tie this into VS Code and attach it to processes that would expose an error stack to the AI agent when things went wrong, and it would then iteratively attempt to find and test solutions for you. I couldn't get the success rate high enough to bother publishing it, but it was really fun and I learned a lot. I also saw that GitHub is essentially doing this and much more with Copilot X, so my solution would never compare anyways. I think this will become a huge time saver in the near future, though. When it worked well, I could have a solution to the error and tests to validate the solution within 30 seconds or so.
Would love to check it out a couple months from now.
My view is, everyone has access to chatgpt and github copilot, and so the idea is to provide value in excess of what chatgpt/copilot can do. Part of that is embedding it in the UI, but (especially for internal tools, which tend to be shorter) the improvement isn't huge over copy/paste or using copilot in vs code.
However, beyond UI integration, we can intelligently pull context on related files, connected DBs/resources, SDKs you're using, and so on. And that's something chatgpt can't do (for now). The quality of response, from what we saw, dramatically improved with the right docs and examples pulled in.
And yes, gpt4 does much better on JS (React specifically) and Python. It's just whatever it's trained on, and there's a ton more JS/Python code out there.
I will fins out the next 2 weeks at least. I hopw this will change how i program
The only similar product i know that has this approche is darklang.
Do we as a industry have a good name for this yet?
(Im not talking about other low code platforms where code is secondary. Both darklang and windmill is code first. Giving them a huge advantage when doing thinks like AI since they can test on real data really fast. Making the turnaround speed and time to production potentially really low)
Since we control the code and the execution we can do a lot of interesting things with sending specific context to the LLMs like this "Fix error" feature.
As an open-source project ourselves, it is pretty obvious next step!
However, OpenAI pinky promises they don't use API data for anything, like training. Maybe that makes you feel a bit safer, although probably it shouldn't.
So I think we shouldn't trust them that much.
It doesn't. I don't trust OpenAI or Sama. Frankly, I'm even hesitant to use VSCode now, even with its customizable privacy/telemetry settings (though I can at least limit its network access).
But not your entire codebase.
Depending on your company thay may or may not be a showstopper.
Langflow/Flowise (especially the former) has almost everything I need within the AI world, and Windmill has everything else. Bridging the gap between them would be pretty great.
I find it very helpful, does code completion too and has a chat interface. You can select code and ask for refactorings, comments and whatever.