421 karma · joined May 1, 2022
yeah this one is a bit weird, you can run the linux binary using WSL and that should work. We have a window flavor build but its not as heavily tested yet (we are figuring out a better testing flow for windows)
We worked quite heavily on the TUI, its written in ratatui and we did try to go the extra mile to make sure that there won't be any regressions from moving to an alt-screen rendering. Lots and lots of small details to manage and get correct, we also tried to get vim keybindings correct while keeping it mouse friendly.
What kind of SDK support are you looking for?
The big difference is the complete loop, each PR gets its own VM with the tool chains installed so the agent can run cargo check or cargo tests etc.
We do find the LLMs of today are not the best elite engineers but very very competent junior engineers. It's been a weird but eye opening workflow to use.
ngl the total expenditure was around $10k, in terms of test-time compute we ran upto 20X agents on the same problem to first understand if the bitter lesson paradigm of "scale is the answer" really holds true.
The final submission which we did ran 5X agents and the decider was based on mean average score of the rewards, per problem the cost was around $20
We are going to push this scaling paradigm a bit more, my honest gut feeling is that swe-bench as a benchmark is prime for saturation real soon
1. These problem statements are in the training data for the LLMs
2. Brute-forcing the answer the way we are doing works and we just proved it, so someone is going to take a better stab at it real soon
I used to think Arc/Rc was a shortcut to avoiding the borrow checker shenanigans, but have evolved that thinking over time.
You do mention it in your comment so wondering if you have anything to share about it
I know my way around this now, which is to literally binary search over the timeline of my edits (commenting out code and then reintroducing it) to see what causes the compiler to trip over (there might be better ways to debug this, and I am all ears)
Most of the times this error is several layers deep in my application so even tho I want to ticket it up, not being able to create a minimal repo for anyone to iterate against feels like a bit of wasted energy on all sides, do let me know if I should change this way of thinking and I can promise myself to start being more proactive.
Having said this, the benefits of borrow checker out weight the shortcomings. I can feel myself writing better code in other languages (I tend to think about the layout and the mutability and lifetimes upfront more now)
My rust code now is very functional, which seems to work best with lifetimes.
I would love to know more about the authors pain, I do hope rustc gets better at lifetime compilation errors cause some of them can be very very gnarly.
I would love to know your workflow, you mention CLI tool or VSCode plugin, which one of them work for you? Whats missing from them where Aide can fill in the gap
We will start including the open file by default in the context very soon (one of the gotchas here, is that the open file could not be related to the question you have)
- looking at git commits - making use of recently accesses files - keyword search
If I set these constraints and allow for maybe around 2 LLM round trips we can get pretty far in terms of performance.
I do wonder what api level access do we get over there as well. For sidecar to run, we need LSP + a web/panel for the ux part (deeper editor layer like undo and redo stack access will also be cool but not totally necessary)
Our first take on solving this is with rollbacks .. which allows you to delete edits up until a point in the conversation.. so if you do notice a bad edit you can do that.
after this, there is the proactive agent ..which checks it's work again and suggests more changes which it needs to do .. you can give feedback and guide it.
With llms we do loose a bit of control but I think the editor should work to solve this
In Aide as well, we realised that the major missing loop was the self-correction one, it needs to iteratively expand and do more
Our proactive agent is our first stab at that, and we also realised that the flow from chat -> edit needs to be very free form and the edits are a bit more high level.
I do think you will find value in Aide, do let me know if you got a chance to try it out
We implicitly take in all the diagnostics on the files https://github.com/codestoryai/sidecar/blob/e5408782a3bfa461...
The readme also talks about how LSP services are not exposed properly yet, my takeaway is that its not complete yet..but surely doable
Do you use any of the AI features which go for editing multiple files or doing a lot more in the same instruction?
What features of cursor were the most compelling to you? I know their autocomplete experience is elite but wondering if there are other features which you use often!