Goose: An open-source, extensible AI agent that goes beyond code suggestions
block.github.io
block.github.io
Here's a short writeup of my notes from trying to use it https://notes.alexkehayias.com/goose-coding-ai-agent/
Haven’t tried other models yet but old like to see how o3-mini performs once it’s been added.
My colleague has working on an MCP server that runs Python code in a sandbox (through https://forevervm.com, shameless plug). I’ve been using Goose a lot since it was announced last week for testing and it’s rough in some spots but the results have been good.
Because the example given was "change the color of a component".
Now, it's obviously fairly impressive that a machine can go from plain text to identifying a react component and editing it...but the process to do so literally doesn't save me any time.
"Can you change the current colour of headercomponent.tsx to <some color> and increase the size vertical to 15% of vh" is a longer to type sentence then the time it would take to just open the file and do that.
Moreover, the example is in a very "standard" format. What happens if I'm not using styled components? What happens if that color is set from a function? In fact none of the examples shown seem gamechanging in anyway (i.e. the Confluence example is also what a basic script could do, or a workflow, or anything else - and is still essentially "two mouseclicks" rather then writing out a longer English sentence and then I would guess, waiting substantially more time for inferrence to run.
I’ve found LLMs most useful for doing things with unfamiliar tooling, where you know what you want to achieve but not exactly how to do it.
On the other hand, it’s an okay test case because you can easily verify the results.
The risk of wasted time is higher than the proposed benefit, for most of my current use cases. I don't do heaps of glue code, it's mostly business logic, and one off fixes, so I have not found LLMs to be useful day to day at work.
Where it has been useful is when I need to do a task with tech I don't use often. I usually know exactly what I want to do but don't have the myriad arcane details. A great example would be needing to do a complex MongoDB query when I don't normally use Mongo.
I'll stub out tests (just a name and `assert true`) and have it fill them in. It usually gets them wrong, but I can fix one and then have it update the rest to match.
Not perfect, but beats writing all the tests myself.
Take it from Aider example: https://github.com/Aider-AI/aider It asked to add a param and typing to a function. Would that save us more time? I don't think so. but it's a good peek of what it can do
just like any other hello world example i suppose
With AI the challenge is that we need to convince the reader that the tool will work. So that calls for a different kind of example.
If the task is not simple then break it into simple tasks. Then each of them is as easy as color change.
git clone a repo
Open goose with that directory
Instruct it to discover what the repo does
Ask it to make changes to the code, being detailed with my instructions.
I haven't tried computerController, only Goose’s main functionality.
1) use better font and size
2) allow to adjust shortcuts and have nice defaults with easy change
3) integrate with local whisper model so I can type with voice triggered with global shortcut
4) change background to blend with default system OS theme so we don't have useless ugly top bar and ugly bottom bar
5) shortcuts buttons to easily copy part of conversion or full conversation, activate web search, star conversion so easy to find in history etc.
They should get more inspiration from raycast/perlexity/chatgpt/arcbrowser/warpai ui/cursor
> Make sure to confirm all changes with me before applying.
https://block.github.io/goose/docs/guides/using-goosehints
So, we're supposed to rely on LLM not hallucinating that it is allowed to do what it wants?
Yes. Frontier models have been moving at light speed over the last year. Hallucinations are almost completely solved, particularly with Anthropic models.
It won't be long before statements like this sound the same as "so you mean I have to trust that my client will always have a stable internet connection to reach out to this remote server for data?".
Besides that, you can absolutely still trick top of the line models: https://embracethered.com/blog/posts/2024/claude-computer-us...
Hallucination might be getting better, gullibility less so.
- prompt from command line directly to Claude
- suggestions dumped into a file under ./tmp/ (ignored by git)
- iterate on those files
- shuttle test results over to Claude
Getting those files merged with the source files is also important, but I’m not confident in a better way than copy-pasting at this point.
Save tens of hours in one commit?! No model is this good yet[1] - especially not Aider with its recommended models. I fully agree with parent - current SoTA models require lots of handholding in the domains I care about, and the AI chat/pairing setup works much better compared to the AI creating entire commits/PRs before a human gets to look at it.
1. If they were, Zuckerberg would have already announced another round of layoffs.
I often get hung up on UI, I can’t make a decision on what I think will look decent and so I just sort of lock up. Aider lets me focus on the logic and then have the LLM spit out the UI. Since I’ve given up on projects before due to the UI aspect (lose interest because I don’t feel like I’m making progress, or get overwhelmed by all the UI I’ll need to write) this is a huge boon to me.
I’m not incapable of writing UI, I’m just slower at it so Aider is like having a wiz junior developer who can crank out UI when I need it. I’m even fine to rewrite every line of the UI by hand before “shipping”, the LLM just helps me not get stuck on what it should look like. It lets me focus on the feature.
Also you can use local models if you want it to be “free”.
The thing is, all side projects have a cost, even if it’s just time. I’m happy to let some of my side projects move forward faster if it means paying a small amount of money.
I find Claude wants to edit files that I don’t like to be edited often. Two ways I deal with that - first, you can import ‘read only’ files, which is super helpful for focusing. Second, you can use a chat mode first to talk over plans, and when you’re happy say “go”.
I think the thing to do is try and use it at a fairly high level, then drop down and critique. Essentially use it as a team of impressive motivation and mid quality.
I do like the idea of letting the model ask for source code.
It’s all about attention / context.
I’m giving myself the option to output collated code to a file, or copy it to clipboard, or just hold onto it for the next prompt.
I know aider does this stuff, but because I’m automating my own workflow, it’s worth doing it myself.
I plan to share it on Github, but waiting for my employer's internal review process to let me open source it as my own project, since they can legally claim all IP developed by me while employed there.
Mainly - tool calling support just merged in llama.cpp (https://github.com/ggerganov/llama.cpp/pull/9639) this week, and it's been a fun exercise to put local LLMs through the wringer to see how they do at handling it.
It's been a mixture of "surprisingly well" and "really badly".
But you didn't bother checking the very next section on side bar, Supported LLM Providers, where ollama is listed.
The attention span issue today is amusing.
I find it rather depressing. I know it's a more complex thing, but it really feels irl like people have no time for anything past a few seconds before moving onto the next thing. Shows in the results of their work too often as well. Some programming requires very long attention span and if you don't have any, it's not going to be good.
So if you're going to market something to me at least do it right. My attention span is low because I don't really give a shit about this.