Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
Still works better than compaction.
On another environment I've been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.
Mentioned by the author in a recent HN thread, I'm also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.
I assume Anthropic & friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.
Same teamates also post 'Sol deleted my git repo!' or 'Sorry, ignore those 300 PR comments i was just looking!' ~once a month.
I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.
The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.
For almost a year they didn't allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It's a little better now, but it's still painfully behind the curve.
Do you know a single product from Infosys / Cognizant / Tata done right?
agy --dangerously-skip-permissions
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.
IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.
accept-edits
also available with shift-tab in agy cli [0]. That is not related to executing commands, only to allowing agy edit files. Unless you set Turbo-mode == yolo-mode, agy gui prompts a zillion times too.With 'some features added' I'm referencing this but don't see change in daily work (still many prompts making getting-work-done impossible): v2.14.0 (September 15, 2026) "New Permissions System" [1]
Details: Introduced the new unified permissions system, presets (Default, Request Review, Turbo), syntax-highlighted permission requests, and restructured the settings under Global Permissions and project-level Inherit Global.
[0] https://antigravity.google/docs/cli/modes/#available-modes [1] https://antigravity.google/docs/changelog/
They've got Zed, VSCode, Jetbrains... But no Emacs or NeoVIM
I would rather not risk my Google account.
I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.
1) Try to integrate agy into a workflow. It can't do standard I/O like: tail -200 app.log | claude -p "Find the problem"
2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.
3) Not open source so I can't fix any of these problems.
4) No skills. In 2026. Yikes.
5) No persistent memory (see Claudes auto memory)
6) No sub-agents or orchestration of any type really.
7) Weird hard coded limits and constant API errors on everything (scaling problems?)
8) No /loop command
9) /btw is weird and ephemeral. No way to merge it back to the conversation.
10) Unstable in general.
11) No way to control it via API.
I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it's so worth it. Then if you want try to go back to agy. Don't worry, agy won't have changed much. It improves at a snails pace.
I'm not sure this list is correct. Number 4 is especially wrong, since Skills are available with the launch of Antigravity 2:
https://antigravity.google/blog/introducing-google-antigravi...
edit: removed persistent memory from list since I realized I'm using a plugin for that and it's apparently not native
Luckily for me both expired yesterday and I was able to subscribe back again (first Youtube family and then Google AI plan).
I like CLI more than TUI, and TUI more than GUI where appropriate. For working with text TUI is better, for e.g. images GUI (GIMP).
All the AI-generated UIs I've seen have been very derivative, certainly not eliminating any of the disadvantages of typical GUI interfaces.
Realistically, the whole "overlapping windows" GUI model, and everything that derives from that, was a metaphor geared towards people who'd never seen a computer before. It fit the increasing consumer focus of computing interfaces. It's no wonder that technical people often prefer TUIs.
Maybe AI will bring real advancements in GUIs, but someone's still going to have to make it happen.
That is what we need, and if you're making the argument that a terminal shell is the best place to provide them then I don't know what to say.
Despite the capitals and trying to come across as someone who SEES "what is COMING" (yay!), you have absolutely no idea what you are talking about.
No wait, I will leave you with a hint. Do whatever you wish to do with that.
Hint: So, you probably want me to use a toolset which is incredibly inferior, as of today, for "me/my usage", just because you feel "IT CAN BE". Right.
Perlis has an aphorism for this, as he does every important problem [0].
In today's world - and idea stated stated is an idea stolen .
It's missing the point. I mean think about the basics, why open the huge bash hole only then to have to close it? If you think about it logically, the only way you can sandbox bash is by writing your own bash implementation specifically for agentic use cases.
This sort of thing alarms me. Having a $bigcorp account becomes a ""social credit"" system where they can ban you from all your personal stuff if they decide that you (or your agents!) are doing stuff they don't like.
I don't want to mess with antigravity because my google account is too entrenched in my life.
without having an entirely separate google account with its own separated bans, theres just no ability to trust those
Yes it's the same with Claude. However, OpenAI allows you to use any harness you like. Which makes sense and that's the primary reason I have their plan now rather than Googles.
EDIT: I checked online and I couldn't find any report of complete banning for using third party harnesses, only Gemini service suspension, but It's never hurts to be careful. If my secondary email gets banned, I should be able to use my main email on agy.
It also makes me think if Google Family with Google One could be abused for extending inference limits.
Not only could it not complete the small task, the code was obviously wrong from looking at it and did not even compile.
When I pointed that out it got pissy and insisted the code was perfect and I didn't know how to use a compiler, or the compiler was buggy. Pasting the compiler errors did not help.
Surreal.
Is that a reference to https://xkcd.com/353/
I have a client app on a very old (for the JS world) version of eleventy using NetlifyCMS (also outdated). Claude has quite easily picked that up to add features to it along the way.
I don't think their point was about knowing the stack, but being able to point a harness at something running on your desktop GUI and say "change this".
Being able to edit and recompile pretty much any part of the OS and userland (often not even needing to reboot!) is not something that can be said about Windows for sure, or even lots of things on Macs too. Or even when the browser is effectively the operating system, the JS/TS others write is also hard to change in your end.
Even if the source code's old, the fact that it is publicly available makes it much easier to train & improve on than if it were walled off.
I used to be afraid of Arch, because I don't want a system that takes work because I'm already busy with work. But now I love it, because the LLM can tweak every knob and fix every issue for me, so it ends up being the OS that takes the least work to use. Get an error message? Tell the LLM and they fix it. Something not working exactly like you like it? Tell the LLM and they tweak it for you. This also works to great effect on Win and Mac, but not to the same extreme degree as it does on Linux and especially Arch.
Anything I want different now, I just tell the LLM to do it for me.
(For example I can now close lots of windows of the same type with 1 click, not 3, my whisker search now finds files and folders and I am able to run games that refused before)
I still ocasionally run into the usual linux driver issues, but not for much longer I suppose. I probably could fix some driver bugs now already if I point fable towards it and pay some attention.
In theory I could do all this before myself, but not just like that in some minutes, but in days/weeks/months ..
Drifted off my point a bit, which I guess was meant to be it’s not necessarily Linux-specific training.
Drivers, Coding environment setup etc. are great too and it's nice to have everything logged so the next (more powerful) agent can come and improve the thing once in a while.
Since LLMs have been so successful at finding exploits it's been clearer than ever that the people so obsessed with enshittifying every system with ineffective (other than pissing you off and wasting your time) "secure boot" functionality were really just too ignorant to succeed at actual novel security work, so they focused on this make believe crap. Well, glad that's over.
I am still very happy with my switch to Linux. But if I didn't have the AI help I would say linux is still unacceptable platform for those not willing, able, and excited to get their hands very dirty.
I'm using it to run overnight tasks, and that's it until my quota runs out.
Canceled my subscription.
its a lot less chatty imo
In the local app interface the winning choice is “efficient” and then turn off the sliders for warmth enthusiasm emoji etc.
It makes openai models so good to talk to i really have trouble switching.
Also canceled my ChatGPT Pro $200/mo subscription. Their Oct 30 price hikes and slow GPT-6.1 model has me looking for alternatives.
- AI is just a tool, like excel; it does what the human operating it tells it to
- next token prediction cannot be true understanding
- models can have no desires and goals, don't anthropomorphize it
However, "hallucination" is very much not one of them
It didn't work. Gemini: "Oh yeah, that obviously cannot work, it's not possible to do it through OBD2" (paraphrasing)
It was quite funny to me, but a bit less so to my colleague.
I'm assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.
> The knowledge cutoff date for Gemini 3.8 Flash is March 2026
https://deepmind.google/models/model-cards/gemini-3-8-flash/
The "some domains" are very narrow. They likely just RL'ed popular queries.
Of course, it's called "flash," and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I'm hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let's see.
Most llm could do it. Claude went from firmware thread -> rtos scheduler -> mcu reference manual -> hardware controller register interface -> vendor sdk -> problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
Including going first for decompiling AGY binary instead of searching the web for documentation...
I've tried Codex as a harness too, and that was nice. I don't find a significant difference between Antigravity and Codex. Codex has more features but I don't use them.
completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.
once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B
I get the feeling the situation is only going to improve longer term so might be a good time to just do it
It's so fucking easy.
From AMD: https://lemonade-server.ai/
Then you can easily throw a openweb-ui container in front, and then connect to the openweb-ui via your mobile app of choice (if you want chat, otherwise you just point your harness of choice at the lemonade server api endpoint).
ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I've been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.
The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn't materialize and the token generation speed decreases.
There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.
hope it's not run by multiple threads and dlsym is not allocating.
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.
Why would the guy who wrote curl share it? We can all build our own now...
Why do the Linux folks need to be so selfless? We can all build our own kernel now...
in amdkfd and hsakmt
Yes, that means AMD is sitting on both sides. They wrote software that doesn't work with their own software.
I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.
In any case, yeah, ~55 tok/s on a high quality model (and massive RAM savings I think?), seems dope.
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
Still impressive that it did so much just to answer my simple question.