I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
This sort of thing alarms me. Having a $bigcorp account becomes a ""social credit"" system where they can ban you from all your personal stuff if they decide that you (or your agents!) are doing stuff they don't like.
Luckily for me both expired yesterday and I was able to subscribe back again (first Youtube family and then Google AI plan).
They've got Zed, VSCode, Jetbrains... But no Emacs or NeoVIM
I would rather not risk my Google account.
It does mean however that it cannot operate as flexibly as it it could with raw API, IMO, but agent-shell is essentially designed towards wrapping the official clients
While it may be technically allowed, I'm not about to risk my account. Google has proven themselves to be capricious and arbitrary when it comes to TOS enforcement, and their appeals system doesn't meaningfully exist in practice.
I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.
The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.
For almost a year they didn't allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It's a little better now, but it's still painfully behind the curve.
Do you know a single product from Infosys / Cognizant / Tata done right?
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
Still works better than compaction.
Mentioned by the author in a recent HN thread, I'm also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.
I assume Anthropic & friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.
On another environment I've been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.
Same teamates also post 'Sol deleted my git repo!' or 'Sorry, ignore those 300 PR comments i was just looking!' ~once a month.
What are people using 1M context window for?
agy --dangerously-skip-permissions
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.
IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.
accept-edits
also available with shift-tab in agy cli [0]. That is not related to executing commands, only to allowing agy edit files. Unless you set Turbo-mode == yolo-mode, agy gui prompts a zillion times too.With 'some features added' I'm referencing this but don't see change in daily work (still many prompts making getting-work-done impossible): v2.14.0 (September 15, 2026) "New Permissions System" [1]
Details: Introduced the new unified permissions system, presets (Default, Request Review, Turbo), syntax-highlighted permission requests, and restructured the settings under Global Permissions and project-level Inherit Global.
[0] https://antigravity.google/docs/cli/modes/#available-modes [1] https://antigravity.google/docs/changelog/
Perlis has an aphorism for this, as he does every important problem [0].
I like CLI more than TUI, and TUI more than GUI where appropriate. For working with text TUI is better, for e.g. images GUI (GIMP).
All the AI-generated UIs I've seen have been very derivative, certainly not eliminating any of the disadvantages of typical GUI interfaces.
Realistically, the whole "overlapping windows" GUI model, and everything that derives from that, was a metaphor geared towards people who'd never seen a computer before. It fit the increasing consumer focus of computing interfaces. It's no wonder that technical people often prefer TUIs.
Maybe AI will bring real advancements in GUIs, but someone's still going to have to make it happen.
That is what we need, and if you're making the argument that a terminal shell is the best place to provide them then I don't know what to say.
It's not that GUIs are inherently worse in principle, but in practice they often are.
The point about overlapping windows is that that "desktop" model permeates the thinking about GUI design, but it's fundamentally limiting and misguided.
> we need complex (but not overlapping windows, sure!) interactivity and rich display capability.
Yep. Pity today's GUIs can't deliver that.
Despite the capitals and trying to come across as someone who SEES "what is COMING" (yay!), you have absolutely no idea what you are talking about.
No wait, I will leave you with a hint. Do whatever you wish to do with that.
Hint: So, you probably want me to use a toolset which is incredibly inferior, as of today, for "me/my usage", just because you feel "IT CAN BE". Right.
In today's world - and idea stated stated is an idea stolen .
What benchmarks are usually good at is showing to what degree new models are better than old models. What they are not good at, by construction, is showing that harnesses are well adapted to how people use them.
Why? Because 4.6 actually talked like a human being. It actually organized its thoughts well, and got the main information across without the wall of text that makes your eyes glaze over. So from the perspective of human-computer interaction and maximizing the productivity of a developer+agent team, 4.7 and 4.8 were regressions. Despite much better benchmark performance.
Even if we consider autonomous agents, that benchmark is not indicative of how well they will interpret *your* requests. Or how well they will interact with other agents in a flock/swarm situation. The benchmark just doesn't cover this. (And the difference can be nontrivial! Sakana AI's published results show two generations of uplifting potential from better harnesses.)
It's missing the point. I mean think about the basics, why open the huge bash hole only then to have to close it? If you think about it logically, the only way you can sandbox bash is by writing your own bash implementation specifically for agentic use cases.
I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.
1) Try to integrate agy into a workflow. It can't do standard I/O like: tail -200 app.log | claude -p "Find the problem"
2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.
3) Not open source so I can't fix any of these problems.
4) No skills. In 2026. Yikes.
5) No persistent memory (see Claudes auto memory)
6) No sub-agents or orchestration of any type really.
7) Weird hard coded limits and constant API errors on everything (scaling problems?)
8) No /loop command
9) /btw is weird and ephemeral. No way to merge it back to the conversation.
10) Unstable in general.
11) No way to control it via API.
I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it's so worth it. Then if you want try to go back to agy. Don't worry, agy won't have changed much. It improves at a snails pace.
edit: removed persistent memory from list since I realized I'm using a plugin for that and it's apparently not native
I'm not sure this list is correct. Number 4 is especially wrong, since Skills are available with the launch of Antigravity 2:
https://antigravity.google/blog/introducing-google-antigravi...
Yes it's the same with Claude. However, OpenAI allows you to use any harness you like. Which makes sense and that's the primary reason I have their plan now rather than Googles.
EDIT: I checked online and I couldn't find any report of complete banning for using third party harnesses, only Gemini service suspension, but It's never hurts to be careful. If my secondary email gets banned, I should be able to use my main email on agy.
It also makes me think if Google Family with Google One could be abused for extending inference limits.
I don't want to mess with antigravity because my google account is too entrenched in my life.
without having an entirely separate google account with its own separated bans, theres just no ability to trust those
Is that a reference to https://xkcd.com/353/
Not only could it not complete the small task, the code was obviously wrong from looking at it and did not even compile.
When I pointed that out it got pissy and insisted the code was perfect and I didn't know how to use a compiler, or the compiler was buggy. Pasting the compiler errors did not help.
Surreal.