107 karma · joined December 30, 2024
Edit: sadface
The notable difference to me is tokens/sec are still much higher on 6.1 Sol
It's been amazing at making sure OOMs for multiple heavy builds on my machine don't happen, adding queues and locks to make sure performance measurements are isolated and gpu stays clean during experiments.
It's also way more able to execute subagent tasks all at once than GPT 6.1 I tried to give it 10 different subtasks all at once that were overlapping and unrelated issues and it did a good job spinning up isolated worktees, agents and then coordinating the merge back together and then verifying them with agents in batches.
https://codeberg.org/ziglang/zig/issues/37060 I opened the issue and can give you the agent generated RCA on the matter if you all want it in issue 604 on antfly's github but I understand that's against policy and totally respect that.
Appreciate all the work you guys do and have been following the whole Zig project since inception fwiw!
So for instance because of the size of our codebase our project has pushed Zig to some of the edges, specifically we end up hitting a bug when using llvm and zig on arm64 (Mac and Linux) where it seems to be caused by some configuration Zig passes through to LLVM. I’ve used codex and claude to help me diagnose and find the bug (we use nix’s glibc zig to circumvent the problem now). I now understand the root cause but am not sure what the proper fix would be. But I’ve not known whether or not even raising the issue would break the terms of contributing? Would raising the issue break the implicit agreement?
We’re quite knowledgeable and thoughtful folks fwiw
I think I’ve seen you are a core team member or contributor? I remember your tag?
At the moment the reverse is true though, the simulator and what we've captured as ground truth for the desired design has been refined enough in tests and specs that the code is often the one implicated, and most of the bugs have been in code related to caching correctness and are only exposed through soak testing.
In opposition to Anthropic/Bun, we mostly used a hands-on approach to the rewrite and took the opportunity to capture the original design of Antfly into specs and any missing tests one subsystem at a time so we didn't strive to be as hands-off as "let Claude hill-climb on the tests". Especially since the system as a whole is far more dynamic and depends more on scalability, distributed systems stuff than Bun required!
Thought it is interesting to juxtapose to the Bun rewrite from Anthropic and wanted to talk about why we went the other way! Would love to talk about our process or the technology!
Benchmarks against are linked in the article but here they are again for posterity https://antfly.io/releases/v0.2
That being said AI is not code, it's a statistical algorithm with non-determinism baked in. You can write code to run them but it's nothing without the evolution of the model weights from the training process. And you can absolutely make the model weights better aligned with intent.