The vram usage is: niri = ~163mb
--
noctalia = ~221mb ashell = ~2mb ironbar = ~10mb
441 karma · joined February 6, 2016
Weapons: • TS/Node.js, Rust, Go • Full-stack problem solving • Scaling, optimization & software design
Interests: • DevOps/Infra • Distributed Systems • Self-contained low-latency system designs: - Embedded databases - Modular monoliths - Stateful apps
Ping me: tasos at eveid.com
The vram usage is: niri = ~163mb
--
noctalia = ~221mb ashell = ~2mb ironbar = ~10mb
TLDR: coding is not solved.
I have 2 projects, one it's a distributed platform, the other one is a general processing engine with an inner workflow engine; Since gpt 5.2 I've tried new models to work in these codebases where the code is of good quality and every time I gave the model a slice of work instead of a single step from that slice the code, the tests, the comments, the docs and everything else has been suboptimal, unmaintainable, complex, bloated and just slop, unless I micro-manage and do many passes.
As a dev when you make a change you consider the broad picture, you consider the user, the codebase, future requirements, maintainability, performance, your team's understanding and some of these you do unconsciously. We are slow but that's for multiple good reasons, you push the organization/understanding forward not just loc of that specific project. I can't count how many PR notes or comments I've added considering teammates or just for a specific team member.
I don't see any way forward for an LLM to reach that unless it reaches general problem solving, my definition of GAI that could tackle software development or "coding" would be a model that doesn't require additional pretraining to solve new tasks or improve how it solves tasks in the future, it would just learn as it's going.
Can everything I mentioned be solved with current generation of LLMs and lot's of markdown and gates? Maybe... but the amount of effort required would be similar to the effort an expert system (pre-llm AI) would require to embed the rules, evolve them, check them everytime... which would require billions or trillions of tokens.
---
off: I really like the discussions around how to prevent slop and bloated code as it's something it would benefit coding even without LLMs and can fit as another piece of automated infra for checking and ensuring code quality, I hope something materializes.
From what I see helix has stopped being developed, probably ppl think its done but i feel its half baked without a plugin system and tons of basic functionality missing.
off: I'm working on my spare time on a code mode lisp alternative (great opportunity to learn lisp) and might switch to it fully as long as I built some simple evals (I'm concerned about token usage, which is why i forked pi the first time)
Is that correct?
ignore the numbers except the classic and keep in mind that classic is based on pi with the only change limiting tool output to 10kb
https://harness.eveid.com/lazy-harness-cost-simulation
* I built this for getting an initial estimate between different checkpoint/ compaction methods for the harness.
It's packed with features, voice mode, browser. I'm not really the target audience..
I've noticed that for me it's much faster to learn by doing exercises targeting the exact thing you need than just trying to apply that to a much bigger chunk of work directly, especially for someone that's a bit of a failing perfectionist.
I was just looking for that and I saw https://github.com/THU-MAIC/OpenMAIC. I haven't tried it as I'm trying a few small experiments to see if I can create a curriculum for a framework with no docs.
The first naive one failed for multiple reasons but mostly because many guardrails are needed and we need a proper process that tracks too many things.
Is no one trying to ship polished things to customers anymore, will each of us have a hacky, bugged version of the same thing with different quirks?
--
My first ever software project even before I worked as a swe had less bugs and frictions than all recent projects where AI was used. Some people like mitchellh seem to know what they are doing (I have not taken a look at ghostty's codebase as it's in zig) so I sometimes get the feeling I'm holding it wrong but in the end everyone around me seems to have similar problems with shipping.
Off: I'm wondering if now when more JS finds place on our machines and bundle size is 2nd place for most, would a revival of prepack or projects in the same vein would be worth it, especially with agents.
I'm trying right now the ACP with my own agent and I'm of mixed opinions but that's maybe because I care how my agent works. I believe that for the agent view a plain buffer with small ui elements would be the best ui for an agent conversation but I may have been spoiled by their text threads. I may spin a personal fork but the thought of tens of mins of compile time isn't that attractive.
Edit: I realized I started moving to terminal based editors like helix due to agents: claude -> codex -> custom pi, with the open sourcing of warp I was considering making a native integration for warp + pi but now I'm thinking zed's text threads (~17k lines) + pi might be a better way, any thoughts or ideas?
The changes I've made locally are:
- Added a discuss mode with almost on tools except read file, ask tool, web search only based no heuristics + being able to switch from discuss to plan mode.
Experiments:
- hashline: it doesn't bring that much benefit over the default with gpt-5.4.
- tried scribe [0]: It seems worth it as it saves context space but in worst case scenarios it fails by reading the whole file, probably worth it but I would need to experiment more with it and probably rewrite some parts.
The nice thing about opencode is that it uses sqlite and you can do experiments and then go through past conversation through code, replay and compare.
I feel that if you want to build a coding agent / harness the first thing you should do is to build an evaluation framework to track performance for coding by having your internal metrics and task performance, instead I see most coding agents just fiddle with adding features that don't improve the core ability of a coding agent.