HNHacker News
TopNewBestAskShowJobs

taosx

441 karma · joined February 6, 2016

Hey there! I'm a Software Engineer crafting robust, efficient solutions.

Weapons: • TS/Node.js, Rust, Go • Full-stack problem solving • Scaling, optimization & software design

Interests: • DevOps/Infra • Distributed Systems • Self-contained low-latency system designs: - Embedded databases - Modular monoliths - Stateful apps

Ping me: tasos at eveid.com

submissionscomments
taosx··on Make Tmux the OS
You are right, I didn't look too much through it and just went ahead and installed/configured it to figure out if it fits.

The vram usage is: niri = ~163mb

--

noctalia = ~221mb ashell = ~2mb ironbar = ~10mb

taosx··on Make Tmux the OS
I just wish I had a polished shell experience that was not a vram hog. Quickshell (multiple big monitors) seems to hog a lot of it, ashell and ironbar are better on resource management but require a lot of customization.
taosx··on ReBarUEFI: Resizable BAR for almost any UEFI system
On my linux machine chrome was crashing almost daily before activating rebar by patching bios.
taosx··on Measuring the sloppiness of code
I have no idea where most people writing code have worked at but in all product and platform teams I worked at the code quality has been much higher than the latest slop SOTA llms can output.

TLDR: coding is not solved.

I have 2 projects, one it's a distributed platform, the other one is a general processing engine with an inner workflow engine; Since gpt 5.2 I've tried new models to work in these codebases where the code is of good quality and every time I gave the model a slice of work instead of a single step from that slice the code, the tests, the comments, the docs and everything else has been suboptimal, unmaintainable, complex, bloated and just slop, unless I micro-manage and do many passes.

As a dev when you make a change you consider the broad picture, you consider the user, the codebase, future requirements, maintainability, performance, your team's understanding and some of these you do unconsciously. We are slow but that's for multiple good reasons, you push the organization/understanding forward not just loc of that specific project. I can't count how many PR notes or comments I've added considering teammates or just for a specific team member.

I don't see any way forward for an LLM to reach that unless it reaches general problem solving, my definition of GAI that could tackle software development or "coding" would be a model that doesn't require additional pretraining to solve new tasks or improve how it solves tasks in the future, it would just learn as it's going.

Can everything I mentioned be solved with current generation of LLMs and lot's of markdown and gates? Maybe... but the amount of effort required would be similar to the effort an expert system (pre-llm AI) would require to embed the rules, evolve them, check them everytime... which would require billions or trillions of tokens.

---

off: I really like the discussions around how to prevent slop and bloated code as it's something it would benefit coding even without LLMs and can fit as another piece of automated infra for checking and ensuring code quality, I hope something materializes.

taosx··on Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces
In the past I managed to get measureable performance optimizing a harness by looking at few traces to see if the traces contained surprised, a lot of text in order to figure out how to use my custom tool, then renamed the tool, changed some parameters and it was already great across around 20 eval tasks in rust/typescript, I repeated the same more recently but I used an llm to look at the traces... didn't achieve the desired result, mostly due to how cost-prohibitive it's for me to run expensive models.
taosx··on Kakoune Code Editor
Does any of them support more advanced functionality that can be provided with plugins? Similar to how, let's say, Emacs has slime for lisp?

From what I see helix has stopped being developed, probably ppl think its done but i feel its half baked without a plugin system and tons of basic functionality missing.

taosx··on Pi coding agent: config folder is out of place on Linux
Fully agree with this. I did that for pi for a while, maintained it and brought merges from upstream while having my own patches on top but then went on holidays and there was that refactor where the agent pipeline failed... now I'm stuck on the version from april/may (works great but I can't use extensions; and I'm too lazy to debug/fix while everything works great).

off: I'm working on my spare time on a code mode lisp alternative (great opportunity to learn lisp) and might switch to it fully as long as I built some simple evals (I'm concerned about token usage, which is why i forked pi the first time)

taosx··on Anthropic Risk August 2026 [pdf]
So I can take that as ~infinite amount of tokens don't don't even get you 2x on any novel tasks? No auto-researcher, no rsi..

Is that correct?

taosx··on DeepSeek V4 Pro 0813
Done, I'll take any other suggestions and apply them later, I will also split it a bit for different usecases as this was initially a throwaway prototype but found it useful. Basically it needs a bit more human touch.
taosx··on Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
Usage? Not exactly. But I tried to make something that can estimate dollars per tokens in actual usage while taking into account multiple factors.

https://harness.eveid.com/lazy-harness-cost-simulation

taosx··on DeepSeek V4 Pro 0813
I created a simulation for coding harnesses based on my own pi sessions. When taking into account all factors, DS-v4-Pro is cheaper than gpt-5.6-luna due to caching. Look at the bill segments difference for cache read cost and uncached cost between deepseek and the other models. At this point is cheaper to use ds-v4-pro than the luna models from openai.

ignore the numbers except the classic and keep in mind that classic is based on pi with the only change limiting tool output to 10kb

https://harness.eveid.com/lazy-harness-cost-simulation

* I built this for getting an initial estimate between different checkpoint/ compaction methods for the harness.

taosx··on OpenAI launches ChatGPT desktop app for Linux
I've been extracting the asar and using it for a while on linux, I rarely start it but it's nice that I can start a session in chatgpt and then move on to codex in the same UI.

It's packed with features, voice mode, browser. I'm not really the target audience..

taosx··on Xbox goes down. You can't play games you own on disc
No owners, no responsability, no incentive to fix it. Ticket grabbers and checklist markers.
taosx··on Ask HN: How are you using AI to learn?
I really want an AI tutor (ai learning harness) that I can give it a topic, it can synthesize information from sources, gauge my level, create lessons that I can read, give me exercises, track repetition optimal timings and track my progress.

I've noticed that for me it's much faster to learn by doing exercises targeting the exact thing you need than just trying to apply that to a much bigger chunk of work directly, especially for someone that's a bit of a failing perfectionist.

I was just looking for that and I saw https://github.com/THU-MAIC/OpenMAIC. I haven't tried it as I'm trying a few small experiments to see if I can create a curriculum for a framework with no docs.

The first naive one failed for multiple reasons but mostly because many guardrails are needed and we need a proper process that tracks too many things.

taosx··on The New AI Superpowers: Focus and Followthrough
I will never believe the premise of 100x boost from AI, why are we still pushing this narative? Yes, AI is amazing in small, constrained, focused pieces of code but the code is nowhere fit for production and the last 10% needed to ship shows that the 90% that's already been done by AI is absolutely trash. Hacks on top of hacks.

Is no one trying to ship polished things to customers anymore, will each of us have a hacky, bugged version of the same thing with different quirks?

--

My first ever software project even before I worked as a swe had less bugs and frictions than all recent projects where AI was used. Some people like mitchellh seem to know what they are doing (I have not taken a look at ghostty's codebase as it's in zig) so I sometimes get the feeling I'm holding it wrong but in the end everyone around me seems to have similar problems with shipping.

taosx··on Qwen 3.8
Exactly, and even when humans make bad decisions there is some friction, I feel that llm's don't have/notice that friction, they just bulldoze without caring about anything else.
taosx··on I burned all my tokens researching how to save tokens
This is exactly why people shouold share what they are doing with ai, so we know whwn it works, the scope, tech...etc. Thank you for sharing.
taosx··on Qwen 3.8
I'm not sure about "useless" but from my experience agentic coding leads to death by a thousand cuts for all projects I've seen so far. Small decisions missed in a codebase that leads to degradation in correctness, reliability and performance. At some point it only takes one engineer to be careless, others skipping PR because they are AI generated...
taosx··on EEG shows brain can simultaneous encode two speech streams
I had the same experience after almost 4+ days going without sleep. A friend came to check up on me after I fell asleep, I woke up and started telling him a whole story that made no sense, but I said it with such importance that for a week he was asking me to explain to him what I meant...which I don't remember exactly, but it was important.
taosx··on Pope Leo XIV’s first encyclical Magnifica humanitas to be published May 25
You assume that exploitation and material improvement can not coexist. You can be exploited just as well, by that I mean you're not getting a fair share for what you contribute to the system.
taosx··on Bun's experimental Rust rewrite hits 99.8% test compatibility on Linux x64 glibc
That's amazing, over time I got a few memory related crashes w/ bun but have deep respect for the performance work put in. Hopefully Rust's compiler will help even more.

Off: I'm wondering if now when more JS finds place on our machines and bundle size is 2nd place for most, would a revival of prepack or projects in the same vein would be worth it, especially with agents.

taosx··on OpenWarp
I feel this is the wrong way to go about things and I agree that it rude. Why not start by engaging with the warp project and see if some of this work could be upstreamed and if you like warp, target longevity?
taosx··on Zed 1.0
Congratz to the team. I really like zed and started using it quite early, loved the text threads and was using them a lot as I don't think llms fit in a box of only agents, they were a nice way to manage conversations, work through them, edit responses to lead the agent better, copy-paste full text, sad to see them go (text threads).

I'm trying right now the ACP with my own agent and I'm of mixed opinions but that's maybe because I care how my agent works. I believe that for the agent view a plain buffer with small ui elements would be the best ui for an agent conversation but I may have been spoiled by their text threads. I may spin a personal fork but the thought of tens of mins of compile time isn't that attractive.

Edit: I realized I started moving to terminal based editors like helix due to agents: claude -> codex -> custom pi, with the open sourcing of warp I was considering making a native integration for warp + pi but now I'm thinking zed's text threads (~17k lines) + pi might be a better way, any thoughts or ideas?

taosx··on DeepSeek v4
MErge? https://news.ycombinator.com/item?id=47885014
taosx··on DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
I estimated that even with heavy usage it would cost your around 30-70$ depending on caching at around 40M tokens. That would give you around double the usage compared to gpt-5.5 on the 200$ sub
taosx··on DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
So the R line (R2) is discontinued or folder back into v4 right?
taosx··on Measuring Claude 4.7's tokenizer costs
Claude seems so frustrating lately to the point where I avoid and completely ignore it. I can't identify a single cause but I believe it's mostly the self-righteousness and leadership that drive all the decisions that make me distrust and disengage with it.
taosx··on Show HN: I took back Video.js after 16 years and we rewrote it to be 88% smaller
Are there any plans to support other frontend frameworks? If I wanted to use it today in something like svelte how should I go about it?
taosx··on OpenCode – Open source AI coding agent
I've forked it locally, to be honest I haven't merged upstream in a while as I haven't seen any commits that I found relevant and would improve my usage, they seem to work on the web and desktop version which I don't use.

The changes I've made locally are:

- Added a discuss mode with almost on tools except read file, ask tool, web search only based no heuristics + being able to switch from discuss to plan mode.

Experiments:

- hashline: it doesn't bring that much benefit over the default with gpt-5.4.

- tried scribe [0]: It seems worth it as it saves context space but in worst case scenarios it fails by reading the whole file, probably worth it but I would need to experiment more with it and probably rewrite some parts.

The nice thing about opencode is that it uses sqlite and you can do experiments and then go through past conversation through code, replay and compare.

[0] https://github.com/sibyllinesoft/scribe

taosx··on OpenCode – Open source AI coding agent
The only thing I'm wondering is if they have eval frameworks (for lack of a better word). Their prompts don't seem to have changed for a while and I find greater success after testing and writing my own system prompts + modification to the harness to have the smallest most concise system prompt + dynamic prompt snippets per project.

I feel that if you want to build a coding agent / harness the first thing you should do is to build an evaluation framework to track performance for coding by having your internal metrics and task performance, instead I see most coding agents just fiddle with adding features that don't improve the core ability of a coding agent.

Page 1 of 7Next →