HNHacker News
TopNewBestAskShowJobs

guybedo

1,123 karma · joined March 3, 2018

https://akalea.com https://kodfactory.com
submissionscomments
guybedo··on Claude Opus 5.5
i think we shouldn't mix things here.

Opus 5.5 isn't the frontier, when they say 'pacing the frontier', it's about internal models not yet released, as they're probably one or two generations ahead already.

guybedo··on Claude Status – Elevated errors for multiple models
problems with Grok too ... so i guess Grok 4.7 tried to escape and took control of Colossus datacenter.
guybedo··on Astra for Law
yes, most of the time i use Astra Low already.
guybedo··on Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
SHA-256-hash-verified sealed package artifact with automatic reconciliation system p95<0.5ms
guybedo··on Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
your load-bearing thesis is probably interesting but it seems i can't read AI written text anymore -- or maybe i just need some more coffee.
guybedo··on Project HydraFusion: Frontier quality via multi-model orchestration
i've been using adversarial critique and reviews for many planning, solution design and implementation steps inside workflows.

It's so effective and helps catching so many design flaws, implementations misses etc ... that i'm wondering how people manage to build complex/large projects with agents without this kind of process. Well, i actually built this thing because i couldn't get good results so i had to find a way.

I'm gonna open source the whole thing but it needs some cleanup, there's a basic landing page here https://kodfactory.com if anyone wants to be notified when it's released on github. Yeah i know, the world really needs another software factory :-)

guybedo··on Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?
LOAD BEARING
guybedo··on Any Human Ever – One life, drawn at random from all who have ever lived
we are probably living inside a more advanced version of this game right now.

Somebody, somewhere, somewhen rolled the dice and here we are.

guybedo··on ChatGPT Is Throwing 404
Civilization IV just took over
guybedo··on GLM-5.3 is now open-weight
I have a dual epyc + 1TB RAM. I could push glm 5.2 to 7 tok/s CPU only.
guybedo··on GLM-5.3-Flash
yeah i mostly agree, especially compared to subsidized subscription cost.

But for a heavy user who has enough work to be done so that the box runs almost 24/7 at say 50tok/sec, the math gets interesting against API prices.

And it can be interesting compared to subscription in the sense that you don't have the quota anymore. That means there's probably a lot of things you're not doing because of the quotas that you could do now.

It depends heavily on the tok/sec obviously and the very best solution financially remains subscriptions. But the idea remains entertaining and not that disconnected from reality

guybedo··on GLM-5.3-Flash
although i initially thought it didn't make sense financially to run this kind of model locally, i did run the numbers and for heavy users this could justify buying $10k worth of hardware with a ROI over a few months, less than a year.

I was looking at my token usage, mostly from subsidized codex/grok subscriptions and i'm a somewhat heavy user. The thing is i would actually use even more tokens if it wasn't for the weekly quotas.

In the end, with a $10k investment and running this kind of model, estimating a 2x increase in token usage because i wouldn't have weekly quotas and comparing to glm api prices, this thing could pay for itself in less than a year.

Obviously i'm paying subscription price right now, so the math doesn't work. Although using local ai removes all weekly quotas. Keep a subscription to have access to frontier models for planning work, and local hardware + glm-5.3 flash for implementation, e2e testing, qa work 24/7.

It's not that crazy of an idea and the numbers aren't that bad.

guybedo··on The Vibe Tax
i'm not sure why people expect agents to one shot everything to perfection with just a prompt.

There's a reason why we talk about software development lifecycle, design, architecture, testing ... It's because it's been the most reliable way to build and ship software. We shouldn't expect discard this and expect agents to perform well outside of this.

I'm treating LLM agents as junior devs who happen to have vast knowledge of software engineering. As their team leader i make them go through planning, implementation, bug sweeping cycles using strict workflows. And it works quite well, i've been working on several large projects (1M+ LOC java,typescript,c/c++) and by any measure the projects are healthy. Sure the code isn't that beautiful, sure i'd have written things differently but it's pretty good nonetheless.

Shameless plug here: i've been also working on https://kodfactory.com, the code factory i've built to work on these large projects with workflows, reviews, etc ... I'm cleaning things up to open source it later.

guybedo··on Eight Myths on Software Engineering and GenAI
although, if i'm out of tokens and have to wait a full day, i won't bother doing some things manually because the day i'll spend doing something won't take more than 1 hour the next day when tokens are available again.
guybedo··on What's the largest software project AI can complete on its own?
i've been working on several rather large projects these past few months, and i'm trying to write as little code as possible.

I don't think i wrote more than 10 lines of code in the largest project i'm working on. Lines of code: Java: 900_635, typescript: 725_418, C++: 180_445, Dart: 96_181.

It's been obvious from the start that no model, as good as it is, can do large(-ish) amounts of work by its own without supervision, control, criticism, etc ... If left unsupervised, models usually do half the work, leaving stubs and todos everywhere.

Quality comes from applying software engineering principles as much as possible, just like you would do with teams of junior devs: planning sessions and implementation sessions with adversarial critiques, specifying as much as possible upfront, planning unit/smoke/integration tests, etc ...

Many systems rely on swarm of agents to build software but i've found it very difficult to get good results without lots of overhead/token waste because of inter agent communications mostly.

So instead i built what is mostly a workflow engine to structure / organize processes into workflows with different agents assigned different roles. I've setup a basic landing page here https://kodfactory.com if anyone wants to follow along.

guybedo··on Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
quite funny to see HN crowd downvoting this.

It looks like some people have a hard time accepting that software written with ai isn't a fad, it's here, it won't go away and it can be interesting and useful for the creator and for users.

Many people on the other hand have moved on and are now team leaders, except their team is mostly AIs instead of junior devs. Trade off: code is usually worse, but in the end AIs are more capable with vast knowledge and speed.

That doesn't mean we have to use AI everywhere, all the time though.

guybedo··on Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
are you using a code editor ? a compiler ? something that you didn't write yourself in pure assembly ? are you a real software engineer then ? where's the limit ?

I don't know why some people are so angry at AI/software writing with AI. It's just like being a team leader with junior(-ish) devs on the team. You don't write the code yourself, you give directions, you help/refactor/optimize where you can, that's the job.

Yes sometimes you need a team to do something, you can't code everything by yourself.

For reference, i'm not OP.

guybedo··on Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
i thought, here on HN, we were past the "oooh it's written by a LLM it's bad!".

I care about the craft, well designed systems, good clean architecture and code, etc...

But i also care about reaching goals. Whether i do it working on my own, or with human coworkers or with AI coworkers doesn't matter that much to me. Yes, the result is sometimes the most important thing.

guybedo··on So you want to use plants to reduce CO₂
it looks like i'm guilty of having never really thought and done some research about this and i'm a bit ashamed to admit i fell in this trap and used to think forests and trees are the most important pieces in the system.

I don't get why media has been fixating so much on trees and forests (Amazonia!) without even mentioning ocean's role in this.

guybedo··on Advancing the price-performance frontier with GPT‑5.6
gpt 5.6 luna was already at the intelligence/cost frontier and it's now even cheaper ...
guybedo··on The new rules of context engineering for Claude 5 generation models
so yeah we should pretty much do as we would with a junior team member:

- we should try to give good non self contradicting guidance

- we should expect the team member to have knowledge of the craft

- we should focus on higher level, taste and preferences

guybedo··on Claude Opus 5
Looking at intelligence vs cost:

- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost

ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...

guybedo··on “We have information that Moonshot distilled Fable for the development of K3”
i have information Anthropic distilled thousands of books, articles, etc ... with their author consent.
guybedo··on Making
i guess there's different point of views here and it depends a lot on what you're trying to do.

As a swe i've always enjoyed the process of designing solutions to problems: design a great architecture, find the right abstractions that make everything fit naturally, write good clean code, etc ... that's what i enjoy doing.

Now there's the part where i've had these ideas for years that i've always though that i'd be cool to work on. But basically it would have required a 10 person team for months, out of reach. Now i have Claude, Codex and Grok and i've already built many projects.

The code isn't pretty, i didn't enjoy it as much as i would have if it were done manually, but i did something i never could have done otherwise.

There's some fun in the process though. I feel like a team manager and i occasionally step in to push for some architectural changes or code rewrites because things become too brittle because of not good enough architecture.

guybedo··on GPT-5.6
it seems terra is pretty much useless, you either want luna max for everyday coding (cheaper and same perf as 5.5 high), or sol xhigh/max for demanding tasks
guybedo··on GPT-5.6
we probably need to use gpt sol max to decide which gpt flavor and effort we need to use per task.
guybedo··on GPT-5.6
It's good to see labs taking into account the cost/task.

Grok 4.5 is interesting because it's smart enough at great price. It seems gpt 5.6 is right there with great efficiency and great pricing.

Working with Fable has been a great experience, but at the end of the day, if you can get only 10% of your work done because it just burns through tokens, that's not that interesting.

I've been mostly using Opus and Fable high for planning and codex 5.5 medium for implementations. Claude is also the only model i can use for design tasks. If gpt 5.6 can finally deliver on the design side, it might be time to ditch the Claude sub and go full Gpt.

guybedo··on ZCode – Harness for GLM-5.2
if you're going to try this one out, don't be surprised to get this message repeatedly, like 4 out of 5 prompts you're trying to send, 24/7, this is gonna be your new friend, then you'll learn to write the only prompt that matters: "retry", "retry", "retry"

Here's the message: "Cannot connect to API: write EPIPE"

guybedo··on The Return of Aspect Oriented Programming
AOP is an interesting pattern but i've mostly tried to stay away from it mostly because:

- code readability and maintainability takes a hit. If you don't know things are defined using AOP in files x,y,z you can read the code and miss a whole lot of things.

- AOP implemented at runtime is a mess when you're trying to debug things

So yeah, instead of having aop defined somewhere else to wrap a function call, i tend to prefer doing it explicitly transaction(function())

guybedo··on GLM-5.2 is a step change for open agents
yeah Kimi K2.7 was doing ok but was painfully slow. The coding plan limits were good though.

I haven't tried deepseek yet, i should check this one out.

Page 1 of 7Next →