HNHacker News
TopNewBestAskShowJobs

lukasco

19 karma · joined March 31, 2023

submissionscomments
lukasco··on Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?
Guilty as charged.
lukasco··on ChatGPT and Codex Is Down
probably the cause, because AGI still can't get releases right.
lukasco··on Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?
Codus interruptus
lukasco··on With OpenRouter, Is Stripe Becoming the Amazon of AI
Interesting. I feel like Cloudflare is also having a good go at the same target, though from the infra direction. Doesn't stop Stripe of course.
lukasco··on DeepSeek Harness developer preview
One thing I found was tool calling in GLM 5.2 on github using Claude Code was mostly failing then fixing. But I didn't compare to z.ai's harness. (Nothing to do with Deepseek, sorry.)
lukasco··on Show HN: Cadre.rocks A local-first board for orchestrating coding agents
Does this solve the problem of which sessions are running what? I have the challenge all the time.
lukasco··on How Claude marks AI-generated content
Interesting points. Building on your "bits of freedom" point, given they are doing this to comply with the EU AI Act, it's also possible that the algorithm is quite weak. And they could play all kinds of games, such as embed it in the session data overall, not just the output (I don't know any details, so just guessing).

And thinking out loud, they could be really horrible and embed by using unicode characters instead of ascii, which would give a lot of flexibility, but would make the result almost unusable (but easy to defeat).

lukasco··on My Agent Setup
I've been building a triage agent for my inbox and whatsapp (it's product shaped), which has ironically left me not building one of these. So even while productizing, I'm getting fomo on the full monty.

I've also been building a harness that maintains my apps which I'm hoping to open source.

Hard agree that these things don't have personal ROI, and are actually quite hard to build reliably.

But it's really fun! And having a bot fix a live error is pretty exciting.

lukasco··on Auto mode is now the default in Claude Code
I've been running auto-mode for quite a while. The main thing that pushed me over the edge was constantly being prompted in the accept-edits mode because of back ticks or backslashes in tool calls. There was no way to approve those preemptively so I was just sitting there saying approve, approve, approve and having to babysit CC.

I've not had any problems with auto-mode so far.

lukasco··on Managing AI Coding Costs at Scale
With self-driving agents, the costs stop being evident as you go, and show up after the fact. I've been building governors that slow the agents down, and then also look at odd things some of the harnesses do, such as constantly making mistakes in tool calls.

But overall, it's quite a task, and you really have to decide what you are prioritizing for. Do I want my agents doing lots of work, or (in my case), leaving some of my subscriptions available for me to do work.

lukasco··on Show HN: The Channels SDK – Bring Any Agent to Any Channel (Slack, MS Teams)
This is one side of the problem: where does your agent show up.

And the other side is: where does the agent run

And the third side ;-): security.

Well done on this bit. I like the smooth onboarding.

lukasco··on Rust-lang/rust is adopting an LLM policy
Thanks, but I did read the policy. This is what they say: "It's fine to use LLMs to answer questions, analyze, distill, refine, check, suggest, review. But not to create."

And later:

"There are very strict guidelines on LLM-generated code changes:

Pre-arranged, non-critical, high-quality, well-tested, and well-reviewed code changes that are originally created by an LLM are allowed, with disclosure."

That's a pretty limited wedge of allowed cases in my opinion.

lukasco··on Rust-lang/rust is adopting an LLM policy
I think LLM policies which ban usage are ultimately self-defeating. They neatly switch concerns about quality of contributions to concerns about "AI slop". These two things are not the same. But the biggest problem is that they ignore the (unevenly distributed) future.

The future: SWEs will no longer write code by hand. The era of hand coding is over. We only saw that for sure in the last 6 - 9 months. But it's over.

Many coders have not yet made this transition, true, but it's happening regardless. You may not like it, but as a coder, you will not be hired if you refuse these tools. You will look as ridiculous as an accountant that refuses to use spreadsheets.

For now, we're not there yet.

As a profession, we are learning how to deal with the massive influx of review work that occurs because of LLMs. The bottleneck has moved. And as of now, there aren't good answers. But there will be. We will figure it out, like we figured out CI/CD and agile.

So personally, I'd focus my policies much more on dealing with the issues that this new era presents, and how to make LLM contributions work, and I'd avoid banning LLMs.

lukasco··on qm – Multiplayer agent harness for work
> Would love to see a 'QM vs Cowork' comparison!

It feels like the big thing they are touting here is the shared company brain. Not clear to me though, how that brain is developed when each person has their own harness. (I did only skim the docs.)

And it still has the issue of: if the agent is acting as me, then security wise it can do anything I can do. Maybe that's why they are recommending for startups.

But yes, comparison would be helpful.

lukasco··on Show HN: How to build and self-host a code review agent
Do you mean using open source models via something like OpenRouter, or do you mean self-hosting the models?

That self-hosting bit is definitely something I'm interested in, but it still seems very fiddly (in general).

I've had good results just plugging in models into Claude Code, however (though not self-hosted).

lukasco··on Show HN: What should the GUI for AI agents look like?
For my autonomous agents, I'm using Github Issues as the UI. And yes, I know this is a total abuse of what a github issue is supposed to be. But it's better than spinning up an ephemeral web site, or hooking into Slack or Telegram (at least for now). We'll see how long it works or if I think of something better.
lukasco··on Show HN: Claude-account – switch Claude Code accounts without logging in again
Can you be logged into two accounts at once? It's not the switching, but the different sessions on different accounts that I find myself doing.

As some have said, easy enough to build yourself, but I think it' nice not to have to test it and make sure it works.

lukasco··on Elon Musk's xAI sues Minnesota over law to ban 'nudify' apps
It's that pesky ban vs enforce ban dichotomy.
lukasco··on I've been building an agent to maintain my apps
Author here: My first version ran alongside my Claude Code subscription, and pretty much banged through my entire weekly subscription in a couple of days. So I bought a GLM subscription and ate through that at the same pace (even though I only used it for the agent).

I've done a bunch of optimizing of the agent since then, like not running in peak times and throttling. But without controls and optimization, the agent will chew through all the tokens you can throw at it.

Put another way: people who say they have agents coding 24/7 are on higher budgets than me.

lukasco··on Ask HN: Is it just me, or is software buggier across the board?
Shipping a lot faster definitely puts a big strain on the QA side of the house. But I'm not at all convinced that things have gotten worse even with "AI slop". CI/CD and releasing features more quickly has long been proven to make software more reliable rather than less. But to work, it requires much much better test coverage. CI/CD and AI code go hand in hand. It's all about the test coverage. Not enough tests, then exactly as you are saying: ship 10x the bugs (which actually means your software will fail catastrophically). And you mention leadership (and others did too). Here's my view: if your company leadership doesn't care about software quality, then it doesn't matter if you used Fable, or agile, or 10 year old agile, the software won't be great. You'll never have the time to do things well.
lukasco··on Towards a harness that can do anything
I still find it hard to develop harnesses because you can't really test many turns with an llm in the middle. I suppose doing LLM as judge is one way to start to tackle this kind of thing.
lukasco··on GitLost: We Tricked GitHub's AI Agent into Leaking Private Repos
I feel like in some ways this problem is starting to self-correct, sadly by creating the dead internet. If there's no business model to creating content since it will get scraped, then no content will get created.
lukasco··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
So many comments here talking about the downsides. The only reason to do a rewrite is because there are massive upsides. Maybe the implicit point is that the upside (memory safety must be the biggest), isn't worth the downside (lots of bugs to be figured out before you trust it).
lukasco··on Postgres rewritten in Rust, now passing 100% of the Postgres regression tests
Quite a lot of projects are trying this "rewrite to a new language using LLM", both internally, or externally (like is here). For me, they confirm some (slightly controversial) takes.

1. human code reviews are dead. We don't yet know what's next. Two reasons they are dead: too much code to review, and code reviewing sucks (who wants to spend their days reviewing code?) 2. Not knowing how to review LLM code is a big barrier to adoption, but bigger regression test suites (testability/evals) is almost certainly the direction. 3. There are a lot of projects that haven't moved to more modern infra because it was too hard. Now it's much easier. Sure stuff will go wrong. Sure it all has to be tested. What's new here? 4. Programming languages for LLMs are coming. 5. Projects that don't allow AI coding will be forced to come around or fade.

Separately, bit off topic:

New projects will often have LLMs built in, so non-determinism will be inherent in the project. No amount of code review will be able to eliminate that.

lukasco··on GitLost: We Tricked GitHub's AI Agent into Leaking Private Repos
This is the real problem with LLMs. There is no way to separate code from data. At best, models could be trained on tokens that indicate untrusted data coming in. But then the untrusted tokens could also be messed with.

I've wondered if it would be possible for there to be two input streams: 1, for prompt, 2 for untrusted data. But I suspect that transformers would still only optionally decide what each one was for. So it would still be a prompt level suggestion, rather than a hard and fast rule.

lukasco··on Ask HN: Is GitHub preparing to go behind a login wall?
I mean, I'm guilty too. Claude code scrapes everything in sight whenever I search for something.

I've written a small tool to query Github to see if a particular feature has been released for an open source project. I had to use an API token to not get immediately blocked. But I don't think rate limits are new for Github.

lukasco··on Zig: All Package Management Functionality Moved from Compiler to Build System
Cross platform building and packaging in C/C++ is such a hot mess. There's so many dimensions to unpack I don't even really know where to start. (I say this as the person who has been packaging GIMP for Mac for the last good number of years.)

OK, here's a few (with a MacOS slant): - compilers (gcc, clang, and their many versions) - libc (and friends) compatibility (I can't say I even ever delved into this one, but it's bit me) - package manager (macports, homebrew) - building for backward compatibility; what's the earliest MacOS version to support (the package managers either like to build for the OS they are running on, or force you -- yes I'm looking at you homebrew) - dependency and dependency version management (love you pkgconfig) - build system for each package (cmake, autotools, meson, ...) - bundling everything into an application - turning that application into a Mac application - code signing and notarization - creating the DMG - debug symbols - crash detection and notification Like right now, libheif on 26 can't be built for 11 and it's not clear why (or maybe they just fixed it...but it's been weeks)

lukasco··on The short leash AI coding method for beating Fable
Same here. Before auto, couldn’t handle the constant stopping because of backslashes no matter how many things I permitted.
lukasco··on Better Models: Worse Tools
Yeah, that looked pretty cool. But I’m always sceptical of these announcements until I try them.
lukasco··on Better Models: Worse Tools
Totally agree. And really, we use the llms to be the universal layer, not the harness. It’s already nigh on impossible to eval a harness with multiple turns (at least as far as I’ve seen), so multiply by specific llm prompt…)
Page 1 of 2Next →