HNHacker News
TopNewBestAskShowJobs

ericol

1,255 karma · joined August 5, 2014

submissionscomments
ericol··on Tell HN: Claude 4.7 is ignoring stop hooks
Stop hooks are a world of pain.

I recently went on a deep dive about them with sonnet / opus.

I wanted to detect if a file or an analysis was the result of the last turn and act upon that.

From my experience, 2 things stand out by looking at the data above:

1. They have changed the schema for the hook reply [1] if this is real stop hook users (And may be users of other hooks) are in for a world of pain (if these schema changes propagate)

2. Opus is caring f*ck all about the response from the hook, and that's not good. Sonnet / Opus 4.6 are very self conscious about the hooks, what they mean and how they should _ act / react_ on them, and because of how complex the hook I set up is I've seen turns with 4 stop hooks looping around until Claude decides to stop the loop.

[1] My comment is in the context of claude code. I cannot make if the post is about that or an API call.

ericol··on Show HN: Tolaria – Open-source macOS app to manage Markdown knowledge bases
if "git versioned" means the .md files themselves, I'm sold. I am actually processing files using a git based workflow in order to tell claude what to look at.

I'll definitely give this a spin.

ericol··on Books are not too expensive
Books _are_ expensive. This article only looks at the monetary side of it.

The cost goes beyond the price tag. Books take up space, and that space compounds as you keep acquiring them. It's space you can't use for anything else, dedicated entirely to objects most people open once or twice and never touch again. And that cost doesn't stay abstract: at some point you're buying more bookshelves, upgrading to a larger one, or worst of all, dragging everything through a move. That last one hits harder the less stable your living situation is, and less stable living situations track pretty closely with lower salaries.

I'm talking about physical books specifically, since that's what the article seems to cover. Ebooks are a different matter.

ericol··on Measuring Claude 4.7's tokenizer costs
> I don't know if you're giving this as something you've actually given Claude, but I don't think it's a good way of using Claude.

That wasn't the full prompt, I trimmed it for clarity, but I agree with everything you said and that's how I actually use it.

I have a proxy logging everything sent to and from Claude in a structured way, which is precisely what let me do that compaction analysis in the first place.

When Claude goes off track, I don't tell it "you did something wrong". I ask it to analyze the tool outputs and the exchange so far and let it reconcile the discrepancy itself. That tends to work better than narrating the error to it.

The venting messages like that one are honestly for me, not for Claude. I know it's a tool. But it also behaves and communicates like a person, and that's a design choice that comes from Anthropic, not from me. What I've found is that writing something like that and then following it with proper instructions works fine in practice: Claude either ignores the venting or briefly acknowledges it and moves on. The actual output isn't affected. It's just how I process frustration without breaking the workflow.

ericol··on Measuring Claude 4.7's tokenizer costs
You are being downvoted but I actually agree with your statement.
ericol··on Claude Opus 4.7 costs 20–30% more per session
Matches what I am experiencing. Makes incredible stupid mistakes.

The weird stuff is yesterday I asked it to test and report back on a 30+ commit branch for a PR and it did that flawlessly.

ericol··on Measuring Claude 4.7's tokenizer costs
I did some work yesterday with Opus and found it amazing.

Today we are almost on non-speaking terms. I'm asking it to do some simple stuff and he's making incredible stupid mistakes:

    This is the third time that I have to ask you to remove the issue that was there for more than 20 hours. What is going on here?
and at the same time the compacting is firing like crazy. (What adds ~4 minute delays every 1 - 15 minutes)

  | # | Time     | Gap before | Session span | API calls |
  |---|----------|-----------|--------------|-----------|
  | 1 | 15:51:13 | 8s        | <1m          | 1         |
  | 2 | 15:54:35 | 48s       | 37m          | 51        |
  | 3 | 16:33:33 | 2s        | 19m          | 42        |
  | 4 | 16:53:44 | 1s        | 9m           | 30        |
  | 5 | 17:04:37 | 1s        | 17m          | 30        |
  # — sequential compaction event number, ordered by time.

  Time — timestamp of the first API call in the resumed session, i.e. when the new context (carrying the compaction summary) was first sent to the
  model.

  Gap before — time between the last API call of the prior session and the first call of this one. Includes any compaction processing time plus user
   think time between the two sessions.

  Session span — how long this compaction-resumed session ran, from its first API call to its last before the next compaction (or end of session).

  API calls — total number of API requests made during this resumed session. Each tool use, each reply, each intermediate step = one request.

Bottomline, I will probably stay on Sonnet until they fix all these issues.
ericol··on Show HN: Baton – A desktop app for developing with AI agents
This looks dangerously close to cmux but with a narrower focus (Just Claude code)

BTW, the claude app kind supports this with the /remote-control command, and that was what made me move away from cmux (I still have to start the sessions there)

ericol··on Inside the M4 Apple Neural Engine, Part 1: Reverse Engineering
It's not exactly that...
ericol··on The engine of Germany's wealth is blocking its future
The future is not evenly distributed.

I tried to search for it, but even the 2 documents that superseded the one from around the time my daughter was at school at not available.

I mean, the site doesn't even have a valid secure certificate so...

In the site below (In Spanish) you can search for 10/2019 and a cursory translation of the document title will show that this is the proper document (For 2019 onwards, the replaced doc 04/2014 isn't available either)

https://koha.chubut.edu.ar/cgi-bin/koha/opac-search.pl?idx=k...

ericol··on The engine of Germany's wealth is blocking its future
When my eldest daughter was in high school (~2010, Argentina) there was a provincial policy where if every single student had a result below a certain score in a test, the scores had to be re assessed against the maximum result.

The resulting situation here was that she was constantly bullied into underperforming. Both cases are actually similar in that each individual has a personal incentive to underperform - the difference is that in your friend's case the policy is granted at the company level so no single employee can defect and break it for the rest, while in my daughter's case one high scorer could invalidate the reassessment for everyone, which is exactly what made defection punishable and the bullying emerge naturally.

ericol··on Claude's Cycles [pdf]
Previous discussion: https://news.ycombinator.com/item?id=47230710
ericol··on Inside the M4 Apple Neural Engine, Part 1: Reverse Engineering
> human intuition driving the exploration

This, a thousand times this.

For me, what AI brings is augmented humans. Just as we don't calculate on paper anymore, what is the reason of doing things by hand when a machine in X times better.

Want to code by hand, as artisans of old? Suit yourself.

I, for one, love the smell of burning chrome.

ericol··on Switch to Claude without starting over
I regularly (Say, once a month) do a comparison of results across all Claude, Gemini and ChatGPT. Just for reasons, not that I want to see if there's any benefit in changing.

It's not "fair" in that I pay for Claude [1] and not for the others, so models availability is not complete except for Claude.

So I did like things at time in the form of how they were presented, I came to really like Sonnet's "voice" a lot over the others.

Take into account Opus doesn't have the same voice, and I don't like it as much.

[1] I pay for the lower tier of their Max offering.

ericol··on Tell HN: YC companies scrape GitHub activity, send spam emails to users
I've had more than a few instances of this over the past 2 years, and my reply is exactly the above.

"What you are doing is against Github's TOS"

ericol··on Writing code is cheap now
He definitely was paying attention.

He had to pause for a second there, arrested by the realization, and was one of the reasons I got an "Exceeds expectations" in one of my KRAs.

ericol··on Writing code is cheap now
> The long-term effect is less clear. If we generate more code, faster, does that reduce cost or just increase the surface area we need to maintain, test, secure, and reason about later?

My take is that the focus is mostly oriented towards code, but in my experience everything around code got cheaper too. In my particular case, I do coding, I do DevOps, I do second level support, I do data analysis. Every single task I have to do is now seriously augmented by AI.

In my last performance review, my manager was actually surprised when I told him that I am now more a manager of my own work than actually doing the work.

This also means my productivity is now probably around 2.5x what it was a couple of years ago.

ericol··on SkillsBench: Benchmarking how well agent skills work across diverse tasks
> Having the LLM write down a skill representing the lessons from the struggle you just had to get something done is more typical (I hope) and quite different from what they're referring to

Just as of last week I had Claude build me a skill when I ask it to help me troubleshoot issues, and it came out quite good.

It did had some issues (Claude tends to o er specify over anecdotal data) but it's a strong step in the right direction.

Also, "skills" are too broad in my opinion. I have one (that Claude wrote) with my personal data that I have available when I analyze my workouts.

I think there's ample room for self-generated skills when you use a rather long exchange on a domain you plan to revisit, _specially_ when it comes to telling Claude what not to do.

ericol··on Porting 100k lines from TypeScript to Rust using Claude Code in a month
I recently had to create a MySQL shim for upgrading a large PHP codebase that currently is running in version 5.6 (Don't ask)

The way I aimed at it (Yes, I know there are already existing shims, but I felt more comfortable vibe coding it than using something that might not cover all my use cases) was to:

1. Extract already existing test suit [1] from the original PHP extensions repo (All .phpt files)

2. Get Claude to iterate over the results of the tests while building the code

3. Extract my complete list of functions called and fill the gaps

3. Profit?

When I finally got to test the shim, the fact that it ran in the first run was rather emotional.

[1] My shim fails quite a lot of tests, but all of them are cosmetics (E.g., no warning for deprecation) rather than functional.

ericol··on The architecture of “not bad”: Decoding the Chinese source code of the void
This applies to many different things, depending on the pair of languages you are using.

In Spanish the closes approximation would be "ni mal ni bien" (Not bad not wrong) but I understand the Chinese expression has a strong lean on "not being wrong".

Not so long ago (I'm 50+, Spanish native speaker, and I've spoken English for the past 30 years almost daily) I learnt about "accountability".

Now before I get a barrage of WTFs, the situation is that in Spanish we only have "Responsabilidad" and that accounts for both responsibility and accountability, with a strong lean on responsibility.

So basically we recognise what is it to be responsible of something, but being accountable is seriously diluted.

The implications of this are enormous, and this particular though exercise I'd leave for people that spend more time thinking about these things than I do.

ericol··on IBM to acquire Confluent
> Read, huge IBM services implementation projects that the company bungled more often than not

Well this is _not_ what they wanted to sell in that talk.

But the implementation shown was über vanilla, and once I got home the documentation was close to un existent (Or, at least, not even trying to be what the docs for such a technology should be).

ericol··on IBM to acquire Confluent
> Watson was intended to solve fuzzy optimization problems.

> Unfortunately, the way it solved fuzzy was 'engineer the problem to fit Watson, then engineer the output to be usable.'

I'm going to review my understanding of fuzzy optimization because this last line doesn't fit the bill in it.

ericol··on IBM to acquire Confluent
I had no idea about what Watson was initially meant to solve.

I do remember they tried to sell it - at least in the meeting I went - as a general purpose chatbot.

I did try briefly to understand how to use it, but the documentation was horrendous (As in, "totally devoid of any technical information")

ericol··on IBM to acquire Confluent
> Ok, so does anyone remember 'Watson'? It was the chatgpt before chatgpt. they built it in house

I do. I remember going to a chat once where they wanted to get people on-board in using it. It was 90 minutes of hot air. They "showed" how Watson worked and how to implement things, and I think every single person in the room knew they were full of it. Imagine we were all engineers and there were no questions at the end.

Comparing Watson to LLMs is like comparing a rock to an AIM-9 Sidewinder.

ericol··on A triangle whose interior angles sum to zero
That's not a fucking triangle.

(It's Friday night people it's a joke and I have no idea what the article is talking about just looked at the picture)

ericol··on Gemini CLI tips and tricks for agentic coding
Even thought the doc _might_ be AI generated, that repo is Addy Osmani's.

Of Addy Osmani fame.

I seriously doubt he went to Gemini and told it "Give me a list of 30 identifiable issues when agentic coding, and tips to solve them".

ericol··on Human brains are preconfigured with instructions for understanding the world
So we come with a system prompt?
ericol··on It's insulting to read AI-generated blog posts
As a side note, i hate posts where they go on and on and use 3 pages to go to the point.

You know what I'm doing? I'm using AI to chase to the point and extract the relevant (For me) info.

ericol··on It's insulting to read AI-generated blog posts
> read something spit out by the equivalent of a lexical bingo machine because you were too lazy to write it yourself.

Ha! That's a very clever spot on insult. Most LLMs would probably be seriously offended by this would thy be rational beings.

> No, don't use it to fix your grammar, or for translations, or for whatever else you think you are incapable of doing. Make the mistake.

OK, you are pushing it buddy. My mandarin is not that good; as a matter of fact, I can handle no mandarin at all. Or french to that matter. But I'm certain a decent LLM can do that without me having to resort to reach out to another person, that might not be available or have enough time to deal with my shenanigans.

I agree that there are way too much AI slop being created and made public, but yet there are way too many cases where the use is fair and used for improving whatever the person is doing.

Yes, AI is being abused. No, I don't agree we should all go taliban against even fair use cases.

ericol··on Retiring Windows 10 and Microsoft's move towards a surveillance state
> It's that Microsoft pretends they are required

I think that's what "artificial limitations" mean. Microsoft pretending they are required when they are not.

← PreviousPage 2 of 13Next →