HNHacker News
TopNewBestAskShowJobs

Calavar

2,763 karma · joined August 23, 2021

submissionscomments
Calavar··on Spinel: Ruby AOT Native Compiler
I'm skeptical of that reasoning because the original C wasn't too clean or performant either. For example emit.c from an earlier commit [1]

It writes a separate call to emit_raw for each line, even though there many successive calls to emit_raw before it runs into any branching or other dynamic logic. What if you change this

    emit_raw(ctx, "#include <stdio.h>\n");
    emit_raw(ctx, "#include <stdlib.h>\n");
    emit_raw(ctx, "#include <string.h>\n");
    emit_raw(ctx, "#include <math.h>\n");
    // And on for dozens more lines
to this

    emit_raw(ctx,
        "#include <stdio.h>\n"
        "#include <stdlib.h>\n"
        "#include <string.h>\n"
        "#include <math.h>\n"
        // And on for dozens more lines
    );
That would leave you with code that is just as readable, but only calls the emit function once, leading to a smaller and faster binary. Again, this is a trivial change to the code, but Claude struggles to get there.

[1] https://github.com/matz/spinel/blob/aba17d8266d72fae3555ec91...

Calavar··on US special forces soldier arrested after allegedly winning $400k on Maduro raid
I know this is tangential to your overall point, but did really they murder everyone in the room? I was under the impression that a few Venezuelan generals kidnapped Maduro themselves, left him at a predetermined point for US forces to pick up, and had their soldiers fire some small arms into the air to make a token show of resistance. There's no way the US would have flown a slow-moving convoy of helicopters into a hostile city unless they knew a priori that Venezuelan air defense missile batteries would be ordered to stand down.
Calavar··on Spinel: Ruby AOT Native Compiler
spinel_codegen.rb is an eldritch horror. I always get spaghetti code like this when using Claude, and I've been wondering if I'm doing something wrong. Now I see an application that looks genuinely interesting (not trivial slop) written by someone I consider to be a top notch programmer, and the code quality is still pretty garbage in some places.

For example infer_comparison_type() [1]. This is far from the worst offender - it's not that hard to read - but what's striking here that there is a better implementation that's so simple and obvious and Claude still fails to get there. Why not replace this with

    COMPARISON_TYPES = Set.new(["<", ">", "<=", ">=", "==", "!=", "!"])

    def infer_comparison_type(mname)
      if COMPARISON_TYPES.include?(mname)
          "bool"
      else 
        ""
      end
      # Or even better, strip the else case
      # (Which would return nil for anything not in the set)
    end
This would be shorter, faster, more readable, and more easily maintainable, but Claude always defaults to an if-return, if-return, if-return pattern. (Even if-else seems to be somewhat alien to Claude.) My own Claude codebases are full of that if-return crap, and now I know I'm not alone.

Other files have much better code quality though. For example, most of the lib directory, which seems to correspond to the ext directory in the mainline Ruby repo. The API is clearly inspired by MRI ruby, even though the implementation differs substantially. I would guess that Matz prompted Claude to mirror parts of the original API and this had a bit of a regularizing effect on the output.

[1] https://github.com/matz/spinel/blob/98d1179670e4d6486bbd1547...

Calavar··on Over-editing refers to a model modifying code beyond what is necessary
What sorts of instructions?
Calavar··on Over-editing refers to a model modifying code beyond what is necessary
It's interesting how variable people's experiences seem to be.

Personally, I tend to get crap quality code out of Claude. Very branchy. Very un-DRY. Consistently fails to understand the conventions of my codebase (e.g. keeps hallucinating that my arena allocator zero initializes memory - it does not). And sometimes after a context compaction it goes haywire and starts creating new regressions everywhere. And while you can prompt to fix these things, it can take an entire afternoon of whack-a-mole prompting to fix the fallout of one bad initial run. I've also tried dumping lessons into a project specific skill file, which sometimes helps, but also sometimes hurts - the skill file can turn into a footgun if it gets out of sync with an evolving codebase.

In terms of limits, I usually find myself hitting the rate limit after two or three requests. On bad days, only one. This has made Claude borderline unusable over the past couple weeks, so I've started hand coding again and using Claude as a code search and debugging tool rather than a code generator.

Calavar··on The abandoned war: Why no one is stopping the genocide in Sudan
It's really hard to cry victim about others misrepresenting Trump's motives for the Iran war as oil, oil, oil when the US did in fact launch a military attack on a country - within the last six months - where the subsequent negotiated agreement on oil rights was quite literally described by the White House press secretary as "the president’s control of Venezuela’s oil" [1] and just a few weeks later the president held a public, televised conference with Chevron and ExxonMobil executives in the White House where he pitched them on investing in the Venezuelan oil industry [2]

[1] https://www.wsj.com/business/energy-oil/trump-venezuela-oil-...

[2] https://youtu.be/sD4x6T-u4XY

Calavar··on The creative software industry has declared war on Adobe
> we're seeing an explosion of brand-new high-polish OSS apps this year

Do you mind sharing a few examples?

Calavar··on Claude Code Found a Linux Vulnerability Hidden for 23 Years
It's not an insightful statement right now, but it was at the peak of cloud hype ca. 2010, when "the cloud" often used in a metaphorical sense. You'd hear things like "it's scalable because it's in the cloud" or "our clients want a cloud based solution." Replacing "the cloud" in those sorts of claims with "another person's computer" showed just how inane those claims were.
Calavar··on MoD sources warn Palantir role at heart of government is threat to UK security
I still don't understand why Theil and Karp decided to name their surveillance tech company after a device that is best known for being used by an evil dark lord to decieve and corrupt. It's like the Mitchell and Webb skit "are we the baddies" except they're the ones who designed the uniforms with skulls on them.
Calavar··on Entities enabling scientific fraud at scale (2025)
This is a good point. It is not humanly possible to verify every claim you read from every source.

Ideally, you should independently verify claims that appear to be particularly consequential or particularly questionable on the surface. But at some point you have to rely on heuristics like chain of trust (it was peer reviewed, it was published in a reputable textbook), or you will never make forward progress on anything.

Calavar··on Judge orders government to begin refunding more than $130B in tariffs
> If they couldn't do anything that gave an appearance of a conflict

This time I won't say maybe - that's a straw man.

I never said Cantor shouldn't be able to do anything that even gives the appearance of a conflict. Or anything even close to that really.

As you said yourself further up the thread, investments of investment bank employees are highly regulated. And not only employees themselves, but also their immediate family members.

Yet that same level of legal regulation doesn't apply to immediate relatives of government officials. We've seen frequently with spouses and children of congressmen, and now we're seeing it with the son of a cabinet member. Yes, this may technically be legal, but legal does not equate to just and desirable. This reads to me like a serious loophole in the law that needs to be closed.

Calavar··on Judge orders government to begin refunding more than $130B in tariffs
> In this case, the idea that Cantor can't do something because the former head is now in a government job is crazy. No one "in the business" thinks Cantor is suddenly hobbled.

That's not the idea, and it almost seems like a straw man to be honest. The actual idea is that the current head of Cantor can't do something because he's a direct relative of a high ranking government official whose powers and job duties present a conflict of interest for this specific set of transactions.

Calavar··on Why XML tags are so fundamental to Claude
It looks like a screenshot from the Claude desktop app, so I don't think the author is trying to disguise the AI origin of the marerial
Calavar··on ChatGPT Health fails to recognise medical emergencies – study
I know you qualified your assertion of three patients an hour with general practice, but there are plenty of specialty practices where six patients an hour is common. Dermatology and ophthalmology clinics often run at that pace (at least in the US). Some surgical clinics can run at that pace for follow up visits (not for initial visits)
Calavar··on Overall, the colorectal cancer story is encouraging
The average American says US healthcare spending, which is 3x to 20x that of other OECD countries on a per capita basis, is way too high.

The average American also thinks they should be provided testing and procedures that their insurance deems medically unnecessary.

Try to reconcile these two beliefs. (Hint: It's impossible)

Calavar··on Lena by qntm (2021)
I agree that not all nudity is porn - nudity is porn if the primary intent of that nudity is sexual gratification. When the nudity in question was a Playboy magazine centerfold, the primary intent is fairly obvious.
Calavar··on Software factories and the agentic moment
It could once upon a time: https://www.bellard.org/tcc/tccboot.html

It can't do that today though. Linux uses C11 features and also many GCC extensions that tcc doesn't implement.

Calavar··on LLMs as the new high level language
I'm not sure I'd call agents an army of juniors. More like a high school summer intern who has infinite time to do deep dives into StackOverflow but doesn't have nearly enough programming experience yet to have developed a "taste" for good code

In my experience, agentic LLMs tend to write code that is very branchy with cyclomatic complexity. They don't follow DRY principles unless you push them very hard in that direction (and even then not always), and sometimes they do things that just fly in the face of common sense. Example of that last part: I was writing some Ruby tests with Opus 4.6 yesterday, and I got dozens of tests that amounted to this:

   x = X.new
   assert x.kind_of?(X)
This is of course an entirely meaningless check. But if you aren't reading the tests and you just run the test job and see hundreds of green check marks and dozens of classes covered, it could give you a false sense of security
Calavar··on Coding agents have replaced every framework I used
Claude already knows who the characters Frodo, Sam, and Gollum are, what their respective character traits are, and how they interacted with each other. This isn't the same as writing something new.
Calavar··on We tasked Opus 4.6 using agent teams to build a C Compiler
I think that's overstating it.

JustHTML is a success in large part because it's a problem that can be solved with 4 digit LOC. The whole codebase can sit in an LLM's context at once. Do LLMs scale beyond that?

I would classify both FastRender and Opus C compiler as interesting failures. They are interesting because they got a non-negligible fraction of the way to feature complete. They are failures because they ended with no clear path for moving the needle forward to 80% feature complete, let alone 100%.

From the original article:

> The resulting compiler has nearly reached the limits of Opus’s abilities. I tried (hard!) to fix several of the above limitations but wasn’t fully successful. New features and bugfixes frequently broke existing functionality.

From the experiments we've seen so far it seems that a large enough agentic code base will inevitably collapse under its own weight.

Calavar··on We tasked Opus 4.6 using agent teams to build a C Compiler
Sure, maybe it's tricky to coerce an LLM into spitting out a near verbatim copy of prior data, but that's orthoginal to whether or not the data to create a near verbatim copy exists in the model weights.
Calavar··on We tasked Opus 4.6 using agent teams to build a C Compiler
The classical definition of a clean room implementation is something that's made by looking at the output of a prior implementation but not at the source.

I agree that having a reference compiler available is a huge caveat though. Even if we completely put training data leakage aside, they're developing against a programmatic checker for a spec that's already had millions of man hours put into it. This is an optimal scenario for agentic coding, but the vast majority of problems that people will want to tackle with agentic coding are not going to look like that.

Calavar··on Y Combinator will let founders receive funds in stablecoins
Sure, you can swap a USD tethered stablecoin to another currency. You can also do that with USD itself. What is the value add of injecting an extra step by first converting your USD to stablecoin?
Calavar··on My Mom and Dr. DeepSeek (2025)
As a physician, I can give further insight. The blood pressure medication the commenter is referring to is almost certainly a beta blocker. The effect on blood sugar levels is generally modest [1]. (It is rare to advise someone with diabetes to stop taking beta blockers, as opposed to say emphysema, where it is common)

They can be used for isolated, treatment of high blood pressure, but they are also used for dual treatment of blood pressure and various heart issues (heart failure, stable angina, arrhythmias). If you have heart failure, beta blockers can reduce your relative annual mortality risk by about 25%.

I would not trust an LLM to weigh the pros and cons appropriately knowing their syncophantic tendencies. I suspect they are going to be biased toward agreeing with whatever concerns the user initially expresses to them.

[1]

Calavar··on AGENTS.md outperforms skills in our agent evals
I thought most of the major AI programming tools were already doing this. Isn't this what subagents are in Claude code?
Calavar··on Tesla ending Models S and X production
Talk about market cap, especially meme stock maket cap, reminds me of that old XKCD comic on extrapolation. Market cap is what you get when you extrapolate the fair market value for the 1% of a company's shares currently on the market all the way out to 100%. But demand doesn't work that way - it doesn't scale linearly.
Calavar··on Iran Protest Death Toll Could Top 30k, According to Local Health Officials
Reza Pahlavi has also had recent interviews with CBS [1], the Economist [2], and CNN [3] (all within the last 30 days). So how is the existence of Reza Pahlavi interviews on Fox evidence that conservative media is covering this issue more than liberal media?

On the topic of Politicians, Democratic congressmen like Dave Min and Jim Hines have also spoken in favor of US intervention in Iran.

[1] https://www.cbsnews.com/video/full-interview-exiled-iranian-...

[2] https://youtu.be/jk6wfvje8Zo

[3] https://www.cnn.com/2026/01/23/world/video/iran-united-state...

Calavar··on Iran Protest Death Toll Could Top 30k, According to Local Health Officials
This is a straw man in my opinion. But regardless of that, your theory doesn't explain why conservative media isn't really covering this either - The Iran protests haven't exactly been front page material on Fox News or OAN or Breitbart
Calavar··on A flawed paper in management science has been cited more than 6k times
Isn't a lack of comments the opposite of the problem you were previously claiming?
Calavar··on A flawed paper in management science has been cited more than 6k times
> All comments, complaints, corrections, and retraction requests? Unmoderated? Einstein articles will be full of comments explaining why he is wrong, from racist to people that can spell Minkowski to save their lives. In /newest there is like one post per week from someone that discover a new physics theory with the help of ChatGPT. Sometimes it's the same guy, sometimes it's a new one.

Judging from PubPeer, which allows people to post all of the above anonymously and with minimal moderation, this is not an issue in practice.

← PreviousPage 2 of 17Next →