HNHacker News
TopNewBestAskShowJobs

cadamsdotcom

2,723 karma · joined December 1, 2014

github.com/cadamsdotcom linkedin.com/in/cadamsdotcom chris at cadams dot com
submissionscomments
cadamsdotcom··on Clean up Claude 5's token vomit with a separate LLM
The better approach is to stop the LLM in its tracks the moment it emits jargon or tortured metaphor and inject a turn that tells it what's expected instead.

The joy of watching a dumb AI-ism be sharply corrected by code you wrote months ago is hard to explain.

cadamsdotcom··on Air Theremin – A browser theremin you play by waving at your webcam
Well then be careful of US/UK spelling!
cadamsdotcom··on Devices with GrapheneOS support should be available in 2027
Stoked to soon be able to give these people money. The moment there's a device available I'll be checking closely to see if it meets my requirements.

Thanks to everyone who worked so hard to get here.

cadamsdotcom··on Opus 5.0 drives incoherence into the stratosphere
I'm voting with my wallet.

My plan downgrade kicks in next month.

Thanks Opus 5 for helping me kick the habit.

cadamsdotcom··on AI usage patterns in software teams
Nice data!

Tim (author) if you're there: it'd be amazing to see the split of which agents people are using, if you have that data.

cadamsdotcom··on Show HN: Automatically detect and patch walking-dead states in Sierra games
Why not rip and replace the parts of the readme then?

It's like making a beautiful chocolate praline and shipping it in a shoe.

cadamsdotcom··on Composable Tests
Why wouldn't you use parameterized tests or compose your setup out of fixtures that run the shared setup code?

Running a test, then running it a second time, then asserting something else.. is just weird.

cadamsdotcom··on Data centers raise nearby temperatures by up to 4 degrees in Phoenix
Can we please state the measuring system in use.

4F is not enough to rise above a noise floor; 4C is getting close.

cadamsdotcom··on Pacing model development in an era of cyber-critical capabilities
What a breath of fresh air.

If 2026's Anthropic did an announcement like that, it'd be so many words it'd crash the browser.

cadamsdotcom··on llms.txt: a proposed standard no major AI platform has confirmed it uses
So an LLM can't just read your home page?

"we seem to keep shooting ourselves in the foot, and the solution we have invented is bulletproof shoes"

cadamsdotcom··on Beware Management Consultants
They got a D in "willingness to pay for a certificate for their shopfront"..
cadamsdotcom··on Superpowers, Not Superintelligence
Your last paragraph is a valid rebuttal. And I've done what you did in the first 3.

But your first 3 paragraphs serve no one but you.

My suggestion for your future, if you'll consider it, is to write those first 3 paragraphs (they feel good to write!), but next, take a deep breath and delete them.

I agree with your point by the way - everything has pros and cons - and there're tipping points. "Too much of a good thing" is true of so much in life.

cadamsdotcom··on Google has acquired the data of failed US airline Spirit
> they promise "by a certain time"

Better call your AI company for a refund. That my friend was an hallucination and you've been promised no such thing.

cadamsdotcom··on Apple announces changes for apps in the European Union
The $99/yr part gives me the ick.
cadamsdotcom··on The Benchmarkpocalypse
Software is a very "spiky domain"; things either work or fail, and there is sharp delineation between and easy verification.

Hm. Two orthogonal properties! This sounds like a 2x2 matrix!

Let's swap hard/easy around & explore the 4 possibilities...

There are domains with sharp delineation and hard verification; they are not at risk until AI gets much better. Humans operate in these domains by applying tremendous deep thought and subjective judgment - our superpower.

Domains with soft delineation and easy verification are most at risk: "it's a picture of a cat" remains true through a wide range of perturbations - eg. skewing the image or moving it across a pixel or correcting its white balance or even changing the cat. AI music? Lots of domains already solved by AI here but they're also not that meaty.

My prediction is the next interesting stuff will happen where verification is hard but there's no sharp delineation. It's the world of "I'll know it when I see it". Good customer service?

cadamsdotcom··on How I over-engineered my book
Custom linting is so cheap and easy, why not!

> "Please add in this repo, a pre-commit hook in the form of a script that prints all filenames and line numbers on which the hyphenated "open-source" string exists, and if any instances are found, will exit nonzero. Then please install it in `.git/hooks` so commits are blocked until me or an agent removes all the instances, so we never again commit the string "open-source"."

1 minute after hitting enter you've created yourself a guarantee your repo will never again see the string "open-source".

cadamsdotcom··on The Benchmarkpocalypse
> expensive, in terms of tokens.

No amount of tokens can come close to my hourly rate.

cadamsdotcom··on The Benchmarkpocalypse
Ungrounded LLM outputs are a bit like your dreams. Without anything to test hypotheses against, stuff can pop in and out of existence and physics is just advice.

Ground your LLM. Tests, documentation, give it many ways to run the thing its reasoning about. It needs to be able to test its hypotheses on its own.

Take yourself out of that loop so you only find out once it's sure.

cadamsdotcom··on Israel creates fake think tank in likely attempt to dupe AI chatbots
This doesn't align with my experience - but happy to be proven wrong if evidence can be produced.

Do you have evidence?

cadamsdotcom··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
It's situational and I suspect there are situations where it would and others where it wouldn't. Would depend on the almost infinite variables of how training was done. You'd be right that it'd be likely to switch but while it's possible it's due to temperature, there are just so many things going on. But it would be one sensible explanation among many.
cadamsdotcom··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
Point taken & let's get back on track.

I agree it's incomplete what I wrote; you were right to call it out.

When you as an individual are getting tens of thousands of dollars of value from a $100 item and someone says there's a slightly-worse version for $1 but can't tell you if you'll find it slightly worse or noticeably worse.. you have to stop what you're doing to evaluate them, pay switching costs etc etc. so the cost difference is less than $99 and might actually be negative, and the difference in margin after paying switching costs is negligible.

So your rational response is to investigate to lower the stakes - but even performing the investigation eats into the time you're spending getting your tens of thousands of dollars of value.

cadamsdotcom··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
That could be how it works, but in practice it takes into account all previous tokens when producing the next-token distribution to sample from. So a switch back is more likely than your explanation supposes.
cadamsdotcom··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
My argument may be wrong, but I myself am just fine, thank you.
cadamsdotcom··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
It's less fungible than you think.

Below a certain threshold prices are effectively all the same.

Companies pay API prices because part of the bundle is a trust anchor / liability shield: "we bought Anthropic's thing! We didn't risk it! We paid for ZDR!"

cadamsdotcom··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
Your comparison is wrong

It's not an expensive car vs. any car.

It's car vs. no car.

cadamsdotcom··on GPS and the Lost Art of Getting Lost
Let's call it the lost art of getting lost when you don't want to.

It's as easy as ever to lose yourself or get lost if you so choose. Probably always will be.

cadamsdotcom··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
200usd/mo for Claude gives me tens of thousands of dollars of value.

API prices are paid by companies getting tens of millions of dollars of value.

In normal life money is the key constraint; buy this don't buy that etc. - whereas in VC funded companies the constraint is time. If you as a founder get funding and don't spend it fast enough you put yourself at serious risk of being replaced.

When enough of the world operates on that principle it creates a highly price insensitive market and that then can support a ton of ideas and experiments, some of which turn out to be really really good. It's a wild way to do innovation but it's been working well for decades.

cadamsdotcom··on What happens when an LLM never sees material beyond fifth grade?
You should not need it to say no.

You can get just as good information by asking its thoughts for and against some issue.

That doesn't force it to stop being sycophantic; in fact it actually exploits sycophancy to give you what you want.

cadamsdotcom··on Ask HN: How do you keep up with HN these days?
You might like a little RSS proxy I vibed purely to let me browse HN and Lobsters' RSS feeds offline. It reads in the articles and comments and serves them up in a new feed you subscribe to instead of the original.

RSS is amazing and being able to zip through your feed without waiting for pages to load is a game changer on top.

https://github.com/cadamsdotcom/offlinerss

cadamsdotcom··on Stop sending me huge PRs; a rant
This to me reads less as there being some objective level when a PR becomes "huge" and more about the tension between a system's ability to absorb change vs. our tools' ability to create change.

By way of analogy consider the relative impact on an ecosystem of one person fishing with a fishing line vs. a commercial fishing boat trawling the ocean. Of course, one person fishing is unlikely to have a huge impact on the ocean so it's generally permitted. Trawling (agentic coding) can be done in a way that's destructive to ecosystems but it can also be done sustainably!

So with the analogy in mind let's bring back the "sensible trawling" idea to agentic coding. What might it look like to solve the "huge PR bad" constraint in another way: by increasing our codebases' ability to absorb change, so what "a huge PR" is, becomes bigger?

Probably needs solves at many levels: assistance quickly comprehending the PR (AI driven walkthroughs, multiple media expected from the PR submitter not just text - eg. a screencast walkthrough of it), it requires rethinking how the code is read (better review tooling); it requires integrations with code-review automation tools (both you home-grown checklist and third-party tools) it requires rigorous testing (comprehensive automated e2e; test-driven; functional tests; etc); it requires putting the actual "in the loop" so post-release fast-follows can be expedited (eg. product signals and Sentry and metric anomalies are fed back in for quick follow up releases)

If you can be so much more responsive to the customer and market. Eg. you can unlaunch features just as easily as you launched them - and you can finally clean up all that tech debt. Better for the business better for the codebase and better for developer happiness.

Not all these strategies work for every situation - you can't do post-release in the loop if the shit needs to work first time! But the whole idea creates so much richness in applying human judgment and engineering solutions and it's all brand new because we never needed to deal with this much change before

Think of it as "releases in the loop".

If opportunities to rethink the stack to support MORE change excite you, congratulations! You're ready for the future that's coming. If you don't like this - get yourself into a job where you can say no a lot, or where shit needs to work first time, and you can be happy. Test-driven, strongly reviewed.. there's ways with agentic coding to also make super high quality stuff. But you can also shoot product from the hip more accurately and more often than ever before.

It won't be applied correctly everywhere - it's still heavily judgmental dependent and we're all fallible - but there'll be a much wider spectrum of options for how to build products. I think this is a really exciting future!

← PreviousPage 2 of 34Next →