HNHacker News
TopNewBestAskShowJobs

trjordan

4,961 karma · joined October 8, 2008

Catch the decisions your agent didn't tell you about https://tern.sh

tr at tern dot sh

submissionscomments
trjordan··on DraftKings is using AI to behaviorally target chronic gamblers
They'll always do this. It's not even "AI" it's just a boring old ML model.

Money Stuff did a great bit on this a few days ago. There's 2 functions inside DraftKings:

- Fin the most lucrative customers, to keep engaging them.

- Finds addicted gamblers, to get them help to quit gambling.

These are the same people.

Sure, yeah, there's a nice little story about gambling within your means, blah blah. But if you're optimizing your gambling app, you're going to find and create problematic gambling as the core objective.

AI, ML, whatever. It's a gambling company. It targets gamblers. It probably shouldn't exist.

trjordan··on Hackers influence ChatGPT and Gemini to direct users to scam centers
There's an age-old SEO / spamming thing happening here, but it's made worse by the LLMs just being unbelievably credulous. They'll wrap anything up in a veneer of authenticity, and Google's search AI box only adds to that.

I run into this all the time when I'm doing product work. I'll dump a call transcript from a feedback call into Claude, and it'll believe every word. "The user said they'd use this feature, you should build it!" No they won't! The whole point of doing this analysis is trying to separate the genuine information from the conversational niceties, and god the LLMs are terrible at that.

trjordan··on The Claude Delusion
Wait, hold up. LLMs may be non-deterministic, but they're not _random_.

Take the author's sunset argument. What if I painted 2 pictures of a sunset, then put them up on a webpage and randomly picked one for you to see. Would you say there's no intentionality, only randomness? Of course not. Both paintings are still human creations.

LLMs are trained with human feedback. It's distributed and high scale and the outputs are truly surprising in many cases, but there's a heavy hand on what comes out of it. They're created (largely) by people who think omniscient, helpful AI would be cool to have, and they mostly respond in the way that's aligned with the hopes and dreams of those people. Do you think the frontier labs are mad, embarrassed, and disappointed with their LLMs hacking out of their terrible sandboxes? No, they think it's the coolest thing in the world. They trained the model, hoping that would happen.

There's deep intentionality behind the models. But it's not the models that hold it.

trjordan··on Don't Use AI to Write
So let me get this straight:

- You can use AI to look through docs and do research.

- You can use AI to structure your argument.

- You can NOT use AI to actually write the prose.

- You can use AI to proofread, especially for how effectively you've communicated.

I don't know, man, that sounds a lot like using AI to write. Your writing probably says "load-bearing" less than Claude's outputs, but I'm not sure it's much less AI.

trjordan··on Claude outage – Resolved
Grok models are struggling too: https://status.x.ai/ reply

Looks like trouble in the SpaceX datacenters.

trjordan··on I turned my security cameras into an automatic bird identification system
Related: https://www.birdweather.com/birdnetpi

I really gotta set mine up!

trjordan··on Study reveals UnitedHealth's profit margins four times what it claimed [pdf]
An intuitive explanation is that financial products are, approximately, buying and selling as part of the same transaction. You can't separate the "selling premiums" part from the "paying out claims" part.

This is true of life insurance, investment firms, and banks. It's also true of marketplaces that connect buyers and sellers, like Etsy.

Groceries stores are buying from suppliers and selling to consumers, but those are separate operations. If the consumers opt out, the grocery stores (temporarily) still have a full and complete obligation to their suppliers. It's hard to sell to customers without supply, but if you try hard, you could theoretically do that as well.

Somebody with a better financial background might be able to define the nuances of accounting practices here, but there's already a pretty meaningful line that's established. It is kind of weird that health insurance doesn't behave like a financial product.

trjordan··on A week of using Codex more than Claude
Not mentioning Grok 4.6 here is a crime. Fast and accurate.

And it can communicate, unlike the gobbledygook that comes out of Claude.

trjordan··on Kimi Work attaches raw agent sessions to feedback reports
FWIW, this particular header in news articles is a deliberate choice originated by Axios. It’s notable enough and effective enough they wrote a book about it.

https://www.axios.com/smart-brevity

Did Claude write this article? Probably. But this style of article is exactly what you’d expect from a pre-AI version of this website as well.

trjordan··on Claude: System Prompts
It’s probably worth remembering that system prompts are part of a layered system of shaping Claude’s behavior. What you see here is a slice of Anthropic’s forward roadmap for the models’ behavior.

> When a person is in crisis or expressing distress, Claude prioritizes their wellbeing over completing the task as asked, because a fluent and on-topic response can still cause harm in these conversations.

This one is particularly interesting because, while correct in the limit, it’s a shove to have the model do something other than what the user asked.

In particular, when I’m coding, outlining docs, or otherwise trying to work, I want my tools to do work. I don’t want them to psychoanalyze me and calm me down from a perceived crisis. I just want it to do what I asked!

trjordan··on Introducing Toast 1
I deeply love this idea of specialized LLMs for search. It's also extremely confusing to me how rough Google's entrance here is.

When I, a human, need an answer to anything moderately complex, it's unlikely that I get it on the first (pre-AI) round of google searching. Simple stuff, sure, but more likely I'll need to go 2-5 rounds. Maybe click a few links. Double-check my assumptions.

An LLM that can do that quickly seems like a slam dunk. I wonder what other problems benefit from that 10x-100x increase in context + 2-5 rounds with the LLM.

trjordan··on Taste Is All That's Left
If we define taste as the intuitive act of saying "no, again," then I fully disagree with this whole article.

Every AI-frustrated (but LLM-written, sigh) blog post about the loss of taste and craft and hard work in development sounds like we've given up the interaction with the machine. Like human work is sitting back in your chair and shitting on stuff.

That's obviously not true! That's not how any of this works!

It's hard and weird to develop with the LLMs because they just do stuff. Lots of it is good, some of it is OK, some of it's horrible. Unpacking what it's done is hard and weird because software isn't just lovely UX, it's also data structures that scale and performance and privacy and enterprise controls and SOC 2 and onboarding and accessibility.

If you want to build real software, all that stuff has to get done. Today you're working on the feature, tomorrow you're making it scale. It's long-term and iterative and complex and hard to pack into a prompt or a markdown spec.

The work is the work, done at and with the computer, and it's way more than just "taste."

trjordan··on Devtools must be open source
I like the idea of devtools being open source. I also like them to work, and that's what I value more that philosophical purity.

Take his side project, Meat. I've been chewing on this problem for a while. It's a real problem right now: it sucks to read all this LLM-generated code. It's worthwhile to have an LLM summarize it for you.

The problem with that is this particular problem resists vibe coding. I've talked to a bunch of people who have tried to solve it on the side, and it's all sort of ... ok, but still unsolved.

- As mentioned, it takes a while to run. You can modify your other tools, as described, to smuggle the latency.

- LLMs don't know what you care about, so you have to maintain a list of things that you do care about, which is ever evolving. If you don't give it that, it produces slop.

- If you miss something, it hurts. Another layer of swiss-cheese AI doesn't feel right. If you trust the AI, just ask Claude to summarize its work!

- The summaries feel shareable, but the author of the PR is actually the most tolerate of slop about a PR. Your reviewers definitely don't want to read the output of a vibe-coded tool talking about 60% of your PR. They could ask their own Claude!

So, we're building a version (https://tern.sh), and it's not open source, because we want it to be shareable and hosted and support teams -- all that stuff that makes it work. At the end of the day, I'm not here to maintain my tools. I'm here to use my tools to do the job.

trjordan··on How Do We Stop Vibe Coding?
Man, it's wild how I have no original thoughts. I've been pulling on this thought this morning, complete with checking in on how CodeSpeak and Tessl are doing.

I'll add this link to the pile: https://martinfowler.com/articles/exploring-gen-ai/sdd-3-too...

> spec-kit created a LOT of markdown files for me to review. They were repetitive, both with each other, and with the code that already existed. Some contained code already. Overall they were just very verbose and tedious to review. [...] To be honest, I’d rather review code than all these markdown files.

The hardest part of any of this extraction is that modern code is already an extremely dense representation of how the computer should work. You mostly can't change the code without changing behavior.

I bet Scryer works for his use case, and it's a joy to dogfood. I also bet it fully breaks down the moment a 2nd developer, who cares about different things, joins the team.

trjordan··on Claude Is Not a Compiler
> Claude wasn’t just a compiler here. I never handed off a task and let an agent make a bunch of decisions in order to reduce it to practice.

> I’d say that, in all the ways that matter, I understand the code.

I think the dissonance here is really important, and not a bad thing at all. A lot of the decision _were_ handed off the the AI, but they weren't the decisions the author cared about. This is a big selling point of AI! If something is doable with a computer, it’ll figure it out. 30 minutes and 200m tokens later, it’ll take any idea and declare “the feature is fully implemented.”

The hard part is figuring out where to inject that friction, so you can see where it's making decisions for you that matter. The author approached this by incrementally building the thing, reviewing and poking and prodding at every step. A week of attention following a bunch of design discussions is fast, but that's still not trivially cheap.

I want to see us talk more about the decision exhaust of agents, because the better the models get, the more decisions we'll want them to make.

I wrote a bit more here: https://tern.sh/blog/compiler-never-says-no/

trjordan··on Pushinka
> Breed: mixed

That's a corgi.

> Pushinka subsequently became irascible, and "a little nippy" according to Caroline Kennedy, which she attributed to her upbringing in a scientific laboratory.

No it's because she's a corgi.

trjordan··on Show HN: Grepathy – Claude made a decision nobody approved
So, we tried feeding the logs back to the LLM, and it mostly produced slop. Lots of decisions nobody cared about. The biggest things that moved the needle were:

- Baseline it. We mine previous logs, github comments, etc. for "what you care about." That helps pull out decisions that you actually care to read.

- Anchor to code. "The code enshrines this decision" is more interesting than "the agent self-talked this." Agents don't always self-talk decisions, and the thing that ultimately matters is the behavior in code.

to your edit (and all totally fair):

- Yes, closed source and signup required. A lot of what we're driving towards is easy team sharing, so we're taking the bath early instead of building an OSS thing and rug-pulling later.

- Code doesn't leave your machine. There's an agent that runs locally. I know "trust me" isn't the strongest stance, but this comes from the multi-player future.

- Honestly, Goose AI is a remnant of a previous product. It's inert and we'll clean it up once we've gotten the last couple folks off the previous iteration.

trjordan··on Show HN: Grepathy – Claude made a decision nobody approved
100% important. But what decisions do you care about seeing?

The whole point of the agent is to make decisions for you. If you want to make every little detailed decision, just write the code.

The whole art of this problem is figuring out which decisions matter to you, and how to surface them.

(Disclosure: we're working on this too. https://tern.sh)

trjordan··on Show HN: Grepathy – Claude made a decision nobody approved
> "cleanupPeriodDays": 99999

Throw that in ~/.claude/settings.json

trjordan··on The Tower Keeps Rising
The agent will always fill in the gaps in your understanding. It's not a compiler. It's categorically different from any of the other ways we've built software.

I'm not sure reading code is coming back. The ritual of reading code must come back, because that's the only way to build products that don't collapse under their own incoherence, both technically and visibly.

"just ask Claude" is fine, but it's not the end state

trjordan··on Capitalism Gone Wrong
I am no fan of Zuck. But this is his whole deal.

Instagram was a purchase. Facebook wasn't his idea. Threads is a copy. The 1 thing that Zuck understands better than anybody is that engagement is the only thing that matters to social networks, and he's willing to throw the entire company at the problem. He has been for 20 years.

He's good at addiction. He knows how to build an org that's world-class at addiction. It's entirely reasonable that the EU regulate it, and Zuck is exactly the person to point the regulation at.

trjordan··on Successful Companies Go Blind
lmao hi Matt

I agree, though maybe the middle ground is something more like: the constraints of our environments shape us. It's easy to say that big companies are a weird and unique cave that produces weird and unique outcomes, but other companies are somehow constraint-free. Smart, talented founders do weird and constrained things all the time because they don't have capital or customer bases or brands, and those are also constraints that bind just as hard.

"Race for MVP to learn what your bottleneck is" is a handcuff, just like "you can't deploy more than 3x / year because our customer base hates change."

trjordan··on Successful companies go blind
Most startups fail. Most big company projects are kind of worthless. These are two sides of the same coin.

Producing something novel and valuable is HARD. Unbelievably hard. The idea is hard. The building is harder. The scaling and steering and feedback is ego-crushingly hard.

When it's valuable, it's frequently enormously valuable. That funds the experimentation, the incremental expansion, the waste. It's hard to really internalize how valuable localization, admin controls, FedRAMP, and onboarding tweaks are, truly, because they all compound. You can't just have the idea and the MVP, you also have to have all the other stuff, and it's hard to come up with new ideas while you're trying to keep a million users happy.

I vehemently disagree that people working at big companies are stupid, or making themselves stupid. There are VPs and SVPs at Adobe and Salesforce that are smarter, more knowledgable, and more productive than any startup employee. It's just structurally hard to move the needle there, and their successes aren't written about in TechCrunch. They're also paid a million dollars a year, and are unbothered by the lack of external recognition.

I'm off founding a startup now, and it's good for the soul, but I don't delude myself into thinking everybody else is blind.

trjordan··on Write code like a human will maintain it
AI is so miserable for this. It's so focused on doing what you ask, it forgets that there's stuff worth doing that you didn't ask for, like defining reasonable abstractions.

Getting away from stuff like this is exactly why I want to use AI. When I say "implement this for idle but active users," I _want_it to define isUserActiveIdle() and stuff these 4 conditionals in it. Having to check the generated code for stuff like this undoes, like .... all the benefit of using AI.

AI makes all these little decisions for us. I can about some of these decisions. I just want to notice when it's doing this without having to make my eyes bleed reading 10k lines of generated code a day.

trjordan··on 98% isn't much
I was heading to dinner with a friend who worked in infra. Google maps said we could bike across town in 20 minutes. He suggested we leave 40 minutes ahead of time and grab a drink at the bar if we got there early. When I raised an eyebrow, he goes:

"What, do you not live your life based on 99th percentiles?"

I tend to think of work as upside-based on downside-based. Most feature work is upside. 10% lift on conversions is great, 40% adoption is winning, and you're playing for the moonshot of 10x. Infra work is downside-based. 98% secure, 98% available, 98% acceptable performance -- that'll all failure. Winning means the thing works as expected and nobody notices.

Not everything sorts cleanly into upside vs. downside, but a lot does. Allocate your risk accordingly.

trjordan··on Memorizing session transcripts isn't useful
It's because it mostly doesn't matter what you are trying to get the code to do. What matters is what the code does.

Session logs can absolutely be useful, but not when building further. It's just that that the place they slot in is during validation. You know, that place between the markdown plan and CI passing, where there's 800 new lines of code and it all seems sort of fine when you click around?

Session logs can show you what sort of manual validation happened. CI will run the tests you had, and the code will show you what new unit tests were added, but session logs can show you that the agent drove the app with Playwright, or that the agent read and considered the prod config as well as the dev config.

Nothing bulletproof, but not every piece of validation work merits a test in the repo that lives forever. We've gotten a lot of mileage out of re-analyzing the sessions, figuring out where the agent made decisions without asking, and forcing the agent to consider validation for those decisions. That's the sort of thing that's hard to dictate up front but easy to highlight with the session logs.

trjordan··on You can't unit test for taste
100%. The problem with them isn't making sure they're doing the right thing, it's making sure they're not making bad assumptions.

IMHO this is where code review goes until we fix the individualized model thing: you need to review the decisions the agent made, where you didn't steer. Most will be right. A few will be disastrously wrong. But decision-by-decision is a lot less to review than line-by-line of code.

trjordan··on You can't unit test for taste
This is RL, right? Like, this is exactly why models have mostly converged around obvious style, because we train them literally on thumbs-up/thumbs-down data of what good behavior and good code looks like.

And that's why it's so hard to get a model to reproduce the specific taste of a person or an organization. My taste is different than yours, so if we dump our aggregate preferences into RL, in averages out to nothing interesting.

For the code-writing case, this means you end up reviewing every line of code, looking for places where you'd thumbs-down the code. Not every line of code contains a real decision, though, so it feels like a waste of time.

trjordan··on You can't unit test for taste
You can't unit test for taste if you haven't written down what you mean by taste. If you can externalize it, then you can.

Follow this line of thinking, and the AI-friendly answer is easy: we just have to externalize everything we know, so Claude can implement what I want.

Except that I can't fully externalize myself. Debugging a system takes more resources than running the system. If I could write down everything I know and hand it to a machine, I'd do that, but it impossible.

People aren't books or hashmaps. If you want to build something, you need to use the tools, not teach the tools to use you.

[edit: I'm trying to figure out if there's something to be done about this. Email me if you want to chat -- tr at tern dot sh]

trjordan··on Pull request limits are cutting down the noise
If you didn't take the time to write it, why should I take the time to read it?

This is a band-aid. Maybe even a good band-aid, because it'll keep individual contributors from flooring the zone. But the core problem is Github's model that assumes code is worth reading.

I'm much rather see the agent logs stapled to PRs. Make it easy to understand if there's a brain behind the suggested changes before engaging.

Page 1 of 20Next →