HNHacker News
TopNewBestAskShowJobs

jumploops

3,235 karma · joined March 1, 2019

username @ gmail
submissionscomments
jumploops··on If AI writes code, should the session be part of the commit?
This is great, giving agents access to logs (dev or prod) tightens the debug flow substantially.

With that said, I often find myself leaning on the debug flow for non-errors e.g. UI/UX regressions that the models are still bad at visualizing.

As an example, I added a "SlopGoo" component to a side project, which uses an animated SVG to produce a "goo" like effect. Ended up going through 8 debug docs[0] until I was satisified.

[0]https://github.com/jumploops/slop.haus/tree/main/debug

jumploops··on If AI writes code, should the session be part of the commit?
I do something similar, but across three doc types: design, plan, and debug

Design works similar to your project.md file, but on a per feature request. I also explicitly ask it to outline open questions/unknowns.

Once the design doc (i.e. design/[feature].md) has been sufficiently iterated on, we move to the plan doc(s).

The plan docs are structured like `plan/[feature]/phase-N-[description].md`

From here, the agent iterates until the plan is "done" only stopping if it encounters some build/install/run limitation.

At this point, I either jump back to new design/plan files, or dive into the debug flow. Similar to the plan prompting, debug is instructed to review the current implementation, and outline N-M hypotheses for what could be wrong.

We review these hypotheses, sometimes iterate, and then tackle them one by one.

An important note for debug flows, similar to manual debugging, it's often better to have the agent instrument logging/traces/etc. to confirm a hypothesis, before moving directly to a fix.

Using this method has led to a 100% vibe-coded success rate both on greenfield and legacy projects.

Note: my main complaint is the sheer number of markdown files over time, but I haven't gotten around to (or needed to) automate this yet, as sometimes these historic planning/debug files are useful for future changes.

jumploops··on If AI writes code, should the session be part of the commit?
I've been experimenting with a few ways to keep the "historical context" of the codebase relevant to future agent sessions.

First, I tried using simple inline comments, but the agents happily (and silently) removed them, even when prompted not to.

The next attempt was to have a parallel markdown file for every code file. This worked OK, but suffered from a few issues:

1. Understanding context beyond the current session

2. Tracking related files/invocations

3. Cold start problem on an existing codebases

To solve 1 and 3, I built a simple "doc agent" that does a poor man's tree traversal of the codebase, noting any unknowns/TODOs, and running until "done."

To solve 2, I explored using the AST directly, but this made the human aspect of the codebase even less pronounced (not to mention a variety of complex edge-cases), and I found the "doc agent" approach good enough for outlining related files/uses.

To improve the "doc agent" cold start flow, I also added a folder level spec/markdown file, which in retrospect seems obvious.

The main benefit of this system, is that when the agent is working, it not only has to change the source code, but it has to reckon with the explanation/rationale behind said source code. I haven't done any rigorous testing, but in my anecdotal experience, the models make fewer mistakes and cause less regressions overall.

I'm currently toying around with a more formal way to mark something as a human decision vs. an agent decision (i.e. this is very important vs. this was just the path of least resistance), however the current approach seems to work well enough.

If anyone is curious what this looks like, I ran the cold start on OpenAI's Codex repo[0].

[0]https://github.com/jumploops/codex/blob/file-specs/codex-rs/...

jumploops··on Show HN: enveil – hide your .env secrets from prAIng eyes
In the context of traditional SaaS, using dynamic secrets loaded at runtime (KMS+Dynamo, etc.).

For agentic tools and pure agents, a proxy is the safest approach. The agent can even think it has a real API key, but said key is worthless outside of the proxy setting.

jumploops··on A few random notes from Claude coding quite a bit last few weeks
Yes, exactly.

The LLM is onboarding to your codebase with each context window, all it knows is what it’s seen already.

jumploops··on A few random notes from Claude coding quite a bit last few weeks
Yeah to be clear it will have the same issues as a flyby contributor if prompted to.

Meaning if you ask it “handle this new condition” it will happily throw in a hacky conditional and get the job done.

I’ve found the most success in having it reason about the current architecture (explicitly), and then to propose a set of changes to accomplish the task (2-5 ways), review, and then implement the changes that best suit the scope of the larger system.

jumploops··on A few random notes from Claude coding quite a bit last few weeks
I’ve found that LLMs seem to work better on LLM-generated codebases.

Commercial codebases, especially private internal ones, are often messy. It seems this is mostly due to the iterative nature of development in response to customer demands.

As a product gets larger, and addresses a wider audience, there’s an ever increasing chance of divergence from the initial assumptions and the new requirements.

We call this tech debt.

Combine this with a revolving door of developers, and you start to see Conway’s law in action, where the system resembles the organization of the developers rather than the “pure” product spec.

With this in mind, I’ve found success in using LLMs to refactor existing codebases to better match the current requirements (i.e. splitting out helpers, modularizing, renaming, etc.).

Once the legacy codebase is “LLMified”, the coding agents seem to perform more predictably.

YMMV here, as it’s hard to do large refactors without tests for correctness.

(Note: I’ve dabbled with a test first refactor approach, but haven’t gone to the lengths to suggest it works, but I believe it could)

jumploops··on Prism
I’ve been “testing” LLM willingness to explore novel ideas/hypotheses for a few random topics[0].

The earlier LLMs were interesting, in that their sycophantic nature eagerly agreed, often lacking criticality.

After reducing said sycophancy, I’ve found that certain LLMs are much more unwilling (especially the reasoning models) to move past the “known” science[1].

I’m curious to see how/if we can strike the right balance with an LLM focused on scientific exploration.

[0]Sediment lubrication due to organic material in specific subduction zones, potential algorithmic basis for colony collapse disorder, potential to evolve anthropomorphic kiwis, etc.

[1]Caveat, it’s very easy for me to tell when an LLM is “off-the-rails” on a topic I know a lot about, much less so, and much more dangerous, for these “tests” where I’m certainly no expert.

jumploops··on Kimi Released Kimi K2.5, Open-Source Visual SOTA-Agentic Model
> For complex tasks, Kimi K2.5 can self-direct an agent swarm with up to 100 sub-agents, executing parallel workflows across up to 1,500 tool calls.

> K2.5 Agent Swarm improves performance on complex tasks through parallel, specialized execution [..] leads to an 80% reduction in end-to-end runtime

Not just RL on tool calling, but RL on agent orchestration, neat!

jumploops··on AI code and software craft
> People have said that software engineering at large tech companies resembles "plumbing"

> AI code [..] may also free up a space for engineers seeking to restore a genuine sense of craft and creative expression

This resonates with me, as someone who joined the industry circa 2013, and discovered that most of the big tech jobs were essentially glorified plumbers.

In the 2000s, the web felt more fun, more unique, more unhinged. Websites were simple, and Flash was rampant, but it felt like the ratio of creators to consumers was higher than now.

With Claude Code/Codex, I've built a bunch of things that usually would die at a domain name purchase or init commit. Now I actually have the bandwidth to ship them!

This ease of dev also means we'll see an explosion in slopware, which we're already starting to see with App Store submissions up 60% over the last year[0].

My hope is that, with the increase of slop, we'll also see an increase in craft. Even if the proportion drops, the scale should make up for it.

We sit in prefab homes, cherishing the cathedrals of yesteryear, often forgetting that we've built skyscrapers the ancient architects could never dream of.

More software is good. Computers finally work the way we always expected them to!

[0]https://www.a16z.news/p/charts-of-the-week-the-almighty-cons...

jumploops··on Unrolling the Codex agent loop
That’s what I used to think, before chatting with the OAI team.

The docs are a bit misleading/opaque, but essentially reasoning persists for multiple sequential assistant turns, but is discarded upon the next user turn[0].

The diagram on that page makes it pretty clear, as does the section on caching.

[0]https://cookbook.openai.com/examples/responses_api/reasoning...

jumploops··on Unrolling the Codex agent loop
I think the delta may be an overloaded use of "turn"? The Responses API does preserve reasoning across multiple "agent turns", but doesn't appear to across multiple "user turns" (as of November, at least).

In either case, the lack of clarity on the Responses API inner-workings isn't great. As a developer, I send all the encrypted reasoning items with the Responses API, and expect them to still matter, not get silently discarded[0]:

> you can choose to include reasoning 1 + 2 + 3 in this request for ease, but we will ignore them and these tokens will not be sent to the model.

[0]https://raw.githubusercontent.com/openai/openai-cookbook/mai...

jumploops··on Unrolling the Codex agent loop
Maybe it's changed, but this is certainly how it was back in November.

I would see my context window jump in size, after each user turn (i.e. from 70 to 85% remaining).

Built a tool to analyze the requests, and sure enough the reasoning tokens were removed from past responses (but only between user turns). Here are the two relevant PRs [0][1].

When trying to get to the bottom of it, someone from OAI reached out and said this was expected and a limitation of the Responses API (interesting sidenote: Codex uses the Responses API, but passes the full context with every request).

This is the relevant part of the docs[2]:

> In turn 2, any reasoning items from turn 1 are ignored and removed, since the model does not reuse reasoning items from previous turns.

[0]https://github.com/openai/codex/pull/5857

[1]https://github.com/openai/codex/pull/5986

[2]https://cookbook.openai.com/examples/responses_api/reasoning...

jumploops··on Unrolling the Codex agent loop
One thing that surprised me when diving into the Codex internals was that the reasoning tokens persist during the agent tool call loop, but are discarded after every user turn.

This helps preserve context over many turns, but it can also mean some context is lost between two related user turns.

A strategy that's helped me here, is having the model write progress updates (along with general plans/specs/debug/etc.) to markdown files, acting as a sort of "snapshot" that works across many context windows.

jumploops··on I was banned from Claude for scaffolding a Claude.md file?
I believe Claude Code recently turned on max reasoning for all requests. Previously you’d have to set it manually or use the word “ultrathink”
jumploops··on Hands-On Introduction to Unikernels
Boot is a misleading term, but you can resume snapshotted VMs in single digit ms

(and without unikernels, though they certainly help)

jumploops··on Gas Town Decoded
I ran the Gas Town intro post through ChatGPT 5.2 Pro[0]

Based on my initial read, and a pass at this summary, it seems mostly right. YMMV

Did some further dives into the little public usage data from Gas Town, and found that most of the "Beads" are tasks that are broken down quite small, almost too small imo.

Super interesting project with the goal of keeping Claude "busy" however it feels more like a casino game than something I'd use for production engineering.

[0]https://gist.github.com/jumploops/2e49032438650426aafee6f43d...

jumploops··on The creator of Claude Code's Claude setup
I’ve found that experienced devs use agentic coding in a more “hands-on” way than beginners and pure vibe-coders.

Vibecoders are the best because they push the models in humorous and unexpected ways.

Junior devs are like “I automated the deploy process via an agent and this markdown file”

Seasoned devs will spend more time writing the prompt for a bug fix, or lazily paste the error and then make the 1-line change themselves.

The current crop of LLMs are more powerful than any of these use cases, and it’s exciting to see experienced devs start to figure that out (I’m not stanning Gas Town[0], but it’s a glimpse of the potential).

[0]https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16d...

jumploops··on Welcome to Gas Town
To be fair, the author says: "Do not use Gas Town."

I started "fully vibecoding" 6 months ago, on a side-project, just to see if it was possible.

It was painful. The models kept breaking existing functionality, overcomplicating things, and generally just making spaghetti ("You're absolutely right! There are 4 helpers across 3 files that have overlapping logic").

A combination of adjusting my process (read: context management) and the models getting better, has led me to prefer "fully vibecoding" for all new side-projects.

Note: I still read the code that gets merged for my "real" work, but it's no longer difficult for me to imagine a future where that's not the case.

jumploops··on Welcome to Gas Town
Curious what fidelity/precision the author finds necessary with Claude 4.5 Opus/GPT 5.2.

Looking at the screenshot of "Tracked Issues", it seems many of the "tasks" are likely overlapping in terms of code locality.

Based on my own experience, I've found the current crop of models to work well at a slightly higher-level of complexity than the tasks listed there, and they often benefit from having a shared context vs. when I've tried to parallelize down to that level of work (individual schema changes/helper creation/etc.).

Maybe I'm still just unclear on the inner workings, but it's my understanding each of those tasks is passed to Claude Code and developed separately?

In either case, I think this project is a glimpse into the future of software development (albeit with a grungy desert punk tinted lens).

For context, I've been "full vibe-coding"[0] for the past 6 months, and though it started painfully, the models are now good enough that not reading the code isn't much of an issue anymore.

jumploops··on Building an internal agent: Code-driven vs. LLM-driven workflows
> why can't there be an LLM that would always give the exact same output for the exact same input

LLMs are inherently deterministic, but LLM providers add randomness through “temperature” and random seeds.

Without the random seed and variable randomness (temperature setting), LLMs will always produce the same output for the same input.

Of course, the context you pass to the LLM also affects the determinism in a production system.

Theoretically, with a detailed enough spec, the LLM would produce the same output, regardless of temp/seed.

Side note: A neat trick to force more “random” output for prompts (when temperature isn’t variable enough), is to add some “noise” data to the input (i.e. off-topic data that the LLM “ignores” in it’s response).

jumploops··on Always bet on text (2014)
> Text is the oldest and most stable communication technology

Minor nit: complex language (i.e. Zipf’s law) is the oldest and most stable communication technology.

Before text, we had oral story telling. It allowed us to communicate one generation’s knowledge to the next, and so on.

Arguably this is present elsewhere in the animal kingdom (orcas, elephants, etc.), but human language proves to be the most complex.

Side note: one of my favorite examples is from the Gunditjmara (a group of Aboriginal Australians) who recall a volcanic eruption from 30k+ years ago [0].

Written language (i.e. text) is unique, in that it allows information to pass across multiple generations, without a man-in-the-middle telephone-like game of storytelling.

But both are similar, text requires you to read, in your own voice, the thoughts of another. Storytelling requires you to hear a story, and then communicate it to others.

In either case, the person is required to retell the knowledge, either as an internal monologue or as an external broadcast.

Always bet on language.

[0]https://en.wikipedia.org/wiki/Budj_Bim

jumploops··on Parasites plagued Roman soldiers at Hadrian's Wall
Can we use the same argument for life among the stars?

Intelligence, even?

jumploops··on Show HN: Turn raw HTML into production-ready images for free
Love the simplicity and “Not MCP” callout (:

Not that it matters, but curious what percentage of this service was “vibe-coded”?

jumploops··on Show HN: HN Wrapped 2025 - an LLM reviews your year on HN
> You’ve mentioned the 1975 book The Mythical Man-Month so many times that I’m starting to think it’s your only personality trait besides complaining about Tailwind CSS.

Ahahaha, not entirely wrong!

jumploops··on OpenAI are quietly adopting skills, now available in ChatGPT and Codex CLI
I think the future is likely one that mixes the kitchen-sink style MCP resources with custom skills.

Services can provide an MCP-like layer that provides semantic definitions of everything you can do with said service (API + docs).

Skills can then be built that combine some subset of the 3rd party interfaces, some bespoke code, etc. and then surface these more context-focused skills to the LLM/agent.

Couldn’t we just use APIs?

Yes, but not every API is documented in the same way. An “MCP-like” registry might be the right abstraction for 3rd parties to expose their services in a semantic-first way.

jumploops··on GPT-5.2
Is that technically not a new pretrained model?

(Also not sure how that would work, but maybe I’ve missed a paper or two!)

jumploops··on My productivity app is a never-ending .txt file (2020)
Not to be the “ai” guy, but LLMs have helped me explore areas of human knowledge that I had postponed otherwise

I am of the age where the internet was pivotal to my education, but the teacher’s still said “don’t trust Wikipedia”

Said another way: I grew up on Google

I think many of us take free access to information for granted

With LLMs, we’ve essentially compressed humanity’s knowledge into a magic mirror

Depending on what you present to the mirror, you get some recombined reflection of the training set out

Is it perfect? No. Does it hallucinate? Yes. It it useful? Extremely.

As a kid that often struggled with questions he didn’t have the words for, Google was my salvation

It allowed me to search with words I did know, to learn about words I didn’t know

These new words both had answer and opened new questions

LLMs are like Google, but you can ask your exact question (and another)

Are they perfect? No.

The benefit of having expertise in some area, means I can see the limits of the technology.

LLMs are not great for novelty, and sometimes struggle with the state of the art (necessarily so).

Their biggest issue is when you walk blindly, LLMs will happily lead the unknowing junior astray.

But so will a blogpost about a new language, a new TS package with a bunch of stars on GitHub, or a new runtime that “simplifies devops”

The biggest tech from the last five years is undoubtedly the magic mirror

Whether it can evolve to Strong AI or not is yet to be seen (and I think unlikely!)

jumploops··on GPT-5.2
It’s possible they’re using some new architecture to get more up-to-date data, but I think that’d be even more of a headline.

My hunch is that this is the same 5.1 post-training on a new pretrained base.

Likely rushed out the door faster than they initially expected/planned.

jumploops··on GPT-5.2
> “a new knowledge cutoff of August 2025”

This (and the price increase) points to a new pretrained model under-the-hood.

GPT-5.1, in contrast, was allegedly using the same pretraining as GPT-4o.

← PreviousPage 5 of 19Next →