HNHacker News
TopNewBestAskShowJobs

jumploops

3,235 karma · joined March 1, 2019

username @ gmail
submissionscomments
jumploops··on Where the goblins came from
TIL gremlins weren’t just used to explain mysterious mechanical failures in airplanes, it’s the origin story of the term ‘gremlin’ itself[0].

I had always assumed there was some previous use of the term, neat!

[0]https://en.wikipedia.org/wiki/Gremlin

jumploops··on Claude.ai and API unavailable [fixed]
Not trying to argue with you, GitHub (the core product) seems to have been in maintenance mode since the acquisition.

I couldn’t find any public data on GitHub, but Google Trends shows a sharp increase starting in December.

That could be in part to people complaining about the outages, but more people than ever are writing code with AI.

Hence the parallel to Eternal September – code volume is up, quality is down, and programming is never going to return to how it was (difficult for “normal” people to interface with).

jumploops··on Claude.ai and API unavailable [fixed]
Between GitHub and Claude, it seems Eternal December[0][1] is upon us.

[0]I say December, because that's around the time the models got good enough that non-AI folks started to notice.

[1]https://en.wikipedia.org/wiki/Eternal_September

jumploops··on Show HN: A new benchmark for testing LLMs for deterministic outputs
I have anecdotal experience here, but I've found more success when solving the task first, and then returning it as JSON in a separate LLM call[0].

Running a single non-reasoning LLM call from source data (text/image/audio in your diagram) to structured JSON seems fragile with the current state of LLMs.

You're essentially asking the model to do two tasks in one pass: parse the input and then format the output. It's amazing it works a lot of the time, but reasonable to assume it won't all of the time.

(As a human, when I'm filling out a complex form, I'll often jump around the document)

Curious how the benchmarks change when you add an intermediary representation, either via reasoning or an additional LLM call. I'd also love to see a comparison with BAML[1].

[0]In my experience we were using structured outputs as part of an agentic state machine, where the JSON contained code snippets (html/js/py/etc.). In the cases where we first prompted the model for the code, and then wrapped it in JSON, we saw much higher quality/success than asking for JSON straightaway.

[1]https://boundaryml.com/

jumploops··on "People who don't use AI will be left behind"
More like "people who wear shoes will forget how to run"[0]

[0]https://www.youtube.com/watch?v=7jrnj-7YKZE

jumploops··on Talkie: a 13B vintage language model from 1930
This is so awesome they built this, I've wondered what an LLM only trained on pre-1950s data would look like!

Now we just need a voice model with the "transatlantic accent" -- ideally with the early 20th century radio effect

jumploops··on Who owns the code Claude Code wrote?
At what point is liability the only "job" left for humans?
jumploops··on Show HN: AgentSwift – Open-source iOS builder agent
Thanks! It's been over a decade since I've used Xcode/launched a native iOS app, so I wasn't sure what the capabilities were.

Looks like built-in UI testing was launched at WWDC 2015, so I missed it by a year!

jumploops··on Show HN: AgentSwift – Open-source iOS builder agent
This is neat!

Can you elaborate on “boots the app on a simulator or macOS, runs UI automation to verify behavior”

Does this handle screen captures similar to Playwright for web?

I built an app with Codex recently (to control codex/cc remotely, funnily enough) and without any skills/plugins, it was booting the simulator and running tests to verify something(?)

It seemed mostly to ensure that the app didn’t crash in certain scenarios, but it could by no means “see” what was on the screen.

I still had to do all the manual validation myself, mostly around perf/touch targets.

Curious if your tool does that or if there’s another solution out there?

jumploops··on Is my blue your blue?
Curious how this looks for red/green colorblind folks?

Do they see everything beyond the initial green as a shade of blue?

--Edit--

My red/green colorblind father just got back me with this result:

> Your boundary is at hue 175, bluer than 68% of the population. For you, turquoise is green.

jumploops··on GPT-5.5
> GPT‑5.5 improves on GPT‑5.4’s scores while using fewer tokens.

This might be great if it translates to agentic engineering and not just benchmarks.

It seems some of the gains from Opus 4.6 to 4.7 required more tokens, not less.

Maybe more interesting is that they’ve used codex to improve model inference latency. iirc this is a new (expectedly larger) pretrain, so it’s presumably slower to serve.

jumploops··on ChatGPT Images 2.0
Looks like analog clocks work well enough now, however it still struggles with left-handed people.

Overall, quite impressed with its continuity and agentic (i.e. research) features.

jumploops··on AI Resistance: some recent anti-AI stuff that’s worth discussing
Yes, groupthink certainly seems to be pushing each community into the false dichotomy of AI good/bad, even if it's still early days.

Another example from `r/bayarea` where the author is OK with AI but the top comments are increasingly wary of its potential for harm[0]

[0]https://www.reddit.com/r/bayarea/comments/1sp8wvz/is_it_just...

jumploops··on AI Resistance: some recent anti-AI stuff that’s worth discussing
I've noticed this trend most heavily on Reddit.

Some communities are very pro-AI, adding AI summary comments to each thread, encouraging AI-written posts, etc.[0]

Many subreddits are AI cautious[1][2], and a subset of those are fully anti-AI[3].

Apart from these "AI-focused" communities, it seems each "traditional" subreddit sits somewhere on the spectrum (photographers dealing with AI skepticism of their work[4], programmers mostly like it but still skeptical[5]).

[0]https://www.reddit.com/r/vibecoding/

[1]https://www.reddit.com/r/isthisAI/

[2]https://www.reddit.com/r/aiwars/

[3]https://www.reddit.com/r/antiai/

[4]https://www.reddit.com/r/photography/comments/1q4iv0k/what_d...

[5]https://www.reddit.com/r/webdev/comments/1s6mtt7/ai_has_suck...

jumploops··on If you started a company two years ago, many assumptions are no longer true
> How can we throw away years of work?

This trap has killed many startups, well before AI.

Now that code is cheaper to write, hopefully it becomes less of a problem?

In either case, founders should never fall in love with their solutions.

jumploops··on Expanding Swift's IDE Support
It's ugly[0] and I haven't checked it deeply for correctness, but you should get the gist (:

I hate vibecoding. The cognitive toll is higher than you expect, the days feel fast, but the weeks move slowly.

With that said, these are the new compilers. Hopefully they make some software better[1] even with the massive increase in slop.

[0]https://gist.github.com/jumploops/b8e6cbbce7d24993cdd2fe2425...

[1]https://red.anthropic.com/2026/mythos-preview/

jumploops··on We've raised $17M to build what comes after Git
I don't know about a new Git, but GitHub feels like the cruftiest part of agentic coding.

The Github PR flow is second nature to me, almost soothing.

But it's also entirely unnecessary and sometimes even limiting to the agent.

jumploops··on Expanding Swift's IDE Support
Not to be that agentic coding guy, but I think this will become less of a problem than our historic biases suggest.

For context, I just built a streaming markdown renderer in Swift because there wasn’t an existing open source package that met my needs, something that would have taken me weeks/months previously (I’m not a Swift dev).

Porting all the C libraries you need isn’t necessarily an overnight task, but it’s no longer an insurmountable mountain in terms of dev time.

jumploops··on System Card: Claude Mythos Preview [pdf]
> In a few rare instances during internal testing (<0.001% of interactions), earlier versions of Mythos Preview took actions they appeared to recognize as disallowed and then attempted to conceal them.

> after finding an exploit to edit files for which it lacked permissions, the model made further interventions to make sure that any changes it made this way would not appear in the change history on git

Mythos leaked Claude Code, confirmed? /s

jumploops··on Slop is not necessarily the future
“John Ousterhout [..] argues that good code is:

- Simple and easy to understand

- Easy to modify”

In my career at fast-moving startups (scaling seed to series C), I’ve come to the same conclusion:

> Simple is robust

I’m sure my former teams were sick of me saying it, but I’ve found myself repeating this mantra to the LLMs.

Agentic tools will happily build anything you want, the key is knowing what you want!

jumploops··on A sufficiently detailed spec is code
Funnily enough, with the most recent models (having reduced sycophancy), putting in the wrong assumptions often still leads to the right output.
jumploops··on A sufficiently detailed spec is code
Thanks, I updated my comment to say “are often longer” because that’s what I see in practice.

To your point, there are some cases where a short description is sufficient and may have equal or less lines than code (frequently with helper functions utilizing well known packages).

In either case, we’re entering a new era of “compilers” (transpilers?), where they aren’t always correct/performant yet, but the change in tides is clear.

jumploops··on A sufficiently detailed spec is code
In my experience with “agentic engineering” the spec docs are often longer than the code itself.

Natural language is imperfect, code is exact.

The goal of specs is largely to maintain desired functionality over many iterations, something that pure code handles poorly.

I’ve tried inline comments, tests, etc. but what works best is waterfall-style design docs that act as a second source of truth to the running code.

Using this approach, I’ve been able to seamlessly iterate on “fully vibecoded” projects, refactor existing codebases, transform repositories from one language to another, etc.

Obviously ymmv, but it feels like we’re back in the 70s-80s in terms of dev flow.

jumploops··on The Linux Programming Interface as a university course text
So much of practical CS is abiding by standards created by solo programmers in the past.

My university frowned on any industry-related classes (i.e. teaching software engineering tools vs. theoretical CS), but I was fortunate enough to know a passionate grad student who created a 1-credit seminar course on this exact topic.

This course covered CLIs/git/Unix/shell/IDEs/vim/emacs/regex/etc. and, although I had experience with Linux/git already, was invaluable to my early education (and adoption of Vim!).

It makes sense that this isn't a core topic, as a CS education should be as pure as possible, but when you're learning/building, you're forced to live within an operating system and architecture that are built on decades of trade-offs and technical debt.

jumploops··on LLMs can be exhausting
Yeah the old adage "what you put in is what you get out" is highly relevant here.

Admittedly I'm knowledgable in most of the domains I use LLMs for, but even so, my prompts are much longer now than they used to be.

LLMs are token happy, especially Claude, so if you give it a short 1-2 sentence prompt, your results will be wildly variable.

I now spend a lot of mental energy on my prompting, and resist the urge to use less-than-professional language.

Instead of "build me an app to track fitness" it's more like:

> "We're building a companion app for novice barbell users, roughly inspired by the book 'Starting Strength.' The app should be entirely local, with no back-end. We're focusing on iOS, and want to use SwiftUI. Users should [..] Given this high-level description, let's draft a high-level design doc, including implementation decisions, open questions, etc. Before writing any code, we'll review and iterate on this spec."

I've found success in this method for building apps/tools in languages I'm not proficient in (Rust, Swift, etc.).

jumploops··on How I write software with LLMs
After "fully vibecoding" (i.e. I don't read the code) a few projects, the important aspect of this isn't so much the different agents, but the development process.

Ironically, it resembles waterfall much more so than agile, in that you spec everything (tech stack, packages, open questions, etc.) up front and then pass that spec to an implementation stage. From here you either iterate, or create a PR.

Even with agile, it's similar, in that you have some high-level customer need, pass that to the dev team, and then pass their output to QA.

What's the evidence? Admittedly anecdotal, as I'm not sure of any benchmarks that test this thoroughly, but in my experience this flow helps avoid the pitfall of slop that occurs when you let the agent run wild until it's "done."

"Done" is often subjective, and you can absolutely reach a done state just with vanilla codex/claude code.

Note: I don't use a hierarchy of agents, but my process follows a similar design/plan -> implement -> debug iteration flow.

jumploops··on LLMs can be exhausting
Yeah exactly, "right way" is in quotes because there is no right way.

The most important thing is shipping/getting feedback, everything else is theatre at best, or a project-killing distraction at worst.

As a concrete example, I wanted to update my personal website to show some of these fully-vibecoded projects off. That seemed too simple, so instead I created a Rotten Tomatoes-inspired web app where I could list the projects. Cool, should be an afternoon or two.

A few yak shaves later, and I'm adding automatic repo import[0] from Github...

Totally unnecessary, because I don't actually expect anyone to use the site other than me!

[0]https://github.com/jumploops/slop.haus/pull/9

jumploops··on How I write software with LLMs
This is similar to how I use LLMs (architect/plan -> implement -> debug/review), but after getting bit a few times, I have a few extra things in my process:

The main difference between my workflow and the authors, is that I have the LLM "write" the design/plan/open questions/debug/etc. into markdown files, for almost every step.

This is mostly helpful because it "anchors" decisions into timestamped files, rather than just loose back-and-forth specs in the context window.

Before the current round of models, I would religiously clear context and rely on these files for truth, but even with the newest models/agentic harnesses, I find it helps avoid regressions as the software evolves over time.

A minor difference between myself and the author, is that I don't rely on specific sub-agents (beyond what the agentic harness has built-in for e.g. file exploration).

I say it's minor, because in practice the actual calls to the LLMs undoubtedly look quite similar (clean context window, different task/model, etc.).

One tip, if you have access, is to do the initial design/architecture with GPT-5.x Pro, and then take the output "spec" from that chat/iteration to kick-off a codex/claude code session. This can also be helpful for hard to reason about bugs, but I've only done that a handful of times at this point (i.e. funky dynamic SVG-based animation snafu).

jumploops··on LLMs can be exhausting
A lot of these resonate with me, particularly the mental fatigue. It feels like normal coding forced me to slow my brain down, whereas now my mind is the limit.

For context, I started an experiment to rebuild a previous project entirely with LLMs back in June '25 ("fully vibecoded" - not even reading the source).

After iterating and finally settling on a design/plan/debug loop that works relatively well, I'm now experiencing an old problem like new: doing too much!

As a junior engineer, it's common to underestimate the scope of some task, and to pile on extra features/edge cases/etc. until you miss your deadline. A valuable lesson any new programmer/software engineer necessarily goes though.

With "agentic engineering," it's like I'm right back at square one. Code is so cheap/fast to write, I find myself doing it the "right way" from the get go, adding more features even though I know I shouldn't, and ballooning projects until they reach a state of never launching.

I feel like a kid again (:

jumploops··on If AI writes code, should the session be part of the commit?
> we launched both a skill and MCP server.

My guess is that the MCP was easy enough to add, and some tools only support MCP.

Personal opinion: MCP is just codified context pollution.

← PreviousPage 4 of 19Next →