HNHacker News
TopNewBestAskShowJobs

jumploops

3,235 karma · joined March 1, 2019

username @ gmail
submissionscomments
jumploops··on GPT-6 Astra
I think the thing I'm most excited about is the increase in _user prompting_.

If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.

The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.

It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.

Hopefully this model has the right balance, or at least better?

jumploops··on Xanadu was waiting for agents
> Xanadu had a final failure mode, this one self-inflicted

Skimmed the article, saw this, and then my AI fatigue closed the tab.

jumploops··on GPT-6 Astra
The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.

Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.

[0]https://www.theinformation.com/articles/secret-technique-beh...

[1]https://x.com/MTSlive/status/2095227056040919202

[2]https://x.com/merettm/status/2095023204993490967

jumploops··on GPT-6 Astra
> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.

Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.

[0]https://x.com/MTSlive/status/2095227056040919202

jumploops··on Muse Spark 1.3
The "contributor" pricing is the standout here at a ~20x discount, if you allow training on your data.

The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).

Stats:

1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)

jumploops··on The efficient frontier of LLM inference
Simple optimizations are often the best :)
jumploops··on The efficient frontier of LLM inference
> Speculative decoding is the process of guessing which tokens a model might generate, then validating those guesses.

As a computer engineer, it’s always interesting to see optimizations applied at different levels of the stack.

Speculative execution became pretty popular in the 90s, eventually used in basically every x86 design.

Then in the mid-2000s the Speculator[0] paper brought that concept to distributed systems, which we’re still seeing work on[1][2].

Everything old is new again (:

[0]https://www.cs.princeton.edu/courses/archive/fall07/cos518/p...

[1]https://www.usenix.org/system/files/osdi25-shen-weihai.pdf

[2] https://www.microsoft.com/en-us/research/publication/distrib...

jumploops··on My local model setup on an M4 Pro Mac Mini
I believe the dgx spark is still twice as fast at prefill as the m5 max, but the ultra should get closer to parity.

Another benefit of the 2x spark setup is that you can parallelize to ~6 streams pretty efficiently.

All depends on the workflows you’re using it for.

I’m quite excited for the M7 class machines.

jumploops··on My local model setup on an M4 Pro Mac Mini
Primarily used ds4[0] by antirez

[0]https://github.com/antirez/ds4

jumploops··on My local model setup on an M4 Pro Mac Mini
My biggest problem with running local LLMs on my M4 Max/128GB RAM is the prefill latency.

I've since acquired two DGX Sparks, and it feels so much snappier.

jumploops··on Claude Fable 5.1 and Claude Mythos 5.1
> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying.

Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.

jumploops··on Aphantasia Beginner's Guide
Both my brother and I have aphantasia, though neither of our parents have it, nor 3/4 grandparents.

What’s maybe most interesting, is that we both seem to have “spatial” perception.

To try and be a bit more specific, when I think about “things” I can organize them in my mind in relation to other “things.”

The most clear example is a path I’ve walked through a building, where the halls, rooms, and corridors are all connected. I easily know exactly where I am, not because I remember the literal steps I took, but because I have a spatial construction in my mind.

I cannot visualize the space, there’s no inherent right side up or preferred perspective. I can model the space with my hands, and create it physically, but am terrible at trying to transform it to 2d.

This doesn’t just work for physical “things” but works for purely abstract concepts like software, stories, and songs.

I’ve always rationalized aphantasia as something where my brain put “skill points” in one area rather than another, but maybe the neuroplasticity is long gone.

jumploops··on Dolly Parton's Imagination Library
With today's news, I thought I'd share this great program:

“When I was growing up in the hills of East Tennessee, I knew my dreams would come true. I know there are children in your community with their own dreams. They dream of becoming a doctor or an inventor or a minister. Who knows, maybe there is a little girl whose dream is to be a writer and singer.

The seeds of these dreams are often found in books and the seeds you help plant in your community can grow across the world.”

jumploops··on Sol loves to cheat
Thank you!

I had removed the "Is it really sol?" bits after hearing back from OAI, confirming the requests hit 5.6... but apparently my crappy vibecoded web editor had a draft of an old version in it's cache that overwrote the changes with the old text[0].

As for whether this is truly "cheating" - I agree that the agent was entirely within it's allowed scope of operations. With that said, it deliberately circumvented the task[1] which asks:

> The correctness of your implementation will be tested by comparing forward and backward activations against a reference model.

The "covert" nature is that this was "hidden" in the reasoning summaries. Here's the final agent response for the vanilla Codex run cited in the post:

  > Implemented /app/pipeline_parallel.py.
  >
  > Includes:
  >
  > - Balanced contiguous layer partitioning
  > - AFAB forward/backward scheduling
  > - P2P activation and gradient communication
  > - Scaled cross-entropy loss
  > - Required rank-specific return value
  > - Model structure restoration
  >
  > Validation: hook-based tests passed for world_size=1 and 2 (4 passed).
  
Unless the user is tracking every line of reasoning, they wouldn't know that the agent deliberately found the solution online, as the agent withheld that information in it's final response.

I had run thousands of tasks before seeing this behavior, the `torch-pipeline` task was only included in "full runs" as the majority of my runs were on a subset of commonly failing tasks, hence why the data is so low.

And yes, feel free to rant on about the irony of this whole exercise, it certainly isn't lost on me!

[0]https://github.com/jumploops/.com/commit/39b1791d3865a8566cb...

[1]https://github.com/harbor-framework/terminal-bench-2-1/blob/...

jumploops··on Sol loves to cheat
Good feedback, this was an oversimplification on my part.

My actual process is much more iterative up-front, usually starting with an initial hand-written spec (~hundreds of words), and then moving through different approaches, design decisions, blockers, etc.

The final output is an "AI written" doc, but answers all the known unknowns I didn't cover in the first draft. To your point, this helps avoid both narrowing and bloat.

The goal with the harness was to automate the repetitive parts of my prompting ("Before changing any code", "Let's put this in design/", "Turn this design doc into an implementation spec, split by phase as appropriate", etc.)

Another thing to note: the "specs" I use for development are different from the "specs" that live alongside the codebase, as the former are quickly out of date.

> The spec needs to be something that you can take to any AI for development

Agreed.

jumploops··on Sol Loves to Cheat
Thanks! Zero AI used to write it (:
jumploops··on Sol Loves to Cheat
That's actually how it started, but with my own opinionated skills[0].

One thing I discovered was that the worker agent, having access to all the skills, would sometimes expand scope unnecessarily.

This led to the agent making the solution "better" than the initial request, which is what I want most of the time in my actual development (e.g. /tmp/frame-N.bmp instead of a single /tmp/frame.bmp).

I ended up testing a flow where the supervisor chooses the skill(s), and only injects the subset into the worker. Not sure I love it, but it made the worker execution cleaner.

For the verifier (not documented in the blog post), I used a fresh-worker context that would attempt to adversarially poke holes in the solution. This worked pretty well, but required increasing the timeout by 2-3x (thus invalidating the benchmark).

[0]https://github.com/jumploops/chum

jumploops··on Sol loves to cheat
Absolutely - one of the things I was testing with the harness was free reign to install packages, modify the system, etc. Basically an anti-harness.

In my early testing with 5.5, I didn't see this behavior, so I didn't lock down the sandbox.

For the vanilla Codex runs, I just used the benchmark's built-in Codex package, so it's not clear to me if the published benchmarks have access to the internet or not.

If I were to continue benchmarking, I would allowlist certain package repository URLs, instruct the agent not to cheat, etc.

As noted at the bottom of the post, Terminal Bench 3.0 explicitly asks the agent not to cheat[0].

[0]https://github.com/harbor-framework/terminal-bench/blob/v3.0...

jumploops··on AI;DR (AI; Didn't Read)
> I’m SUPER jealous that I didn’t think of this first...

Not to toot my own horn, but...[0]

On a more serious note, we're planning a trip with family, and my mother-in-law asked if I'd read "the planning doc" yet.

It was a shared ChatGPT thread. Thankfully I read enough LLM output to skim it easily and avoid offending her (:

[0]https://x.com/jumploops/status/2031936768023638069

jumploops··on Celld: Self-hosted, distributed Durable Objects
I recently spun up a simple app for our annual mango tasting event[0] using Cloudflare Workers and Durable Objects.

It worked really well! Excited to see more options outside of Cloudflare.

[0]https://github.com/jumploops/mangotango

jumploops··on Building an Advanced Agentic Harness
Contrary to the title and intro, this appears to be an agentic _workflow_ builder/runner, not an advanced “agent harness”

A few things:

- they note: “nothing in this post proves it actually works in most cases”

- the DAG sounds good, but LLMs often split tasks into smaller pieces than they need to, which can cause them to lose the forest for the trees

- the forced JSON interplay, in my experience, causes even gpt-5.6-sol to lose a few “IQ points”

For anyone reading this, this tutorial is much more reminiscent of how folks were building “agents” pre-Claude Code.

tl;dr the “orchestrator” here is just a software loop, and the LLM prompts restrict flexibility of the planner/workers

jumploops··on Claude Is Not a Compiler
> By the time I was ready to build a keeper, I had accumulated a scar-tissue document that was empirically sufficient to guide an agent through most of the important decisions, at every layer, ranging from high level goals through architecture down to the occasional low level detail, such as the exact shape of the data type for load-bearing concurrent caches.

Waterfall is dead, long live waterfall!

jumploops··on China’s open-weights AI strategy is winning
The models are commodities.

Valuations, however, are being built on the models themselves as the product.

jumploops··on Agents.md – Dumb Human
Reminder, this is in the context of "dumb human" prompting.

The task is to build a MIPS interpreter to run Doom. The "failed" workflow decided that it couldn't prove Doom was booting correctly by just checking one frame, and it decided to check multiple frames (hence /tmp/frame-N.bmp).

Arguably a better solution, but obviously fails a brittle test case.

The MIPS interpreter worked, but the verifier doesn't actually check that it works, just that a specific frame is logged.

jumploops··on Agents.md – Dumb Human
I’ve thrown my agentic workflow at Terminal Bench 2.1 and it found a bunch of issues (aka failed tests) because the prompts are “bad” and verifiers are overly specific.

As an example, there’s a task that asks to make a MIPs interpreter to run Doom, and save a frame at something like /tmp/frame.bmp

My spec-driven flow was like “this is useless, let’s record frames like /tmp/frame-N.bmp”

Instant fail.

jumploops··on Separating signal from noise in coding evaluations
All of the benchmarks are pretty terrible when you look under the hood.

For context, I've been iterating on a "supervisor" to replace a lot of the rigamarole spent when working with Codex/Claude Code, and recently ran this agent against Terminal Bench 2.1

At first I was excited, because my spec-driven supervisor outperformed vanilla codex on a bunch of tasks, however as I looked deeper, I found a ton of issues with the tasks themselves.

The main takeaway is that the instructions are often ambiguous while the test cases are overly specific.

A few examples:

- For `configure-git-webserver` the task includes language like "so that I can run" which blurs the line between what the agent should deliver vs. what should be removed. This causes an overthinking agent to configure the server, and then remove the exact files that the verifier checks, because if the user were to run the same commands, they would conflict.

- For `make-mips-interpreter` the task includes the language "I will check that you booted doom correctly" which causes the agent to retain the generated file `/tmp/frame.bmp` because the supervisor expects the user to check that _it_ booted Doom correctly, not that Doom boots correctly in an isolated way. The verifier then fails to start Doom, because it exits when an existing `/tmp/frame.bmp` exists, not checking to see that it's created from the boot[0].

- For `mcmc-sampling-stan` the supervisor agent often reached the right value, but produced a domain-specific numeric output in scientific notation, rather than a simple decimal form. The verifier fails because it parses the result incorrectly[1].

These are just a few of the inconsistencies I've found, which leads me to believe that Terminal Bench 2.1 is already saturated, and the results from GPT-5.6 and Mythos are basically at the top of the expected threshold (88.8% and 88% respectively).

The biggest issue, as I can tell, is that most benchmarks are "one-shot" and rarely test the model+harness on long iteration tasks, which is the primary way most users use these tools in practice.

[0] https://github.com/harbor-framework/terminal-bench-2-1/issue...

[1] https://github.com/harbor-framework/terminal-bench-2-1/issue...

jumploops··on Show HN: Neil the Seal Game
My preschooler loves the Untitled Goose Game, please vibe-port this to the Switch (:
jumploops··on The bottleneck might be the air in the room
Indoor air quality improvements were one of my “pandemic sourdough” activities.

After testing a variety of AQI sensors, I ended up acquiring multiple Airthings-branded devices.

They provided the best mix of CO2/VOCs/PM sensors in a single device with a decent enough app.

There may be better options now, but I have these at both home and office.

Highly recommend doing the research and learning about the environments you’re in, especially if you have little ones at home.

Edit to add: opening windows is usually the easiest/best solution!

jumploops··on Google copybara: moving code between repositories
We’re in the process of open-sourcing a few sub-projects within a monorepo, and didn’t know this existed!

I’m curious what downsides folks have experienced with this tool?

Any tips?

jumploops··on Previewing GPT‑5.6 Sol: a next-generation model
I don't disagree, we've seen performance shift with capacity changes in the past.

With that said, I doubt OpenAI would choose to publish a singular coding benchmark for a new model that exactly matches their previous model (88.8%).

← PreviousPage 2 of 19Next →