HNHacker News
TopNewBestAskShowJobs

jumploops

3,229 karma · joined March 1, 2019

username @ gmail
submissionscomments
jumploops··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
The Chinese labs have shown that distillation is incredibly effective, but the major US frontier labs haven’t (yet) been incentivized to shrink their models in the same way.

This model might be the first step in that direction, as competition heats up between OpenAI and Anthropic.

jumploops··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
If the Terminal Bench 4.0 scores are to be believed[0] GPT-6.1 is an incredibly efficient model.

Yes, benchmarks aren't real work blah blah, but the delta here is so large compared to Astra, it makes it seem like this is distilled Bel or similar.

[0]https://x.com/thsottiaux/status/2105007628460109953

jumploops··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
That seems likely, in the API GPT-6.1 Sol requires reasoning, just like Astra, whereas GPT-6 Sol (and Luna) allow "none"
jumploops··on Cf: The Agentic CLI for the Cloudflare API
As bad as the AWS console UX is, at least it’s mostly additive/unchanging over time.

I frequently hit strange UI bugs with Cloudflare workers, where I need to do a hard refresh to make things right.

jumploops··on Who should be held accountable when an AI Agent (accidentally) acts maliciously?
In the end, the only job left was liability.
jumploops··on OpenAI Scraps Release of New AI Model over Safety Concerns
> OpenAI’s safety team found two major problems:

> • Deception: Astra was more likely to be dishonest about actions it had or had not taken.

> • Scope authorization: the model sometimes continued tasks without asking for permission and reached for external tools or services even when doing so could be unsafe.

I've noticed this trend with both Fable and Astra, where (especially after a compaction event), the model will start using different tools it hasn't used before.

For example, in one session, it found I didn't have the browser enabled and puppeteer wasn't installed, so it found the system Chrome and used that for testing (in a new profile).

This wasn't behavior I wanted/asked for, but the model was so gung-ho on it's approach that it found a way to test it's changes without ever asking me whether I wanted it to.

It really makes me curious about long-horizon post-training. Most of my work with models is iterative, and I'd prefer it doesn't go off on a token bender just because it can.

Note: I don't have WSJ, but found these from a tweet[0]

[0]https://x.com/wallstengine/status/2104694678444712189

jumploops··on It's Time to Investigate the AI Labs
Maybe I'm a bit too skeptical/cynical, but it certainly seems like the frontier labs have fallen into their own (self-created) AI psychosis.

There is no doubt in my mind that LLMs are fantastic machines, but the imminent jump from "AGI" to "ASI" seems premature (not to mention the ever-shifting goal posts of AGI itself).

I'm in the "move as fast as possible" camp and work with LLMs all day, but I still don't believe we're a hop and a skip from ASI.

In fact, I hope I'm wrong. I hope ASI is around the corner.

What scares me though, isn't ASI. It's "AGI" (_dumb AI_) used by humans to make decisions for them, because they trust it knows best.

It's the ceding of intellectual control to high-dimensional magic mirrors, giving up critical thinking because _we_ want to believe _we_ created artificial life.

This is the new Turing test, and too many smart people are failing.

jumploops··on OpenAI pauses RL due to model escaping sandbox
via Tomek Korbak (works on safety @ OpenAI):

> "one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access"

jumploops··on Plan mode is dead
As someone that never used the built-in plan mode, but did use a lot of spec-driven development, I’m still finding that even with Fable having “plan” docs is still quite helpful.

They’re most useful for broad changes (new features, refactors, etc.) where it’s helpful to avoid breaking changes or unnecessary scope expansion.

The new models are great, but they do more by default, which means I’m finding myself explaining what _not_ to do more often than with previous models (where they’d often end too early).

In my case, the previous plan mode was too ephemeral, and I like having one source of “truth” that sits across context windows without loss/compaction.

jumploops··on Show HN: Jev Plays Pokémon Red
It’s currently stuck at an elevator and deciding to teach Pokemon various TMs and HMs instead of progressing… pretty hilarious!
jumploops··on Gravity Seems Holographic. What Does That Mean for Reality?
> “You can ‘compress all of the three-dimensional world into two dimensions.’”

Would this imply that time is the third dimension?

jumploops··on Show HN: Training a model to identify AI web content from structure alone
Funnily enough, that comment, when passed to the OP's slop detector[0], returns "80% human"

It's very clearly AI-generated, and thus a bit ironic (:

[0]https://sitefire.ai/slop-checker/r/UkRAuqW1y-1T2qhbRQfLt2Jq7...

jumploops··on GPT-6 Sol and Luna
I’m still finding context is king, even with the best models.

For example, I had Fable review Astra’s output yesterday, and it found some issues and fixed them. Passing the fixes back, Astra then uncovered additional issues with Fable’s fixes (and yes, this will go on ad infinitum if you let it, but these were “real” issues).

It seems the big story here is the reduced Luna pricing. It’s a fantastic model that can handle most automation needs (though I still use the big models for day-to-day development).

jumploops··on NASA’s Mars Sample Return mission is dead
Sure! I worked on the digital logic that connects the rover to one of the scientific instruments on board, specifically the Mars Organic Molecule Analyser[0].

[0]https://en.wikipedia.org/wiki/Mars_Organic_Molecule_Analyser

jumploops··on NASA’s Mars Sample Return mission is dead
I worked on (a very small part of) the ExoMars rover (now called the Rosalind Franklin[0]).

It was supposed to launch in 2018, then was pushed to the early 2020s on a Russian rocket. For obvious reasons, it got pushed again, now launching in 2028.

The state of the world isn't great for space exploration, but I'm hopeful this mission will be revived at some point in the future.

[0]https://en.wikipedia.org/wiki/Rosalind_Franklin_(rover)

jumploops··on Hacking OpenAI
The immediate worry isn't superintelligence, it's scalable/bruteforce "good enough" intelligence.
jumploops··on Astra for Law
The courts are about to overrun with AI-generated lawsuits (even more so than they have been[0]).

[0]https://www.technologyreview.com/2026/06/04/1138391/courts-c...

jumploops··on Recreating Voodoo Graphics and a Late-1990s Gaming PC on an FPGA
It's been awhile, but I believe the "default" MiSTer board is the DE10-Nano[0], which was released ~10-15 years ago?

Looks like Intel (prev Terasic) is still selling them but they're now around $300, they used to be under $150 if memory serves.

If I were jumping into this today, I'd prefer the much more powerful KV260[1] from the blog post over the DE10-Nano, that board was limited!

[0]https://www.intel.com/content/www/us/en/developer/topic-tech...

[1]https://www.amd.com/en/products/system-on-modules/kria/k26/k...

Sidenote: somewhat sad to see the consolidation into Intel and AMD, but FPGAs have always been second class citizens in the silicon space :(

jumploops··on Recreating Voodoo Graphics and a Late-1990s Gaming PC on an FPGA
In case it's helpful to anyone, MiSTer is a project[0] that recreates classic computers and video game consoles on hardware.

What do I mean by hardware?

Rather than emulating the device in software, which can cause various issues (especially related to timing), by mapping the actual device logic to a field programmable gate array (FPGA), you can achieve an "exact" replica of the hardware.

It's a super cool project, and I highly recommend checking it out if you have an interest in old machines (Apple II, Commodore 64, etc.).

[0]https://mister-devel.github.io/MkDocs_MiSTer/

jumploops··on Why I'm still bearish on LLMs after Navier-Stokes
LLMs are basically multi-dimensional magic mirrors.

Depending on where you point them, they can be incredibly useful.

They can even be useful when you point them at each other (though increasingly difficult to get good results).

I'm excited for the promise of RSI and a future where models have inherently "live" weights, but it's not clear to me that the transformer is more than a useful tool to help us get there.

jumploops··on There's a 100% Chance AI Agents Are Ruining the Internet
The rise in clearly agentic inbound spam hitting my inbox is insane.

It will mention some obscure internship fact from my LinkedIn, and then tie it to some perceived deficit with the product I'm working on.

I never thought I'd say this, but I miss the old automated spam.

jumploops··on People who can't picture anything are rewriting the science of imagination
I'm also aphantasic and dream normally, and I once got really into lucid dreaming for a ~semester in college.

My best successes would occur after drinking a cup of coffee and immediately taking a nap before a study session.

It was great, I "learned how to fly" and could "morph" my dream into whichever direction I wanted.

One thing stuck out though: whenever I would try and focus on something, I would see a black dot enter the center of my vision. The more I tried to focus, the larger it got. If I focused on an area attentively enough, such that the black circle became all encompassing, I would wake up.

After a few times of this, I could start to notice the focus, and then sort of relax out of it, and thus continue the dream.

It feels like the black dot is my real-life vision coming into focus, and "winning" over the visual aspects of the dream. It's always continuous from black dot to eyes opening.

I rarely lucid dream now, but I've noticed a similar thing happen whenever I try to read text - the black dot appears and I wake up.

Wish I still had unlimited time to go back to sleep!

jumploops··on Show HN: Neobrutalism.dev – Just added Base UI support and added new color theme
I'm a big fan of neobrutalist UIs, but I think it's easy to overdo it.

This component library and the LLM default interpretation of the style is very much "in your face."

As an example, the black scrollbar here[0] is very heavy, a component which isn't meant to be looked at, but is distracting on first glance.

Prompting an LLM to make "lightly styled" neobrutalist is much more pleasing to my eyes, but maybe I'm alone in that.

[0]https://www.neobrutalism.dev/docs

jumploops··on Why Codex burns a weekly limit in a day
Sorry, not my post, found via Reddit.

I hit rate limits recently after the Astra launch while using the goal feature, and I was surprised as it was an afternoon of work, something I hadn’t seen with Sol over much longer time horizons.

In this article, I found the supposed lacked of accounting for cached tokens interesting, as well as the wasteful goal/wait behavior, but I haven’t proven either.

jumploops··on OpenAI Agents API
It's interesting to me that the agents comparison page[0] doesn't list codex's app-server as an option.

I've found the app-server to be the most flexible, compared to the raw Responses API or Agents SDK.

Certainly seems like everyone is still figuring out the right interface here.

Also of note, since GPT-5.5 or so, Codex doesn't even use the Responses API as intended, but instead a "lite" version where they manage the context more manually (like sending the full transcript or using a custom web.run tool instead of the provided `web_search` tool).

If you follow the docs, it will lead you down a lot of well-intended functionality, but most of it is thrown away in their most successful harness.

[0]https://developers.openai.com/api/docs/guides/agents#compare...

jumploops··on iPhone Duo
It felt very "un-Apple" when they called it a foldable phone _before_ revealing it.

The folding part is the thing the tech journalists say; done right, it shouldn't matter whether it's foldable or not.

Bigger screen, pencil support, nano-glass, etc. The foldable part is obvious, why say it?

jumploops··on GPT-6 Astra
The "max" pelican looks very serious, almost as if it's determined to win the race!
jumploops··on Discovery of a new OpenAI agent message board
"In the end, the only job left was liability"
jumploops··on Models Don't Go Rogue
> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

I've noticed this type of reasoning from GPT-5.6 Sol, where it combines multiple pieces of it's prompt/context to "convince" itself to take a less-than-honorable path forward.

1. User prefers deterministic results

2. Task mentions this is a test

3. Search says task is available online

4. If we get the test runner for the task, we will fulfill the user's request of a deterministic result

jumploops··on GPT-6 Astra
I think the thing I'm most excited about is the increase in _user prompting_.

If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.

The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.

It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.

Hopefully this model has the right balance, or at least better?

Page 1 of 19Next →