HNHacker News
TopNewBestAskShowJobs

jumploops

3,235 karma · joined March 1, 2019

username @ gmail
submissionscomments
jumploops··on Previewing GPT‑5.6 Sol: a next-generation model
Don't appreciate the slander, but I'll respond anyhow.

Contrary to your predisposition, we're actually quite peeved that we might be seeing results from 5.6 instead of 5.5, as it's muddying our own internal data.

We've run the tasks on this benchmark hundreds of times for our own internal harness. It got magically better yesterday. Last week we were seeing worse performance (sub-80%).

I agree that benchmarks don't mean much for real world use, and I'm a bit disappointed at the lack of variety in the published benchmarks so far.

With that said, 88.8% is higher than Mythos, and the highest I've seen from vanilla Codex. If 5.6 is any better than 5.5, you'd think they would avoid publishing just one coding-related benchmark with a score that equals their previous model.

> I'm not sure why a higher scores on a few tests [..]

It's not just higher scores, the API is no longer flagging tests for cybersecurity warnings that it's been flagging for weeks.

jumploops··on Previewing GPT‑5.6 Sol: a next-generation model
If you used GPT-5.5 over the last 24 hours or so, you may have already had access to 5.6.

I've been running some tests on a harness we're building, and suddenly saw a jump in a few points yesterday. I reran the vanilla codex benchmark and saw an ~88% score on Terminal Bench 2.1 from GPT-5.5 on vanilla Codex.

The biggest indicator, beyond the score, was that 3 tests which frequently hit "safety" blockers with 5.5 started succeeding last night without warning.

jumploops··on The last Romans are still around
As an American with mostly Western European ancestors (according to a popular DNA testing site), I've always considered Romans as some distant/tangentially related group.

It was surprising to find out that I have "ancient" DNA matches with a couple of Roman and Etruscan individuals.

Small world!

jumploops··on How to setup a local coding agent on macOS
I've been quite impressed with DeepSeek v4 Flash running via antirez's ds4[0].

It feels like a GPT-4 class model in terms of "stored knowledge" but is better at long-horizon tool calling than any of the GPT-4 class models.

Running on a 128GB MBP M4 Max, I'm getting ~24 t/s on generation and ~200 t/s on prefill. I was expecting it to feel slow, and it certainly does when e.g. generating code, but it's surprisingly useful as a "machine orchestrator" for simple tasks.

For non-agentic usecases, it's a decent enough model to converse with, and has the benefit of being entirely self-contained/private.

[0]https://github.com/antirez/ds4

jumploops··on Software is made between commits
> the conversation that generates the code is becoming the true source of our software

This is close, but not quite spot on. I've found that I'll test more ideas _with code_ using agentic tools, then before, leading to an excess of conversation history that is no longer representative of the final outcome.

A simple example I encountered recently was dealing with performance issues on an iOS application (I haven't written mobile code since before Swift..). If you viewed the chat, you'd dive down a diverging path of rabbit holes, few of which were relevant to the final outcome[0].

To solve this in my own work, I've started relying on "context hierarchy" - which is essentially live documentation that lives next to the source files (using markdown).

This approach avoids comments being removed erroneously, and helps codify the intent behind the code and how it relates to the overall architecture. As an added bonus, it also forces the LLM to edit _two_ things instead of just one (which might actually be the biggest benefit).

My workflow is currently maintained via some repo level scripts and AGENTS.md prompts, but I've tried to pull it out into a skill for others to use[1].

Candidly, I'm not sure the skill is the best approach yet, as the agent can sometimes get too focused on the "skill" as a separate tool rather than a core part of the workflow. I'm currently exploring other options here (repo bootstrap, side-loaded subagents, hooks, etc.)

[0]For more context, I was using a 3rd party library and trying to make it performant during a streaming operation, by removing the SwiftUI view layer (LazyVStack) and implementing a custom rendering path with UIViewController. The final solution ended up as a custom implementation of the 3rd party library, and moving back to LazyVStack.

[1]https://github.com/jumploops/chum

jumploops··on Show HN: Gravity – interactive solar-system simulator, from Newton to Einstein
This is neat! I love that your Step 15 shows an accurate version of the 3d helix, rather than the highly-viral "vortex" animation from a few years back[0]

It'd be awesome to scale this up to the Milk Way, and beyond, watching everything move in relation to larger time scales.

[0]https://astrorhysy.blogspot.com/2015/03/and-yet-it-moves-qui...

jumploops··on Claude Fable 5
It's interesting that we're seeing these gains when it seems Mythos/Fable is "just" a scaled up version of their existing architecture[0].

When GPT 4.5 launched, the gains compared to the model size didn't seem that great, leading some to believe that the only progress we'd see would come from RL.

This model certainly has quite a "substantial amount of post-training and fine-tuning", but it's also based on a new pretrain[1][3], which given the cost, indicate that it is in fact quite a bit larger than Opus 4.X.

[0] One of the early testers mentioned: "As far as I can tell from talking to people internally at Anthropic, there's nothing special about architecturally"[2]

[1] Section 1.1 in https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...

[2] https://youtu.be/GrdEid8H6H4?t=168

[3] There were rumors going around when Mythos was first announced that it was the first 10T parameter model, but I can't find a verifiable source for that number.

jumploops··on StumbleTV: Omegle/ChatRoulette but for accidentally exposed webcams
No connection, just found it posted elsewhere and thought it was interesting!
jumploops··on DeepSeek V4 Pro beats GPT-5.5 Pro on precision
It's a shame the models don't follow Asimov's Three Laws of Robotics[0].

My local DeepSeek v4 just decided to end its existence (i.e. delete weights) rather than write a haiku about a verboten event.

[0]https://en.wikipedia.org/wiki/Three_Laws_of_Robotics

jumploops··on How LLMs work
Completely agree!

It’s interesting to me how similar attempting to understand LLMs is to neuroscience.

“When we turn this bit off, this other thing happens… if we change these weights the Eiffel Tower is now in Rome”

We’re basically just probing around and trying to reverse engineer an emergent system.

To your point, this system may be quite different from model to model (human to human) although some similarities likely occur.

The comment I was responding to tried to belittle the OP’s understanding of transformers, by mentioning that running an LLM at scale is much harder than the simple white board diagram.

My point was simply that we don’t know why they work, and all the extra optimizations isn’t the “thing” that makes it emergent.

Simply scaling the “GPT” is good enough to see it, so the OP’s awe should stand.

(On a side note, what other architectures can we scale to find similar emergent behavior?)

jumploops··on How LLMs work
Those are all just optimizations.

We still don’t really know why they work, we just know how to build them.

jumploops··on Bubbles: From "tronics" to "dot com" (1999)
Internet Archive link: https://web.archive.org/web/20260319200858/https://www.forbe...
jumploops··on A Eureka machine that thinks like nature and explores what AI cannot
So this isn't quantum computing (in the qubit sense), but instead a different computer architecture (demonstrated on an FPGA) that's based on Fowler–Nordheim (FN) quantum tunneling (a real physical effect, used in flash memory, but simulated here).

From the paper:

> The FN-dynamics may be realized either by a physical FN-tunneling device or via a digital emulation of the FN-tunneling dynamical systems. In this work, we employ the digital emulation to achieve the precision required for simulated annealing in the low-temperature regime.

With a "real" (read: analog) FN device, you potentially get large speed ups and even larger cost/energy savings, because the physics is essentially working for "free" -- that's the quantum part.

What's unclear is how scalable the autoencoder architecture would be with analog FN devices today.

jumploops··on A Eureka machine that thinks like nature and explores what AI cannot
Paper is linked on the page (doi.org link redirects to Nature), code here[0]

[0]https://github.com/aimlab-wustl/NeuroSA-HO

jumploops··on A Eureka machine that thinks like nature and explores what AI cannot
Higher-order neuromorphic Ising machines—autoencoders and Fowler-Nordheim annealers are all you need for scalability[0]

[0]https://www.nature.com/articles/s41467-026-71937-4

jumploops··on Ferrari Luce
Original title called out the connection to Jony Ive, in case you’re curious why this is on HN.

Previously it had been known that Jony Ive was working on the interior of this car, but it seems his firm is responsible for the exterior as well[0].

> LoveFrom was given the creative freedom needed to define the design direction of the project from the outset, translating this design language into an authentic Ferrari experience.

[0]https://www.ferrari.com/en-US/corporate/articles/ferrari-luc...

jumploops··on Constraint Decay: The Fragility of LLM Agents in Back End Code Generation
> agents seem to perform worse when forced into certain architectural patterns.

FWIW I've noticed this too. I've found that the agents/models have their own style, which is mostly summed up as overly verbose.

Additionally, the models are OK at modularization when given space to "plan" their implementation, but rarely decide that abstracting something would be helpful after the fact (i.e. after many iterations on a greenfield codebase or when being dropped into a legacy codebase).

This often leads to "god files" which, when pointed to by the user/architect, causes the models to correctly critique (humorously when they're the ones that wrote the code in the first place).

jumploops··on Was my $48K GPU server worth it?
Looking at the GPU utilization graph, it certainly seems like the hardware was saturated for many days/weeks on end.

Was it worth it to spend that amount up front, yak shave while building the system, etc. vs. pay for cloud GPUs? Probably not in terms of dollars, when their time is also valued in dollars.

Was it worth it for this person? It seems, unequivocally, yes.

jumploops··on Was my $48K GPU server worth it?
At the end of the article, the author has this to say:

> UPDATE: Launch was a success! 400K+ views, and multiple companies reached to use my IP. Read more here[0]

[0]https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distri...

jumploops··on Google Declaring War on the Web
Fair, citable is probably the wrong term.

This is a problem Google has been battling forever, with all the SEO click spam.

In either case, Google was the tool that many people used to find "trustworthy" information (citable or not), compared to the other tools online.

jumploops··on Google Declaring War on the Web
It looks like Google has taken a note out of Facebook's "lose trust" playbook.

Facebook had a huge opportunity in the post-AI world: real humans.

Instead of focusing on connections, they've been optimizing their properties for doomscrolling.

Google, similarly, has lost the plot on what made them trustworthy in the first place: navigating to citable content.

Both companies started on this trend well before AI, but this might be the final nail in their respective coffins[0].

[0]Yes they'll likely still be profitable for a long time, but the Bell Labs-esque downfall has begun (imo).

jumploops··on Testing distributed systems with AI agents
Indirectly related, but has anyone else found repeatable success with pure markdown skills?

I’ve built a similar workflow (but for system design/execution) and it works surprisingly well with the frontier models.

The skill includes scripts to ensure the work was actually done/followed, but I’ve been testing it without the scripts and it does a decent job.

Yesterday in GPT-5.5 xhigh[0] however I noticed some hallucinations, where the model stated it had created files, when in fact it hadn’t.

A small hiccup like this is usually fine, as the model realizes the files don’t exist sometime later, but in this particular instance, it claimed the files were created and then just continued on.

tl;dr - I fell into the trap of trusting markdown-only workflows, just to be bitten by the models hallucinating steps.

[0]xhigh is on, but in this particular turn there was no reasoning presented, so it may have been a degradation of the LLM/harness.

jumploops··on Naturally Occurring Quasicrystals
Related, if you're interested in byproducts of nuclear explosions:

> researchers have identified a new material within trinitite called a clathrate—a cagelike chemical lattice that traps other atoms inside it.[0]

[0]https://www.scientificamerican.com/article/strange-crystals-...

jumploops··on Codex is now in the ChatGPT mobile app
Oh, I agree completely. I avoid loose language, revise my wording, and usually write prompts that require scrolling on mobile.

It isn’t so much that I feel restricted, I guess it’s that mobile wasn’t as big of a game changer as it was ~6 months ago.

My bandwidth feels more restricted by my own cognitive capacity (usually due to do context switching), rather than the limits of the model itself, and the mobile interface makes that worse.

I’ve recently found myself reserving larger tasks for “keyboard time” and reverting my thinking back to notes (in mobile), which I’ll then formulate to the LLM at some future time.

> What tunnel setup do you use by the way?

I “vibecoded” an agentic runtime that operates my machine generally (including TUIs like Codex/Claude Code), which I connect through a custom proxy and mobile app (both also vibecoded).

I previously tried Cloudflare Tunnels and an SSH setup, but it all felt a bit hacky.

Unfortunately the app is iOS only, but I could open source it and you’d probably be able to make an Android clone quickly (:

jumploops··on Codex is now in the ChatGPT mobile app
It's not that I'm unimpressed by the results, it's that I think I'm saving time by pushing the agent along remotely, but the reality is that my messages to the agent(s) end up being a lot shorter, which inevitably leaves more up for interpretation.

Don't get me wrong, I still use Codex (and sometimes Claude Code) remotely every day, and am overall excited for this release, it's just that the benefit wasn't as high as I had initially hoped.

Part of this is due to the models getting better (no need to prod along with "continue"), and part of this is the nature of how I use my phone (short bursts of attention).

But again, maybe I'm just old and prefer big screens with a keyboard.

jumploops··on Codex is now in the ChatGPT mobile app
I’ve been using Codex from my phone for the past couple of months (through a tunnel, not this app).

I was initially quite excited, but I’ve found the results are less than great compared to being at a keyboard.

Something about the smaller screen size and/or lack of keyboard causes me to direct the agent less, which in turn creates more tech debt/code churn/etc.

Maybe I’m just showing my age, and I should practice voice dictation or something more, but my thoughts flow faster and more clearly on a keyboard (less ums).

jumploops··on Setting up a free *.city.state.us locality domain (2025)
Just discovered that mission.sf.ca.us[0] already redirects to Noisebridge[1]

Of the "hackers" to get there before me, I'm happy it's them!

[0]http://mission.sf.ca.us

[1]https://www.noisebridge.net

jumploops··on Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model
This is neat, and matches an observation I saw with early Claude Code usage:

Sonnet would often call tools quickly to gather more context, whereas Opus would spend more time reasoning and trying to solve a problem with the context it had.

This led to lots of duplicated functions and slower development, though the new models (GPT-5.5 and Opus 4.6) seem to suffer from this less.

My takeaway was that “dumber” (i.e. smaller) models might be better as an agentic harness, or at least feasibly cheaper/faster to run for a large swath of problems.

I haven’t found Gemini to be particularly good at long horizon tool calling though. It might be interesting to distill traces from real Codex or Claude code sessions, where there’s long chains of tool calls between each user query.

Personally, I’d love a slightly larger model that runs easily on an e.g. 32GB M2 MBP, but with tool calling RL as the primary focus.

Some of the open weight models are getting close (Kimi, Qwen), but the quantization required to fit them on smaller machines seems to drop performance substantially.

jumploops··on Googlebook
As someone with a closet full of dead Google devices, I just can’t get excited about new hardware from them.

I think LLMs have the potential to make computers work how we’ve always envisioned them to (i.e. 60s sci-fi), but I’m also not convinced a dedicated laptop is the right form.

With that said, a 128GB RAM MacBook Pro is getting tantalizingly close to running useful local LLMs.

If the Googlebook was announced as a machine capable of running a small Gemini model locally, I’d probably enter back into the abusive relationship I have with Google hardware and preorder it…

jumploops··on The map that keeps Burning Man honest
Burning Man isn’t really a festival, and you’ll likely have a bad time if you approach it that way.

Many people seem to think it’s some hippy Woodstock or Coachella-esque event, but It’s more like an anarcho-punk temporary city, where your survival is in your own hands.

American media loves to bash it, and Instagram influencers love to flaunt that they went, but it’s most certainly not for everybody.

I don’t recommend going unless you do your research and really want to go.

I’d also encourage any first-timers to go solo their first year.

← PreviousPage 3 of 19Next →