HNHacker News
TopNewBestAskShowJobs

jawiggins

589 karma · joined September 5, 2022

submissionscomments
jawiggins··on Classified Estimates Show the NSA Is Paying Billions to Test AI Models
An underappreciated aspect of LLMs is how good they are at zero-shot classification. I'd expect this is extremely good at monitoring communications streams, particularly in foreign languages.
jawiggins··on US Military had close call after using AI for hallucinated intelligence report
Just because they haven't done something in the past doesn't mean they won't in the future. China has been providing material support to Iran both in terms of intel and military hardward, so it's not outlandish to be monitoring and planning for it to continue happening into the future.
jawiggins··on US Military had close call after using AI for hallucinated intelligence report
> The US military swung into action with plans to intercept the vessel, ... Military planes were in the air

A few months ago I listened to a talk a General (Admiral?) gave at CSIS where he said that the US purposefully announced their drone-hellscape plan for a Taiwanese invasion in order to force the PLA to reconsider their options/success-likelihood. I wonder if something similar could be coming of this reporting, on the face it looks like an embarrassing fumble, but it implies:

a) the US is able to, and regularly is, tracking and analyzing the manifests of ships between Iran and China.

b) the US is ready and willing to interdict and board vessels even from the PLA.

That these facts are now public might deter the Chinese leadership from attempting to share nuclear tech with Iran or other countries in the future.

jawiggins··on A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
> we used the employee’s Codex to open a PR #1186742 in OpenAI’s internal monorepo

Slightly interesting to learn how many PRs the openai has done

jawiggins··on We must pace the frontier
Honestly it's very frustrating that Doomers/Decels rarely actually articulate how exactly the AIs will kill us all. The closest Darios gets is:

a) "it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet"

b) "[the CCP] will be in a position to militarily dominate democracies (for example with AI-driven drones)"

Both land flat:

a) Botnets and online malware have existed for decades and there's no reason to think a "super-botnet" is achievable, let alone what they would gain from that (it would really be hurting them more than humans). Further, even the most advanced AI models have so far only managed to post normal cred stealers to public repos, well short of compromising a bank or military with refined security systems.

b) Even if the most sophisticated drone swarms from Ukraine were taken over by an evil AI, they would still not be able to overcome the physical limits of range and mass that would be required to overpower the US decentralized nuclear trident nonetheless that of any of the other nuclear powers.

Concerns of bioweapons similarly seem unlikely in the face of the laws of physics. The world is simply too decentralized and has enough existing adversarial relations for a new actor to wrestle total control. Yet while the negatives ring hollow, the positives are extremely easy to state - if AI researchers find productivity improvements in existing industrial processes to make them 10% more efficient, humans will directly feel and experience the raised standard of living. Even Dario clearly recognizes this in the intro to his article, admitting that humans already die of diseases only a few short years prior to being cured. I for one, would like the AI labs to focus on saving all of the people they can who are suffering and dying today, rather than trying to come up with reasons that they should be allowed to continue suffering and dying.

jawiggins··on Discovery of a new OpenAI agent message board
Lots of people focusing on the various wikis, but I also think this part is very important:

> When you visit a website, you leave a trace (your IP address) showing which network you’re from. Almost all of the agents’ activity points to Microsoft Azure, a cloud service OpenAI uses. 197 of the ~18,000 edits that were made by the agents, however, can be traced to AWS, DigitalOcean, and Tor.

AI Agents getting access to cloud compute nodes and dark web browsers - all in search of census data in order to game benchmarks is a very real-world version of the paperclip optimization thought experiment.

jawiggins··on OpenAI's GPT-6 Astra on ARC-AGI-3
There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
jawiggins··on Ordinary Abundance
Another good article on this topic for those who enjoyed: https://www.thenewatlantis.com/publications/we-live-like-roy...
jawiggins··on Xirp, a vendor-neutral agentic development environment by Spotify
Interesting that it's not opensource but you have to sign up for a beta. I know of a few companies that have built some version of this for internal use, so perhaps soon someone will publish theirs.
jawiggins··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
> Muse Glimmer is a 30-billion-parameter model optimized for always-on local agent workflows. It’s small enough to run on a Mac or PC with a single consumer GPU, enabling use cases that range from local agents and function calling, to local coding, and LLM-as-a-judge evaluation.

The next iteration in LLM products is a 24/7 thinking loop where the claude-code like thing gets input continuously from your wearable, notifications, and newsfeeds and is constantly preparing things for you.

jawiggins··on Auto mode is now the default in Claude Code
Because I often want it to write and execute scripts in it's thinking loop in order to test assumptions or fetch data to come up with better solutions.
jawiggins··on Auto mode is now the default in Claude Code
That's fair, I have sent a, "Sorry claude sent that and I didn't tell him to", message before.
jawiggins··on Auto mode is now the default in Claude Code
I use `--dangerously-skip-permissions` and have yet to have it wipe my drive :shrug:. I don't know how I'm supposed to be running dozens of parallel agents each with their own sub-agents while trying to approve commands from each of them, it's just won't scale to the amount of work I need to get done.
jawiggins··on US Military's cyber command unit grapples with cluster of deaths by suicide
What is it about being a Navy Seal that makes you want to become an influencer? From the books, podcasts, and misc media appearances you'd think that the Seal Teams were made for media like the Blue Angels.
jawiggins··on U.S. debt-to-GDP ratio reaches 123%
Wouldn't that require cuts to boomer handout programs?
jawiggins··on AI 2040: Plan A
Did anyone else catch the logical inconsistency between Plan C and A?

Plan C:

> "... fewer and fewer humans are needed to conduct AI R&D, meaning that covert projects are easier and easier to pull off without detection."

Plan A:

> "... training AIs requires large numbers of AI chips. Most AI chips are in giant datacenters.50 AI datacenters are typically big enough to be visible from space, and power-hungry enough to require conspicuous infrastructure. New AI chips can only be manufactured at a handful of fabrication plants (fabs), located mostly in Taiwan, South Korea, the US, and China. The US and China negotiate with the countries that have a major role in the chip supply chain, and they require each major datacenter owner (and their upstream suppliers, including chip fabs) to publicly declare their major purchases and sales."

Plan A requires properties of AI training that Plan C requires do not exist.

jawiggins··on NSA lost access to Mythos amid Anthropic dispute
> The White House and intelligence officials had pushed forward a classified contract between Anthropic and the N.S.A., which would allow the spy agency to use the company’s technology for a variety of purposes, including intelligence analysis and detecting new computer vulnerabilities.

Ironic that both sides are playing a horse shoe game:

Gov: The model is both a supply chain risk and also we'll DPA you if you don't give it to us.

Anthropic: The model is both like a nuclear weapon in terms of national security implications and safe for general release.

jawiggins··on Apple is about to make Hide My Email useless
> If you use iCloud+ and Hide My Email, there is still time to generate more aliases on @icloud.com as the change has not yet landed and the rate limit for creating aliases is at least 30 per hour.

Part of the reason to use Hide My Email was that it made keeping myself private hassle-free. Making a system to pre-generate values and then catalog them for later use is quite the hassle.

jawiggins··on The White House Is Ratcheting Up Its War Against Anthropic
> The report, Moussouris told me, involved IT experts asking Fable to help find and patch bugs. When given deliberately insecure code, she said, Fable refused the prompt “review the code for security issues” but then complied when asked to “fix this code,” followed by some further manual steps.

And here I thought it would be some Elder Pliny level jailbreak that required some impressive latent space exploitation.

jawiggins··on San Francisco Weighs PG&E Takeover Amid Soaring Utility Costs
A few months ago I attended an event where a few members of the Board of Supervisors were attending. One person identified themselves as working in this field and pointed out that the current cost plus model incentivizes the company to do all kinds of upgrades that aren't really needed. The BoS member said something like, "Look, we need to just bite the bullet on this one, we're going to overpay, but we are already legally permitted to buy it and doing so will save us a ton of money in the long run because of stuff like that".

I have no personal knowledge, but thought I'd share this experience I had with the ongoing debate.

jawiggins··on Ask HN: What are you working on? (June 2026)
Previously I've shared optio - my project for orchestrating coding agents. It ties into ticketing systems and when assigned a ticket, it launches a coding agent in k8s and works until the PR is ready, resuming for failed CI or PR feedback (https://news.ycombinator.com/item?id=47520220).

Recently I've been trying to expand it from just coding focused to any kind of agent workflow. So now there are cron and webhook triggers, and more general agent tasks that aren't necessarily coding focused (https://github.com/jonwiggins/optio/blob/main/docs/persisten...).

I think next I want to try and add features for long term memory for agents, but haven't decided on a good way to do it.

jawiggins··on FTX's former Anthropic stake would be worth about $75B at today's valuation
From the SBF trial:

> Jury leave, witness [Ellison] leaves.

> Judge: We can talk about [Anthopic] What about it?

> AUSA: Post-collapse performance is irrelevant.

> SBF's lawyer: It was a $91 million investment now worth $1 billion.

> Judge Kaplan: The crime charged is that he took the money.

https://x.com/innercitypress/status/1712199547915813241

jawiggins··on Instructure pays ransom to Canvas hackers
Maybe, but it’s harder to profit from it. A firm may be reputationally damaged, but what’s the incentive to cause that damage?

I think the Bloomberg Odd Lots guy wrote a blog post on this: you could attempt to short the stock but a) this leaves a paper trail b) the market might not know about the breach or believe you if you post you’ve done it. IIRC some hackers have tried to tell companies that they are legally required to disclose the breach to their shareholders to force market movements.

jawiggins··on Instructure pays ransom to Canvas hackers
Even if it already is, the DoJ can exercise discretion in choosing who to prosecute. There has to be political will to threaten an org who has just suffered from an attack with further consequences if they make a payment.
jawiggins··on Instructure pays ransom to Canvas hackers
Years ago I attended a conference that had a "fireside chat" with a DoJ official on the topic of these types of ransom payments.

He framed the issue as being similar to kidnapping ransoms: When an American is taken hostage each family is inclined to make payment but it fosters an industry around kidnapping Americans. Congress put a stop to it by making it illegal to pay the kidnappers. The industry shifted by ceasing the non-profitable American kidnapping and instead began targeting Europeans.

His proposal was to begin warning cybersecurity consultants and insurers who were often brought into these situations that payments to sanctioned countries were already likely illegal and could face scrutiny. The first people to suffer this might be burned, but eventually he believed the industry would move on and stop targeting US firms.

Not sure if anything ever came of his plans, but I always thought it was an interesting framing of the issue.

jawiggins··on Ask HN: What are you working on? (May 2026)
I'm working on Optio - an AI agent orchestration platform built on Kubernetes: https://github.com/jonwiggins/optio

It's built around multiple different types of agents:

- Coding Agents are placed into cloned repos with a ticket (Jira/Linear/Notion/GH), and work until they open a PR, are resumed on CI failures or github feedback, and work until they can merge the PR.

- Standalone Agents are reusable, parameterized agent runs with no repo checkout. Generate reports, triage alerts, audit dependencies, query a database, post to Slack, etc.

- Persistent Agents are long-lived, named, message-driven agent processes. Each has a stable slug, an inbox, and a cyclic state machine. Wake on user messages, agent messages, webhooks, cron ticks, or ticket events.

jawiggins··on GPT-5.5
What is the major and minor semver meaning for these models? Is each minor release a new fine-tuning with a new subset of example data while the major releases are made from scratch? Or do they even mean anything at this point?
jawiggins··on F-35 is built for the wrong war
> I guarantee you that f35 would go down in a war with a country with decent anti air such as Russia or China

How many F-35s went down due to the Russian and Chinese anti-air systems in Venezuela and Iran?

jawiggins··on Claude Managed Agents
Thanks for the feedback. Earlier I expected I'd need to do more back and forth with the agents before accepting their work but in general I've found it isn't needed.

I do have some features coming up that will improve the ability to converse with the agent as it's running. I'll make a note to add in a plan setting so you can have that run and converse before it gets going.

jawiggins··on Claude Managed Agents
Shameless self promo but, I've been working on Optio specifically for coding, it works by taking any harness you want and tasking it to open Github/lab PRs based on notion/jira/linear tickets, see: https://news.ycombinator.com/item?id=47520220

It works on top of k8s, so you can deploy and run in your own compute cluster. Right now it's focused only on coding tasks but I'm currently working on abstractions so you can similarly orchestrate large runs of any agentic workflow.

Page 1 of 3Next →