HNHacker News
TopNewBestAskShowJobs

KerrickStaley

2,141 karma · joined September 7, 2012

I live in San Francisco and do AI safety research at OpenAI. https://www.kerrickstaley.com/about
submissionscomments
KerrickStaley··on OpenAI's Misalignment Framework: A Tactical Bid to Preempt Global AI Governance
I worked on an early draft of the OpenAI misalignment reporting framework, and my immediate coworkers are the authors behind the first batch of reports that have come out through this process.

The primary reason for putting this process in place was to allow more transparency. There was a sense that the DseWiki incident should have been disclosed, before outside researchers had to disclose it for us.

There was no meta gaming about regulation that I was aware of. I would personally be excited if there were regulation mandating this disclosure process, which allows anyone at the company to raise an issue and shepherd it through the reporting process.

KerrickStaley··on We must pace the frontier
Also endorsed by Sam Altman https://x.com/sama/status/2098811563415150910 and Elon Musk https://x.com/elonmusk/status/2098789109980332057
KerrickStaley··on Ask HN: How do you manage skills files?
Codex and Claude Code both respect ~/.agents/skills; you don't need to have ~/.codex/skills and ~/.claude/skills .
KerrickStaley··on Show HN: Make your Framework 12 sound like a creaky door
I think someone also made this for MacBooks: https://github.com/samhenrigold/LidAngleSensor
KerrickStaley··on The bottleneck might be the air in the room
I don't think this is correct. The concentration of CO2 in air is about 0.04%, whereas the concentration of oxygen is 20%, so the partial pressure of oxygen is about 500x higher. This means that if, for example, 10% of the oxygen in a room spontaneously disappeared, it would be replaced about sqrt(500) = 22x faster through leaks in the room than a 10% spontaneous CO2 increase would dissipate. (This ignores a small effect due to the different density of the two gases).

So in practice the oxygen level can never drift meaningfully far from the atmospheric pressure, whereas carbon dioxide easily can because the pressures involved are so low.

KerrickStaley··on Your ePub Is fine
I use KOreader on my Kobo, which is an alternative open-source ebook renderer. It installs pretty easily without rooting the device. In addition to better standards compliance it also is a lot more performant and has niceties like auto cropping of PDFs.
KerrickStaley··on If you are asking for human attention, demonstrate human effort
A somewhat related experience: I asked for advice on Twitter about something and got two unhelpful AI-generated responses (from accounts I have never heard of / don’t follow) and no human responses. The thing is that I already asked multiple frontier AIs the same question and didn’t get a satisfying answer. I specifically went to Twitter because AI did not have the answers I was looking for. Providing an AI answer to a human question assumes that the asker hasn’t already done their homework and tried asking an AI.
KerrickStaley··on Danish Pension Blacklists SpaceX over 'Catastrophic Governance'
Good point; I didn't realize that GOOG/META/AMZN were outside the scope of VGT. Is there a good alternative to QQQ/VGT that you would recommend?
KerrickStaley··on Danish Pension Blacklists SpaceX over 'Catastrophic Governance'
I think VGT is a good QQQ replacement. It is based [1] on the MSCI US Investable Market Information Technology 25/50 Index which is free-float adjusted [2] [3], meaning that SpaceX will have a lower weight due to its lower free float. Also, VGT has a substantially lower expense ratio (9 bps / year [4]) than QQQ (18 bps / year [5]). You can compare VGT and QQQ's holdings on these pages [6] [7].

[1] https://fund-docs.vanguard.com/F0958.pdf

[2] https://www.msci.com/indexes/documents/methodology/2_MSCI_25...

[3] https://www.msci.com/documents/10199/6bafd9e3-0474-f03b-16bd...

[4] https://investor.vanguard.com/investment-products/etfs/profi...

[5] https://www.invesco.com/qqq-etf/en/about.html

[6] https://stockanalysis.com/etf/vgt/holdings/

[7] https://stockanalysis.com/etf/qqq/holdings/

KerrickStaley··on Open source Kanban desktop app that runs parallel agents on every card
This reminds me of Vibe Kanban (https://vibekanban.com/) which I use to manage coding agents on most of my projects.

The Vibe Kanban developers unfortunately decided that they didn't see a path to profitability and have stopped investing in the project. It's open source and so you can run it locally / fork it, but it has stopped improving and there are still annoying bugs that need to be fixed (and I don't have time to maintain it personally). This makes me sad because I would be willing to pay for Vibe Kanban, but I didn't need the features their paid plan offered (in retrospect maybe I should have paid anyway).

I'll give Kanbots a go :) I'd recommend liberally copying features from Vibe Kanban. In particular the remote support and "Open in VS Code" button (which in my case opens a local VSCode client pointing to a remote VSCode server) are critical for me.

KerrickStaley··on Make ZIP files smaller with ZIP Shrinker
You can also make ZIP files smaller by switching the compression from Deflate to Zstandard. In the one case I tried this, this resulted in a 60% file size decrease [1]. Unfortunately Info-ZIP which provides the unzip command hasn't had a release in 18 years, so it doesn't support this newer compression/decompression method. You have to use 7-Zip instead.

[1] https://github.com/UKGovernmentBEIS/inspect_ai/pull/3145

KerrickStaley··on Show HN: Gaussian Splat of a Strawberry
Here’s a good 2 minute explainer https://youtu.be/HVv_IQKlafQ
KerrickStaley··on Fc, a lossless compressor for floating-point streams
Another library in this space is pcodec; I'd appreciate a comparison of the two.
KerrickStaley··on Cursor Camp
I love this! It reminds me of https://cursordanceparty.com/ which was built by a friend about 15 years ago and is still online :)
KerrickStaley··on "cat readme.txt" is not safe if you use iTerm2
> At the time of writing, the fix has not yet reached stable releases.

Why was this disclosed before the hole was patched in the stable release?

It's only been 18 days since the bug was reported to upstream, which is much shorter than typical vulnerability disclosure deadlines. The upstream commit (https://github.com/gnachman/iTerm2/commit/a9e745993c2e2cbb30...) has way less information than this blog post, so I think releasing this blog post now materially increases the chance that this will be exploited in the wild.

Update: The author was able to develop an exploit by prompting an LLM with just the upstream commit, but I still think this blog post raises the visibility of the vulnerability.

KerrickStaley··on Muse Spark: Scaling towards personal superintelligence
"Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting."

- Hacker News Guidelines https://news.ycombinator.com/newsguidelines.html

KerrickStaley··on Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
I think most people can speak faster than 120 WPM. For example this site says I speak at 343 WPM https://www.typingmaster.com/speech-speed-test/, and I self-measure 222 WPM on dense technical text.
KerrickStaley··on Ollama is now powered by MLX on Apple Silicon in preview
I think (without having done extensive research) that some sort of Apple hardware is your best bet right now. Apple hasn’t raised RAM upgrade prices [1] (although to be fair their RAM upgrades were hugely inflated before the crunch) and their high memory bandwidth means they do inference faster than most consumer GPUs.

I have an M4 MacBook Air with 24 GB RAM and it doesn’t feel sufficient to run a substantial coding model (in addition to all my desktop apps). I’m thinking about upgrading to an M5 MacBook Pro with much more RAM, but I think the capabilities of cloud-hosted models will always run ahead of local models and it might never be that useful to do local inference. In the cloud you can run multiple models in parallel (e.g. to work on different problems in parallel) but locally you only have a fixed amount of memory bandwidth so running multiple model instances in parallel is slower.

[1] https://9to5mac.com/2026/03/03/apple-macbook-price-increase-...

KerrickStaley··on Proton Mail Helped FBI Unmask Anonymous 'Stop Cop City' Protester
https://archive.is/cGvKG
KerrickStaley··on Google Workspace CLI
Tried this out today and it feels half-baked unfortunately. I can't get auth working (https://github.com/googleworkspace/cli/issues/198).

The decision to pass all params as a JSON string to --params makes it unfriendly for humans to experiment with, although Claude Code managed to one-shot the right command for me, so I guess this is fine. This is an intentional design per https://justin.poehnelt.com/posts/rewrite-your-cli-for-ai-ag...

KerrickStaley··on My Favorite 39C3 Talks
Side note, a lot happens at C3 other than the talks! Art, electronic gizmos and demos of all kinds, people hacking in realtime on projects, impromptu meetups, and bumping techno music :) I'd encourage people to attend in person if they get a chance; just watching the talks online is only a fraction of the experience.
KerrickStaley··on Show HN: Vibe Code your 3D Models
I recently designed an eval to see if LLMs can produce usable CAD models: https://kerrickstaley.com/2026/02/22/can-frontier-llms-solve...

Claude 4.6 Opus and Gemini 3.1 Pro can to some degree, although the 3D models they produce are often deficient in some way that my eval didn't capture.

My eval used OpenSCAD simply due to familiarity and not having time to experiment with build123d/CadQuery. There is an academic paper where they were successful at fine-tuning a small VLM to do CadQuery: https://arxiv.org/pdf/2505.14646

KerrickStaley··on Can frontier LLMs solve CAD tasks?
Cool project, thanks for sharing!

The simulator lets the LLM request renders from different angles/times, so the LLM can get visual feedback. For failures, the simulator also returns status codes like `object_fell` or `mount_initially_collided_with_object` depending on what happened. You can see what the tool call looks like by looking at the Transcript tab, e.g. here https://kerrickstaley.com/ai-cad-design-mount-viz/gso__mug__...

I agree it's not clear how much benefit models get from iteration. Many of the successful runs are one-shots. You can see some examples of basic spatial reasoning e.g. here https://kerrickstaley.com/ai-cad-design-mount-viz/gso__mug__... :

> The initial collision is because the mount was positioned at the same height as the mug's body center (z=-22), causing overlap. I need to lower the mount significantly so the mug starts above it and drops into the cradle.

KerrickStaley··on Chained Assignment in Python Bytecode

  a = b = []
has the same semantics here as

  b = []
  a = b
which I don't find surprising.
KerrickStaley··on Suicide Linux (2009)
A fun way to play this game with less downside is to run `set -euo pipefail` in an interactive session. Then, whenever you execute a command that returns a non-zero exit code, your shell will exit immediately.

Unfortunately certain commands like `rg` will return non-zero by design when there are no matches, which could be an intentional outcome.

KerrickStaley··on Gaussian Splatting – A$AP Rocky "Helicopter" music video
This 2-minute video is a great intro to the topic https://www.youtube.com/watch?v=HVv_IQKlafQ

I think this tech has become "production-ready" recently due to a combination of research progress (the seminal paper was published in 2023 https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/) and improvements to differentiable programming libraries (e.g. PyTorch) and GPU hardware.

KerrickStaley··on Claude Code On-the-Go
Does anyone know what the "Poke" service that this blog mentions is? I'm having trouble finding it on Google.
KerrickStaley··on Ask HN: What Are You Working On? (December 2025)
I'm experimenting to see if frontier LLMs can do practical CAD modeling. I'm starting with a single task: designing a wall mount for my bike pump in OpenSCAD or CadQuery (two code-based CAD systems).

None of the frontier LLMs (Gemini, ChatGPT, Claude) produce usable designs when just prompted with some photos of the pump and a written description of the mount. I'm now building a simulator in Mujoco that the LLMs can use to test and iterate on their designs to see if they can do better in this setting.

I'm hoping to make an interesting blog post of it and maybe end up with a usable wall mount design.

KerrickStaley··on Getting a Gemini API key is an exercise in frustration
In my personal experience, OpenRouter makes it easy to call Gemini 3 Pro Preview and other frontier LLMs with very little setup. It’s great for projects where you want to compare different LLMs or have the flexibility to switch. It charges a 5.5% fee on top of the base API price so at scale you would want to switch to directly calling the provider.
KerrickStaley··on I bought a £16 smartwatch just because it used USB-C
This problem seems prevalent on cheaper devices. When I buy a device and discover it has this problem I always return it. I've seen it on the Hypervolt Go 2 (which I returned and replaced with a Theragun Mini) and on the Hitachi Magic Wand Micro (which I replaced with a Dame Dip).

Like the post mentions, I think this happens because the devices are missing two resistors that are needed to indicate, when connected via a USB-C to USB-C cable to a charging brick, that the device wants 5V power. Resistors are cheap and I think the only reason they get dropped is carelessness.

The whole point of USB-C is that you can charge any device with any power supply.

Page 1 of 11Next →