HNHacker News
TopNewBestAskShowJobs

cjbprime

11,924 karma · joined January 3, 2010

Chris Ball. Senior Principal Security Engineer on Zoom's Offensive Security team, posting personal opinions only.

Previously: keybase.io, gittorrent.org at Recurse Center, VP of Engineering at FlightCar (YC W13), Linux kernel SD card subsystem maintainer.

https://printf.net/

[ my public key: https://keybase.io/cjb; my proof: https://keybase.io/cjb/sigs/r1jKbK2XHT3K67jMZN5eZydp5Y3bnwbrv7Eqkm1-wqU ]

submissionscomments
cjbprime··on DeepSeek Harness Desktop for macOS and Windows
It has particularly good observability (the ability to see the full content of every prompt and response and tool call) compared to other harnesses, that's the main thing that stood out to me.
cjbprime··on Astra for Coding: Why Are We Doing This Again?
> I’m more and more convinced that all of AI engineering is Neijuan (内卷, meaning curl inwards). In China it describes a system that demands ever more effort and competition without improving output

I don't know what to say, except that articles exactly like this one have been showing up constantly for the last three years, and literally all of them were obviously outdated and irrelevant within about a month.

cjbprime··on NTSB issues investigative update on B-767 runway excursion accident in Miami
It could be literally automated. And I'm talking about while they were still minutes away on approach, not during final landing.
cjbprime··on NTSB issues investigative update on B-767 runway excursion accident in Miami
And yet I am! Care to elaborate?
cjbprime··on NTSB issues investigative update on B-767 runway excursion accident in Miami
I mean, complete ethical nihilism is one way to try to avoid losing this particular argument, but I don't think you'll get many people to agree with you on it.
cjbprime··on NTSB issues investigative update on B-767 runway excursion accident in Miami
Maybe ATC could be expected to revoke landing clearance on extremely obviously unstable approaches? I wonder why that didn't happen.
cjbprime··on NTSB issues investigative update on B-767 runway excursion accident in Miami
They made a series of at least ten (and that's kind of charitable) this-must-never-happen mistakes that likely rise to the level of extraordinary criminal negligence, including ignoring checklists, alarms, configuring the plane for landing in general, and basically everything related to safety.

You don't have to extend sympathy, just as you don't have to extend sympathy to drunk drivers who kill people.

cjbprime··on AI Responsibility – OpenAI and Anthropic
> LLMs can't read binary directly without a disassembler.

They totally can. They're remarkably competent at disassembly.

cjbprime··on VMs won't contain cyber-capable agents
The tokens spent to find the vulnerabilities will cost money, and the attackers will usually find themselves more financially incentivized to spend money on finding the vulnerabilities than the defenders.
cjbprime··on The Commodore 77, an All-New Cyberpunk 2077 Collaboration
If it was a mechanical keyboard with that case instead of an entire computer they could charge twice as much for it.
cjbprime··on VS Code in the Terminal
Waaait what? That's both crazy and crazy useful.
cjbprime··on Memory prices climb 500% in 12 months
Not sure about that, V100s are all selling for less than $1k despite being 24-32GB VRAM around 1TB/s memory bandwidth, which is the same price point that 4090s and 5090s are commanding $4k-$5k for. Hardly hotcakes.
cjbprime··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
> strong recommendation: ignore that default. Run Qwen 3.8 27B on low or even no reasoning levels at first. It’s a great model, but wow that default setting is a bad place to start.

Has anyone tried asking the model to choose and emit the most appropriate reasoning level for each prompt, as the first part of answering it?

cjbprime··on Qwen 3.8 27B
> General knowledge: I usually ask 2 questions many small models get wrong: summarize Operation Trojan Horse by John Keel, give publication year. Summarize the Ariel school incident of 1994. Qwen3.8 got the first question right, along correct publication year, gave glorious detail on Keel's theory, but got the second wrong. It thought that school was located in the USA. Ah well.

I'm not sure that this means anything. You're asking a ~27GB file to have losslessly compressed the entire training set (which apparently is a large chunk of the entire internet). That's not possible. Whether it happened to encode these particularly obscure facts losslessly or vaguely isn't really telling you anything about how good a model it is.

cjbprime··on GLM-5.3: Frontier coding with emergent cyber capabilities
Why not? The web uses TLS, how's it different security-wise compared to a package download?
cjbprime··on Qwen 3.8 27B
Does anyone know how to get this working with Claude Code via llama-server? I'm getting a jinja template error about the system prompt not being the first message.
cjbprime··on Qwen 3.8 27B
Hm, I have a 4090 as well, and:

$ build/bin/llama-server -m Qwen3.8-27B-IQ4_NL.gguf --mmproj mmproj-BF16.gguf -c 170000 --parallel 1 -ngl -1 --cache-type-k q8_0 --cache-type-v q8_0 -b 1024 -ub 512 --flash-attn on --no-context-shift --no-mmproj-offload --spec-type draft-mtp --spec-draft-n-max 5 --spec-default --cache-type-k-draft q4_0 --cache-type-v-draft q4_0 --threads 24 --jinja --reasoning on -fit off

0.02.993.689 E ggml_backend_cuda_buffer_type_alloc_buffer: allocating 911.53 MiB on device 0: cudaMalloc failed: out of memory

Update: Oh, it works after I stop Xorg. But nvidia-smi only showed Xorg using 200M out of the 24G, so why would a 911M alloc fail?

cjbprime··on Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Does it not use Metal, on macOS? Would it be faster if it did?
cjbprime··on Show HN: Getting GLM 5.2 running on my slow computer
It's a very conservative warning. The application does not perform writes, so the application doesn't actually wear your SSD at all. The rest is just application-independent general hygiene.
cjbprime··on GLM-5.2 – How to Run Locally
You still have a core misunderstanding. Only one layer of weights is required in memory at a time. A forward pass can be over-simplified as a matrix multiplication of each layer, one at a time.

There is no swapping of working RAM. We're just talking about loading the weights read-only data into RAM on-demand for each layer. It is only as slow as your storage interface.

cjbprime··on GLM-5.2 – How to Run Locally
Both sentences are likely wrong. It's a written lifespan (technically an "erased" one), not a read lifespan. The weights are only needed from disk read-only. And Mac NVMe interfaces are surprisingly fast.

Edit: Oh, I think you maybe thought I meant swapping working RAM off disk? I didn't. I meant swapping weights off disk into RAM on-demand.

cjbprime··on GLM-5.2 – How to Run Locally
I've got access to a 192GB RAM Mac Studio, which is below the stated minimum RAM. Can swapping off fast disk be used to make it work out, especially since it's MoE?
cjbprime··on Ask HN: What was your "oh shit" moment with GenAI?
ChatGPT reconstructing idiomatic Python source code from Python bytecode was definitely up there. That is not something humans have written a great deal about online. It requires simulating the Python VM.

I remember also having a massive wtf reaction to realizing that original ChatGPT was pretty good at decoding long random/unique base64 strings.

cjbprime··on Guitar tuner that uses phone accelerometer
> If it detected any harmonics it would be too high.

I think it's not that simple. A tuner is "hearing" the fundamental and all of the harmonic overtones combined. It has to guess at which frequency is the fundamental, even if the overtones are actually stronger than it amplitude-wise, and it does that by looking at the nature of the repeating overtone pattern and extrapolating back to the fundamental.

I think you can end up an octave too low (half the actual frequency) if the waveform repeats in a way that implies a different overtone repetition pattern, for example if there's an every-other-cycle artifact to the waveform.

cjbprime··on We got 207 tok/s with Qwen3.5-27B on an RTX 3090
Inference (not training) is bottlenecked by memory access speed, not compute. Having special hardware wouldn't make it faster unless you somehow found a faster memory controller than the GPU has.
cjbprime··on Preliminary report into Air India crash released
I don't agree with the "twice". A frequently performed manipulation like the fuel cutoff (usually performed after landing) collapses down to a single intention that is carried out by muscle memory, not two consciously selected actions.
cjbprime··on Preliminary report into Air India crash released
The prelim report states these pilots were indeed breathalyzed before takeoff.
cjbprime··on Preliminary report into Air India crash released
I meant philosophical toggle switches, not physical ones. The gear can go between down and up. The fuel can go between run and cutoff. Given enough practice, the brain takes care of the physical actions that manipulate those philosophical toggles without conscious thought about performing them.
cjbprime··on Preliminary report into Air India crash released
Now I'm trying to remember if I've ever picked up my razor and accidentally begun tooth brushing motions with it. Probably!

More relevantly, you seem to me to be unduly confident about what this pilot's associative triggers might and might not be.

cjbprime··on Preliminary report into Air India crash released
Sometimes? If you have enough altitude to trade for speed then after the cutoff you could glide to a hypothetical miraculously-placed runway right in front of you, vs. having fire quickly consume the entire plane if you don't cutoff..
Page 1 of 34Next →