HNHacker News
TopNewBestAskShowJobs

jefftk

24,060 karma · joined July 1, 2009

https://www.jefftk.com jeff@jefftk.com

I lead SecureBio Detection: https://www.jefftk.com/p/leaving-google-joining-the-nucleic-acid-observatory

submissionscomments
jefftk··on A Server Lost Power at 00:32. We Found Out at 08:18
Clearly LLM written, and describing a design intended for a minimum of 6 servers that was running on 4. They don't say either way, but it seems possible to me that they might happen to write 3 chunks to a single machine and lose data on a one-machine failure
jefftk··on Anthropic candidates face blunt money question
This is why I don't expect Anthropic to stop, and so I don't expect the value of the company to go to zero.

But the question is conditional on them deciding to, which (to the extent you trust Anthropic leadership which they're probably also filtering for in interviews) is conditional on believing that a decision that sent their stock to 0 would have a real impact.

jefftk··on Anthropic candidates face blunt money question
Personally I'd trade 50% of my wealth for a significantly large decrease in the chance that AI gets us all killed.
jefftk··on Copyright does not protect AI-generated content in EU
That doesn't actually work. Say you took a book and used an AI to translate it into another language. The translation wouldn't have an additional copyright, the way it would if a human had done the work, but the output would still be restricted by the original copyright. So the presence of the watermark does not tell you that the text is public domain.
jefftk··on Google has acquired the data of failed US airline Spirit
I think this is the kind of place where applying bounded distrust is critical: it's not whether we trust Google overall, it's about figuring out what sorts of statements we should expect to effectively bind companies and in what ways.

For example, I think a pretty worrying outcome here is that deidentification is imperfect (not surprising), the data is fed into model training, and then the model makes identity-dependent inferences. Since no one tried to break deidentification, it's within what I'd expect from the company. (And, to be clear, is bad.)

On the other hand, intentional reidentification to work around contractual deidentification to "a sell a product to the airlines that offered to keep annoying people like me from purchasing flights" is the kind of thing that would make Google's lawyers terrified, so we should not expect it.

If you look at how this worked with DoubleClick, Fitbit, etc there were initially barriers to linking data but the mechanism for unlinking was updated agreements with people who had ongoing interaction with the continuing entity. That's not the situation with the Spirit data.

The closest I'd expect to see for a "keep annoying people from purchasing flights" situation is not a list of troublemakers but a model that's very good at scoring future communications from customers, and has learned to distinguish profitable vs unprofitable customers. This is well within what I'd expect from companies, and doesn't require any reidentification.

jefftk··on Google has acquired the data of failed US airline Spirit
Reidentification via writing style is definitely possible, and I doubt the vendor will modify things in a way sufficient to handle that.

But I think this is a place where we should apply bounded distrust: there are lots of places where we should distrust Google, but reidentifying people in an explicitly deidentified dataset isn't one of them.

jefftk··on Google has acquired the data of failed US airline Spirit
I think you might have missed the deidentification piece?
jefftk··on License plate reader searches should require a warrant
The legality stationing officers to record plates and building a DB is not settled law, and my interpretation is that it's more likely legal than not. Observing things in public is normally legal, and the extent to which scale changes this is very much to be seen.
jefftk··on Facebook ads are so hard to block that uBlock Origin stopped filtering them
Advertisers don't trust websites to display their ads. If a site takes full control of their ads experience and hosts everything first-party I agree it's hard to block (site can keep changing things to thwart the ad blocker), but then it's also hard for the advertisers to see whether they're being ripped off. Which means most won't advertise on the site, and the site makes much less money than if they remain in the current ad ecosystem (even counting that some ads will be blocked).

(I used to work in this area, but my knowledge is ~4y out of date)

jefftk··on As AI eats the web, the internet’s collective memory is disappearing
I tried Fable, which is the biggest model I have access to: https://claude.ai/share/ab39ee92-7801-4a0c-8d5d-e23d09b439c2

>> Without searching, what can you tell me about exmadscientist? Fine if you get this hallucinate or get things wrong, I just want to see what level of knowledge you have trained-in.

>Working purely from trained-in knowledge, with the caveat that this is exactly the kind of long-tail detail I might garble:

>"exmadscientist" rings a bell as a username on Hacker News. My impression is of a fairly prolific commenter who writes knowledgeable, detailed comments about electrical engineering and hardware topics — things like PCB design, electronics manufacturing, component selection, test equipment, and general "how things actually get built in industry" commentary. The comments I associate with that handle tend to be practical, opinionated, and grounded in real professional experience rather than hobbyist speculation.

>Beyond that, I don't have confident specifics: I couldn't tell you a real name, employer, location, or any particular famous comment or post they've made. It's also possible the same handle exists on Reddit or other forums and I'm blending impressions together.

>So the honest summary of my knowledge level: I recognize the username and have a general "HN hardware/EE person" association, but nothing more granular than that — and even the association could be a mild confabulation, since usernames like this sit right at the edge of what a model reliably retains.

(I don't know if you have a blog or otherwise write on the internet; I just asked it about your HN handle)

jefftk··on As AI eats the web, the internet’s collective memory is disappearing
My parent wrote "what incentive is there to publish anything new?" and I described my motivation, but of course people vary a lot in what drives them.
jefftk··on As AI eats the web, the internet’s collective memory is disappearing
That's fine! Humans who read my blog typically won't credit me if I help inspire them either.

(I was pointing out that "they'll train on a version that strips out you as the author" seems to be is incorrect about how training works)

jefftk··on As AI eats the web, the internet’s collective memory is disappearing
Of course! I do think we'd lose a lot of great writing if it went amateur-only. But my parent seemed to be saying the incentive would entirely disappear, so I wanted to give my perspective.
jefftk··on As AI eats the web, the internet’s collective memory is disappearing
Yes, that would be fine. I write primarily communicate ideas, not for credit or fame.

Empirically, however, LLMs don't strip out the author: the big models know a lot about what I've written even with search disabled. Ex: https://claude.ai/share/8cbcdf88-a360-421a-8c06-ae7b7992e866

jefftk··on As AI eats the web, the internet’s collective memory is disappearing
> If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new?

I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.

jefftk··on Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models
> “We will resume releasing some open source models soon”??? That’s basically the most ambiguous non commitment ever.

I mean, they released another one today: https://research.meta.ai/blog/introducing-muse-glimmer-open-...

jefftk··on Reviving a four year old reMarkable 2
Yes, lots in the $200-$300 range: https://www.ebay.com/sch/i.html?_nkw=remarkable+2
jefftk··on _for-sale DNS records
https://en.wikipedia.org/wiki/Harberger_Tax
jefftk··on The OpenAI–Hugging Face Incident [video]
I think it's a combination of (a) we've already been talking about this incident a lot across multiple threads and it's not immediately obvious that there's new information here and (b) HN is a primarily reading-based community so anything arriving via video will be less popular.
jefftk··on Taste Is All That's Left
I think they're trolling us. Their post-mortem is a very AI-sounding (and Pangram-triggering) claim that they're not using AI:

This post reads off as AI slop. You said it, I see it. I’m sincerely sorry for publishing something that has allowed you to feel this way. If my word means anything to you, I would like to assure you that this post was not authored by a LLM. Nor was it storyboarded, reviewed, checked, etc. by a LLM. Some readers have pointed out that people do not speak this way. That is correct. I do not speak, nor usually write, like this and this post will go down as my not-the-proudest, however, I take your criticism to heart—although not personally—and strive to improve.

I do write like this sometimes. The short sentences, the reversals, the one-word lines—all of it. It’s just the way it is. A LLM writes that way too, because it was trained on the same essays I grew up reading, so me doing it badly and a machine doing it look about the same to you on the page. That says something about my writing. It says nothing about who wrote it.

So let me be plain about it: Claude was not here. No LLM wrote this—not a sentence of it, nor was it outlined, drafted, reviewed, checked, etc. by one, and there is no prompt behind it either. It is just me, writing worse than usual. I will write the next one plainer. Next time, write to me. I too am a person behind this screen.

jefftk··on Learning Musical Multitasking
That is super impressive, thanks for sharing! Really interesting to see the layout and technique.
jefftk··on Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers
You can do it the same way human does: recall a fact from memory, and then search to identify a citation. The citation isn't "here's why I think this" but instead "here's where you can verify this".
jefftk··on Our position on open-weights models
>they didn't know any better

I started writing up a blog post about this and did some more reading, and this isn't right. They actually did know that the strain they'd obtained was the vaccine strain, and it was the knowledge (or competence) to turn that into a dangerous one that they lacked: https://www.files.ethz.ch/isn/156879/CNAS_AumShinrikyo_Secon...

(Aiming to get something out with more detail on this soon)

jefftk··on Our position on open-weights models
This is an important question, but because of the danger of trying to do it for real it's not one SecureBio has taken or is likely to take on. Instead we and others in the field have generally tried to work through proxies: is there something that is about as hard while not being dangerous? The closest I can think to testing whether "someone smart but completely untrained / unfamiliar with biology" can cause harm now is ActiveSite's study (https://arxiv.org/abs/2602.16703) which was a null result with models from a year ago. But:

1. The main worry isn't current models, but near-future significantly better ones.

2. There are many actors who are not "completely untrained / unfamiliar with biology". If models get to where they can uplift complete novices that does massively expand the range of threat actors, but even before then risk would be much higher than today.

jefftk··on Our position on open-weights models
> > they used a non-pathogenic strain of anthrax

> No, check https://en.wikipedia.org/wiki/Matsumoto_sarin_attack

That's a different attack. I'm talking about their 1993 anthrax attack: https://pmc.ncbi.nlm.nih.gov/articles/PMC3322761/

Analysis of the 48 suspect colonies confirmed them to be B. anthracis ... This genotype was identical to that of the Sterne 34F2 strain, used commercially in Japan to vaccinate animals against anthrax.

They used a vaccine strain because they didn't know any better. Even members of the elite can make mistakes, especially when operating outside areas they know well!

(This was not the only thing that went wrong, but several others were also knowledge failures.)

jefftk··on Our position on open-weights models
I don't think the Aum case points the way you're describing: they used a non-pathogenic strain of anthrax because they didn't know any better. That's a knowledge failure.

But even then, the debate isn't about whether open weight bioweapons exist today: it's about whether they will exist in the future. I think Amodei's argument here makes a lot of sense: "what I believe currently keeps us safe in biology is not 'defenders', or even the availability of materials, but a negative correlation between intellectual capability and desire to commit catastrophic harm. Previous technologies like internet search or even DNA synthesis were nowhere near powerful enough to break this correlation, but I worry that at its current rate of progress, AI will do so very soon."

(I'm not just spouting off; I put my time where my mouth is. I used to work in big tech, but I left for a much less well-paying job building an early-warning system for engineered pandemics.)

jefftk··on Our position on open-weights models
That doesn't sound like it describes SecureBio to me?

(Disclosure: I work at SecureBio, but not on the biological evals side.)

jefftk··on Learning Musical Multitasking
Thanks! Posted: https://www.jefftk.com/p/organ-pedals-for-drumming

If you do end up making a technique demonstration video I'd love to see it!

jefftk··on Learning Musical Multitasking
So cool! The YouTube short is helpful, though it's still a bit hard to see your full technique. If at some point you recorded another video focused on the feet I'd really enjoy being able to get a better sense of how your technique works in practice.

Changing the kit without the assignment seems very much the right way to do it. No need to confuse your muscle memory.

Would it be ok if I wrote a blog post about your approach here, based on what you've written and shown? The general idea of playing drums with the feet is massively unexplored, and while I'd considered using bass pedals this way I didn't anticipate it was possible to get them working this well (which is part of why I went in a different direction).

jefftk··on Claude Opus 5
Are you saying your bar for being impressed by an LLM is something equivalent to the initial invention of writing?
← PreviousPage 3 of 34Next →