LLM-hacker-news: LLM plugin for pulling content from Hacker News
github.com
github.com
Normally fragments are specified using filename or URLs:
llm -f https://simonwillison.net/robots.txt "explain this policy"
Or: llm -f setup.py "convert to pyproject.toml" -m claude-3.7-sonnet
I also added a plugin hook that lets you do this: llm install llm-hacker-news
llm -f hn:43615912 -s 'summary with illustrative direct quotes'
Here the plugin acts on that hn: prefix and fetches data from the Hacker News API, then applies the specified system prompt against LLM's default model (gpt-4o-mini, unless you configure a different default).I wrote more about the Hacker News plugin here: https://simonwillison.net/2025/Apr/8/llm-hacker-news/
It uses the Algolia JSON API https://hn.algolia.com/api/v1/items/43615912 and then converts that into a (hopefully) more LLM-friendly text format.
Another neat fragments plugin is this one, which grabs a full clone of the specified GitHub repository and dumps every non-binary file in as a fragment at once: https://github.com/simonw/llm-fragments-github
Example usage:
llm install llm-fragments-github
llm -f github:simonw/files-to-prompt 'suggest new features for this tool' llm -f https://news.ycombinator.com/reply?id=43621396&goto=item%3Fid%3D43620125%2343621396 "summarize the comment"Let me know what I should improve.
---
Compiling the Linux kernel to get more stability it's done with the '-O3 -ffast-math -fno-strict-overflow' CFLAGS.
Run your window manager with 'exec nice -19 icewm-session' at ~/.xinitrc to get amazing speeds.
What's your concern here - is it not wanting LLMs to train future models on your content, or a more general dislike of the technology as a whole?
The "not train on my content" thing is unfortunately complicated. OpenAI and Anthropic don't train on content sent to their APIs but some other providers do under certain circumstances - Gemini in particular use data sent to their free tier "to improve our products" but not data sent to their paid tiers.
This has the weird result that it's rude to copy and paste other people's content into some LLMs but not others!
I've not seen anyone explicitly say "please don't share my content with LLMs that train on their input" because almost nobody will have the LLM literacy to follow that instruction!
The same way Google and others have been crawling and capturing all your public posts for decades to power their search engines. Now the data is being used to power LLMs.
Were you able to opt out of being part of the search index (and I don't mean at the site level with a robots.txt file)?
I think your choice here is "don't post on a publicly accessible website", unfortunately.
If you're in the EU, then yes, as "Right to be Forgotten" is a thing: https://en.wikipedia.org/wiki/Right_to_be_forgotten#European...
But in general I agree, the expectation of something remaining "private" and "owned by you" after you publish it on the public internet, should be just about zero. Don't publish stuff you don't want others to read/store/redistribute/archive.
Others have said it already but when you are posting here on a public website, I would argue that you are effectively consenting that your content is now available for consumption by site visitors.
"Site visitors" may include people, systems, software, etc..
I think it would be pretty impractical for every visitor to the site to have to seek consent from each poster before making use of the content. That would literally break the Internet.
> Except as expressly authorized by Y Combinator, you agree not to modify, copy, frame, scrape, [...] or create derivative works based on the Site or the Site Content, in whole or in part, [...]. In connection with your use of the Site you will not engage in or use any data mining, robots, scraping or similar data gathering or extraction methods
Though I guess this is a tool to produce such content, rather than the author doing this themselves, its ok?
This whole thing is a Pandora's box. We can regulate, forbid, anything, but we all already have models downloaded locally (you did too, right..?). So unless there's some client-side "computer says no" we will never be able to block this anymore.
The local models were mostly unusably weak until about six months ago when they suddenly got useful: Qwen Coder 2.5, Llama 3.3 70B, Mistral Small 3 and Gemma 3 have all really impressed me on a 64GB Mac and I expect Mistral Small 3 would work in 32GB.
Meanwhile this years's Gemini Pro 2.5, Claude 3.7 Sonnet and the most recent GPT-4o API models (or o3-mini high for coding) are significantly better than what we were using last year.
I'm fluent in Python and JavaScript, but these days I'm using LLMs to help me get started writing code in AppleScript, Bash, Go, jq, ffmpeg (that command-line interface is practically a programming language just on its own) and more. I'm learning a ton along the way - previously I wouldn't have been able to get up the energy (or time) to climb the initial learning curve for all of those.
Presumably you mean this code here? https://github.com/simonw/llm-hacker-news/blob/e945c84e825f4...
I reviewed it. For this particular application (turning JSON from https://hn.algolia.com/api/v1/items/43615912 into a more concise format suitable for feeding into an LLM like https://github.com/simonw/llm-hacker-news/blob/e945c84e825f4...) it's perfectly adequate (or "good enough" to quote your comment). You're welcome to convince me otherwise!
On the contrary, it seems there's a lot of people on the Internet who think copyright means something different than it actually does, and therefore justifies them their Dog in the Manger attitude.
One of the prompts was: "make the comments even shorter, and have everyone involved be a pelican (a bit subtle though)"
See also my notes here: https://simonwillison.net/2025/Apr/8/llm-hacker-news/
I mean, to get something like https://hackernewsletter.com/, but personalized for my tastes and interests.
for result in results: fetch_content |> send_to_openai