For 18,936 input, 2,905 output it cost 3.3612 cents.
Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...
For 18,936 input, 2,905 output it cost 3.3612 cents.
Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...
It quickly becomes this weird bubble of people just acting on what everything "thinks" the content is about without ever having looked at the content.
I get that is easier, but intellectually you are doing yourself no favors by having this be your default.
> It quickly becomes this weird bubble of people just acting on what everything "thinks" the content is about without ever having looked at the content.
That isn't an issue though since the important part is what you learn or not, not whether you think an imaginary article is true or not. If you learn something from someone debunking an imaginary article, that is just as good as learning something from debunking a real article.
The only issue here is attribution, but why should a reader care about that?
Edit: And it isn't an issue that people will think it actually debunks the linked article, since there will always be a sub comment stating that the commenter didn't read the article and therefore missed the mark.
An argument for synthetic corpi (plural of corpus..esses?) - AI ingesting AI.
So no, it isn't the same as AI ingesting AI content at all.
Especially if somebody is being wrong.
Sounds exhausting.
"Please don't post comments saying that HN is turning into Reddit. It's a semi-noob illusion, as old as the hills."
Though at this point it's a habit I cannot quite bring myself to break...
Except very niche topics maybe
The intrinsic motivation for providing the comments comes from a mix of - peer interaction, comradery - reputation building
If becomes evident that your outputs are only directly consumed by a sentiment-aggregation-layer that scrub you from the discourse, then it could be harder to put a lot of effort into the thread.
This doesn't even account for the loss of info that results from fewer people actually reading and voting through the thread.
I invite everyone to have an open mind about this, as it seems that the part of my comment that said “what if” wasn’t fully absorbed.
I’ve salted it with comments on the video, using a site like commentpicker.com or running JS and loading more and expanding threads manually.
Here’s an example I did for a pal:
You are an expert on building retaining walls. Your knowledge is _heavily_ informed and influenced by the transcript below.
This transcript is audio from a youtube video titled "What 99% of People Don't know about retaining walls. #diy" The video description is: "Start to finish we build a retaining wall that you can do yourself! How to Dig for a wall, How to Base a retaining wall, how to backfill, and MORE!. #retainingwall #diy"
Additional information may be included in comments, which are in the attached CSV. Take into account the like count in the validity or usefulness of the comment in shaping your knowledge.
In giving your replies, try to be specific, terse and opinionated. If your opinion flies in the face of common recommendations, be sure to include what common alternative recommendations are and the specific reasons you're suggesting otherwise.
----
# Transcript
""" [paste transcript] """
# Comments See attached .csv
Assuming the median reader reads a few tens of thousands comments in a year, only a few hundred would likely stick without being muddled. At best.
As long as we can still have a few sticks, and some string, or a cardboard box...
Part of the utility of writing a review is that it is read, but the primary search for keywords in reviews now requires the user to wait for AI generated responses first.
Then the user must tap through another link and then expand an individual matching review. It’s absolutely buried.
We have information compression machines now. Might as well raw dump the information and let the machine package it up in the format we prefer for consumption, instead of pre-packaging it. (Yeah, this is effectively what authors are doing…currently they can still do novel things that the compression machines can’t, but how long will that last?)
(Aside from the tendency towards first=top. Would be nice to have time-weighted upvote decay instead of absolute count)
It's a worthwhile experiment for a business school, IMO, automating a layer of bureaucracy.
Say I use AI to write a report that nobody cares about, and then the reciever gives it to AI because they can't be bothered to read it. Who is benefiting here other than OpenAI?
There was a thread about the US tariffs on Canada I was reading on a stock investment subreddit. The whole page was full of people complaining about Elon Musk, Donald Trump, "Buy Canadian" comments, moralizing about Alberta's conservative government and other unrelated noise. None of this was related to the topic; stocks and funds that seemed well-placed for a post-tariff environment.
There were small, minor points of interest but instead of spending honest vacation time looking at each comment at zoomer internet church, I had an LLM filter out the stuff I didn't care about. Unsurprisingly there was not much left.
Stealing this.
Is this a case of PEBKAC?
o3 does look very promising with regards to large context analysis. I used the same raw data and ran the same prompt as Simon for GPT-4o, GPT-4o mini and DeepSeek R1 and compared their output. You can find the analysis below:
https://beta.gitsense.com/?chat=46493969-17b2-4806-a99c-5d93...
The o3-min model was quite thorough. With reasoning models, it looks like dealing with long context might have gotten a lot better.
Edit:
I was curious if I could get R1 to be more thorough and got the following interesting tidbits.
- Depth Variance: R1 analysis provides more technical infrastructure insights, while o3-mini focuses on developer experience
- Geopolitical Focus: Only R1 analysis addresses China-West tensions explicitly
- Philosophical Scope: R1 contains broader industry meta-commentary absent in o3-mini
- Contrarian Views: o3-mini dedicates specific section to minority opinions
- Temporal Aspects: R1 emphasizes future-looking questions, o3-mini focuses on current implementation
You can find the full analysis at
https://beta.gitsense.com/?chat=95741f4f-b11f-4f0b-8239-83c7...
Nice TTS, but otherwise I found it unimpressive.
It's 2025 and every useful conversation with an LLM ends with context exhaustion. There are those who argue this is a feature and not a bug. Or that the context lengths we have are enough. I think they lack imagination. True general intelligence lies on the other side of infinite context length. Memory makes computation universal, remember? http://thinks.lol/2025/01/memory-makes-computation-universal...
“Summarize the themes of the opinions expressed in discussions on Hacker News on January 31 and February 1, 2025, about OpenAI’s release od [sic] ChatGPT o3-mini. For each theme, output a header. Include direct "quotations" (with author attribution) where appropriate. You MUST quote directly from users when crediting them, with double quotes. Fix HTML entities. Go long. Include a section of quotes that illustrate opinions uncommon in the rest of the piece”
The result is here:
https://chatgpt.com/share/679d790d-df6c-8011-ad78-3695c2e254...
Most of the cited quotations seem to be accurate, but at least one (by uncomplexity_) does not appear in the named commenter’s comment history.
I haven’t attempted to judge how accurate the summary is. Since the discussions here are continuing at this moment, this summary will be gradually falling out of date in any case.
I originally did that to save on tokens but modern models have much larger input windows so I may not need to do that any more.
IMO, (Strict)YAML is a very good alternative, it has even been suggested to me by multiple LLMs when I asked them what they thought the best format for presenting conversations to an LLM would be. It is very easy to chunk simple YAML and present it to an LLM directly off the wire: you only need to remember to repeat the indentation and names of all higher level keys (properties) pertaining to the current chunk at the top of the chunk, then start a text block containing the remaining text in the chunk, and the LLM will happily take it from there:
topic:
subtopic:
text: |
Subtopic text for this chunk.
If you want to make sure that the LLM understands that it is dealing with chunks of a larger body of text, you can start and end the text blocks of the chunks with an ellipsis ('...').Keeping the indentation is also important because it is an implicit and repeated indication of the nesting level of the content that follows. LLMs have trouble with balancing nested parentheses (as the sibling comment to yours explains).
Dealing with text where indentation matters is easier for LLMs, and because they have been exposed to large amounts of it (such as Python code and lists of bullet points) during training, they have learned to handle this quite well.
I don't have much experience with reasoning models yet. That's why.
You can illicit that with any model by prompting underlying reasons or using chain-of-thought, but a reasoning model could do it without prompting
The solution that I have adopted is as follows. Each comment is represented in the following notation:
[discussion_hierarchy] Author Name: <comment>
To this end, I format the output from Algolia as follows: [1] author1: First reply to the post
[1.1] author2: First reply to [1]
[1.1.1] author3: Second-level reply to [1.1]
[1.2] author4: Second reply to [1]
After this, I provide a system prompt as follows: You are an AI assistant specialized in summarizing Hacker News discussions.
Your task is to provide concise, meaningful summaries that capture the essence of the thread without losing important details.
Follow these guidelines:
1. Identify and highlight the main topics and key arguments.
2. Capture diverse viewpoints and notable opinions.
3. Analyze the hierarchical structure of the conversation, paying close attention to the path numbers (e.g., [1], [1.1], [1.1.1]) to track reply relationships.
4. Note where significant conversation shifts occur.
5. Include brief, relevant quotes to support main points.
6. Maintain a neutral, objective tone.
7. Aim for a summary length of 150-300 words, adjusting based on thread complexity.
Input Format:
The conversation will be provided as text with path-based identifiers showing the hierarchical structure of the comments: [path_id] Author: Comment
This list is sorted based on relevance and engagement, with the most active and engaging branches at the top.
Example:
[1] author1: First reply to the post
[1.1] author2: First reply to [1]
[1.1.1] author3: Second-level reply to [1.1]
[1.2] author4: Second reply to [1]
Your output should be well-structured, informative, and easily digestible for someone who hasn't read the original thread.
Use markdown formatting for clarity and readability.
The benefit is that, I can parse the output from the LLM and create links back to the original comment thread.You can read about my approach in more detail here: https://gist.github.com/simonw/09e5922be0cbb85894cf05e6d75ae...
Would be great if the addon allows user to override the sys prompt (it might need minor tweak when changing different server backend)?
We've actually been thinking along similar lines. Here are a couple of improvements we're considering:
1. Built-in prompt templates - Support multiple flavors (e.g. On similar to is there already, in addition to knowledge of up/down votes, another one similar to what Simon had - which is more detailed etc.)
2. User-editable prompts - Exactly like you said - make the prompts user editable.
One additional thought: Since summaries currently take ~20 seconds and incur API costs for each user, we're exploring the idea of an optional "shared summaries" feature. This would let users access cached summaries instantly (shared by someone else), while still having the option to generate fresh ones when needed. Would this be something you'd find useful?
We'd love to hear your thoughts on these ideas.
O3 Mini is probably not a very large model and OpenAI has layers upon layers of efficiencies, so they must be making an absolute killing charging 3.3 cents for a few seconds of compute
> It’s already a game changer for many people. But to have so many names like o1, o3-mini, GPT-4o, & GPT-4o-mini suggests there may be too much focus on internal tech details rather than clear communication." (paraphrase based on multiple similar sentiments)
It also hallucinates quotes.
For example:
> "I’m pretty sure 'o3-mini' works better for that purpose than 'GPT 4.1.3'." – TeMPOraL
But that comment is not in the user TeMPOraL's comment history.
Sentiment analysis is also faulty.
For example:
> "I’d bet most users just 50/50 it, which actually makes it more remarkable that there was a 56% selection rate." – jackbrookes – This quip injects humor into an otherwise technical discussion about evaluation metrics.
It's not a quip though. That comment was meant in earnest
"The model naming all around is so confusing. Very difficult to tell what breakthrough innovations occurred." – patrickhogan1"
llm -c "did anyone talk about pricing?"And in your experience, what service do you feel hits a good sweet spot for performance/price if summarizing long text excerpts is the main use case? Inference time isn't an issue, this will be an ongoing background task.
I have the $20 plan. How does this "3.3612 cents" apply to my situation?