Mapping Hacker News to find who knows what in the HN community
blog.wilsonl.in
blog.wilsonl.in
Also, I feel like this tool selects for active commenters, not for knowledgeable experts. Not to mention throwaway accounts.
Still a cool project.
HN does not let you do that though. At least last time I asked about it, they sent me a response saying that they’d notify me when it became possible to delete an account. And they provided some reasons for why they don’t delete comments and accounts.
For me I would prioritize the wishes of the contributors.
Also my wish as a contributer is that the threads stay like this. So I can come back later and reread them. I often gained value in reading a linked old thread.
My advice for people not wishing for that, would be simply to stop commenting instead of demanding the site should change.
I know this means someone could still use stylometry to try to reaggregate the posts, but that's less reliable than HN actually telling us that the same authenticated user made all these posts over a decade.
-Twilight Zone Episode: Time Enough at Last-
If you want to add value and not bloody public spectacle rank comments instead of users.
I have a bunch of low quality posts here when idiots piss me off, but also share world first research and breakthroughs I've been involved with the rare time the counter party is worth talking to.
> I have a bunch of low quality posts here when idiots piss me off
It’s ironic because in the second of the parts I quoted you on here you are basically yourself generalizing the users instead of the comments.
Is it really “idiots” that piss you off (the users)? Or is it the specific things they said in isolated cases (the comments) that piss you off? Wasn’t your point exactly that this kind of distinction is important?
We're thinking we could uplevel the social layer
head => palmCurrently, HN is the only place on the internet I am willing to interact with others _because_ it lacks the "social layer" you are recommending.
The focus on user comments that are thoughtful, relevant, and respectful _is_ the social connection I value.
But at the same time I think the OP has made an awesome project! Well done
Not the negative space, but a list of untolerated content.
I have absolutely no idea who I've had an argument with on HN ever. I'm sure I've had a few.
This is basically 99% of hacker news comments in a nutshell.
Apart from sheer volume (I have been a message board nerd since FIDONet when I was a teenager, it's second nature for me) I attribute what success I've had here to engaging principally on security topics, which is where I've spent my whole career. But I also just like to shoot the shit about stuff! I'm wrong all the time! (Including about security stuff).
A fan? A stalker? Just a rando with too much time on their hands?
Obviously this applies to a negligible percentage of total commenters, but as I only comment on certain topics I’m more likely to encounter friends with the same interests/experiences.
And sure, ideally you wouldn't need such a list of trusted experts but just focus on content. There even was a time when this worked - you could just type "what is the best database to use" into a search engine, and get a helpful result. Not anymore. On HN it may still be better than elsewhere, but ultimately it's a similar issue.
If you zoom and pan the red cursor over that map location/text, your source comment will appear underneath the map.
Kind of like how when the pandemic started, sales of guitars and home gym equipment went up as people worked on at-home hobbies.
Weirdly, almost no mentions of Apple stuff (which is supremely ironic given my blog).
So you get an impression of what people discuss here, but not necessarily what they know :)
(I wonder if the embedding was generated using IDF, placing maybe too much emphasis on rare word tuples?)
Mine has a ton of stuff:
Show HN: Exploring HN by mapping and analyzing 40M posts and comments for fun - https://news.ycombinator.com/item?id=40307519 - May 2024 (159 comments)
Is there an opt-out of this thing?
Could this tool be repurposed for that? Presumably the “map” rendered in each user’s avatar could be encoded as a vector and then compared to that of another user.
EDIT: Wait, I just realized it already does this… (or at least I think so - it’s not immediately obvious if “Explore More Users” is ranked by similarity.)
It's probably less likely to work because people often use alts to participate in discussions that they wouldn't want to associate the primary identity with. Whether that be discussions about their employer, about politics, or something else, the subjects an alt particulates in will likely be different.
Style works because few people are capable of fundamentally altering their style even if they tried.
nemotron-4-340b: "One might find it intriguing to observe if you have the capability to disguise your statements by utilizing a common LLM as a filter. It's plausible to believe that you could."
gemma-2-27b got verbose on us: Here are a few ways to rewrite the sentence "I would be amused to see whether or not you could effectively disguise your comments by utilizing a common LLM" to reflect different styles:
*More formal:*
* "It is hypothesized that a commodity LLM could be employed to effectively obfuscate your comments." * "Utilizing a commodity LLM for comment anonymization is a plausible strategy."
*More casual:*
* "I wonder if you could fool a regular LLM with your comments?" * "Could a basic LLM make your comments look like they came from someone else?"
*More skeptical:*
* "While it's possible, it's unlikely a commodity LLM could effectively mask your comments to the point of complete anonymity." * "Are you sure a simple LLM can truly hide your identity? It might not be enough to protect you."
*More direct:*
* "Using a commodity LLM for comment anonymization is a viable option." * "Can commodity LLMs effectively anonymize your comments? Probably."
The best way to rewrite the sentence depends on the context and the tone you want to achieve.
----
All of the above courtesy of perplexity.ai's playground.
Ironically, it doesn’t use embeddings even though at the time and since I’ve basically exclusively used embeddings.
Main issue with embeddings is they don’t change well overtime (you can adjust, but not as easily as other methods). Language is filled with industry specific acronyms and models are generally not great at adapting to changing phrasing and acronyms.
Anyway, eventually decided to put some more pieces together and built it into a company - https://ipcopilot.ai
Now we discover ideas on hacker news for our demo https://m.youtube.com/watch?v=B5ymgh-ZDiI&pp=ygUKSXAgY29waWx...
Like user123 is moderately funny, articulate, interested in typescript and web development, and rated high on conscientiousness and extraversion.
"Phædrus was a master with this knife, and used it with dexterity and a sense of power. With a single stroke of analytic thought he split the whole world into parts of his own choosing, split the parts and split the fragments of the parts, finer and finer and finer until he had reduced it to what he wanted it to be. Even the special use of the terms "classic" and "romantic" are examples of his knifemanship."
In a bit of nominative determinism, or perhaps just having chosen the name because I know myself (or maybe I just over use these words), my keywords include: "part, system, level, language, article, object," etc.
What I'm saying is: I like the focus on the content and that it's not about who said it.
It got me to remove my twitter handle from my bio though. If you could update that in your app I would be thankful.
I have a google chrome extension that lights up when there are HN comments for the page I'm on. If I organically discover a blog post and that lights up, then I will often read the comments even though it might have been posted years ago. There is often useful insight or context to be gained.
But I guess I'm still confused. If a vendor provides software that doesn't honor the right, how many times can they do that before they get in trouble?
In HN case, I'm sure an email would sort it out, however.
Information flows freely worldwide. Regulations do not.
If your username is your full legal name and city, I would say that's your own stupid fault (and I live in a GDPR country).
I know how you feel.
I felt that way, too.
But what I found was that 2 clicks away from any comment is a query like: https://news.ycombinator.com/threads?id=RamblingCTO
Hacker News itself has for many years provided a free near-realtime API and a variety of large datasets to the world for the specific purpose of making it easy for anyone to analyze our comments here. (https://github.com/HackerNews/API)
No comments you write on the Internet have ever been temporary. They stopped even pretending to be so 20+ years ago.
Rather than organizing the world's information, what if we could organize the world's people?
I do not wish to be organized.I suspect a bias towards more common topics might be occurring.
Keywords the site actually associates with me: language, English, article, team, book. To be fair, at some zoom levels I do get "chess".
https://hn2.wilsonl.in/user/dang
Why does it say "Marion Milner" and why are there only so few posts on the map?
https://hn2.wilsonl.in/user/dep_b
I guess that's why I still like posting anonymously?
I think this is an extremely cool idea both on HN specifically and generally on the Internet. Bluesky does a bit of the thing where you can mix and match your content to your ranker/recommender system.
I hope you folks keep working on it, this is a refreshingly cool hack in the space.
Also I imagine that my area is pretty niche (not many people talking about documentation on HN) so I may have less competition
Fallacy: appeal to authority. Practically, just because someone generates great content on subject A doesn't mean their take on subject B is any better than random. A well-reasoned self-consistent argument informed by accurate data is far more valuble than 'trust this expert opinion because this expert can be trusted' approaches - although it may require more work on the part of the reader. Don't get lazy.
Awesome project, fantastic UI.
I find myself dragged to Apple discusions more than I would like to admit, but they are generally negative in the last years.
I wonder if you could make a video of how an individual could use it effectively in a pragmatic sense e.g. networking, research
(after more playing around I get the UX, very powerful indeed)
(feedback: add some controls for filtering (dates etc), relevancy, color islands slightly, just general UX improvements and I would probably pay a few dollars a month for this)
(feedback: when looking at a user, you have to zoom really far in before results start appearing. it should probably show labels/results if less than X amount of data points are in your viewport.)
(feedback: Is there "Explore more users" section the more related users to the current profile? It would seem so, but not immediately clear. And if it is, none of my related users have me as a related user in their profile.
===
This is really inspiring work, hats off!
Could you share some info about how you generated the 2D space embedding/visualization?
How can that map be used to determine who is knowledgeable about what? Looking up myself, I can't connect what I see with my own areas of knowledge.
I think I need an ELI5 for this. Again, this isn't a criticism of the effort at all, it's a public admission of my own ignorance.
Then you encode a search query the same way and return the nearest N vectors/users. That's vector search.
It seems they did the above for HN comments.
We'd love to hear your feedback as you play with it more and we improve all of the above.
To be fair.. you can’t. Unless you think that mentioning specific words or phrases associated with a specific topic a high number of times means that you’re knowledgeable at it.
I'll admit, I am a generalist, and I comment more than I post, but I would have thought/hoped my knowledge base would trend towards neurotech, sleep, mental health, health, wellness, etc, as I feel that is the "less general" stuff I comment on.
Though I recognize this is how I like to see myself, I do often comment on business models, branding, etc etc, which are areas I have interest in.
I wonder if topics that have more activity from a user, but are in a less popular topic area would improve this?
I'm not sure how you surface the more nuanced understanding of each users knowledge or interests.
The few things I'm thinking are 1) more weight to more recent content than earlier. We change and grow as people. I used to be in the music industry, then 3d mapping, and now neurotech. My music experience is now old, but neuro is fresh. That may be more interesting/valuable.
2) commonality with the viewing user. Where do we overlap, and again, where do we overlap where others don't? How often would a person be searching for someone whom they don't have overlap with? We are likely interested in the same things. Though perhaps somebody who knows nothing about marketing, is looking for marketing help, so I'm not sure how you surface that.
I guess I'm wondering what you imagine the ultimate use case being? Or the early adopter use case. You may be prioritizing the algorithm for uses I'm not expecting.
It would be nice if map showed titles of all articles on extreme close up.
Overall it's interesting tool but I think for people with small amount of comments this is not much informative.
Cool visualization and analysis, really well made!
Email random users (control group) the same
Ask after a month who's still in touch
I'd participate but please send me some of each group and don't tell me which is which :D
Also, thanks a lot for attempting to put it in decoded form on another website. If I now get spam on that address, at least I know whose idea that was and that it's not yet spammers who got this clever, but rather it's due to well-meaning hackers with an idea: the most dangerous kind! :P
I guess having over a decade in startups counts for something, but it's crazy to score higher than sama or all the other founder/investors here who make 1000000x more from startups.
My favourite terms here are apparently "inclusive pregnancy emoji proposal", "Controversial tweet on China issue", and "enhancing pasta flavor with salt". I do remember all of these conversations, it's interesting to see it brought up.
Anyway, I'm gonna bookmark this for the next time I look for a job on HN lol
Checked RIP Aaron Swartz. Nothing meaningful. What am I missing?
Now if this tool could distinguish posts that sound like they were written by a layman from those made by an expert...
I checked my profile earlier and there was none, though my HN profile has scads of info. That's why I was curious.
Totally fair to say "Word limit!" or something, but it means your reply doesn't seem to jibe with my observation.
Oh, would some Power the gift give us / To see ourselves as others see us!
Looks fantastic. Very cool project.
An example would be that users place a small opt-out snipped within their user profile, e.g. [no-robots] ?
Perhaps OC might consider never commenting outside of 4ch..? Except even glowies train LLMs =D
1. I want the knowledge/experience/anecdotes/opinions of others. 2. I want to share my own knowledge/experience/anecdotes/opinions with others. 3. Contributing low-bandwidth, high-quality* content is (hopefully) doing my small part in encouraging others to do the same.
My primary concern: You are providing a map of individual user interests over a period of time. This is feels like a violation of privacy to some and is outright dangerous to others.
Why? 1. People have many interests. The Marketing Manager likes birdwatching, Bollywood, hacking podcasts, amateur chemistry, traditional Ukrainian breadmaking, ASMR, erotic independent film, rock climbing, pole-dancing, gunsmithing, and horse racing. 2. People are public with some interests, and private with others. The degree of public vs. private disclosure can come from external pressure (ie: political oppression, public servant expectations) or internal motivation (ie: embarrassment, pride). 3. People change over time. The 27 year old Marketing Manager may want to distance themselves from their poor political comments as a student.
Where to now? If you stop work now, someone else will pick up where you left off. So how do make a system which balances the user AND allows you to continue?
1. Easy and simple opt-out option. 2. User profiles. Allow me to share /which/ interests are associated with me. I am a nuclear engineer by trade, but I'll be damned if I want to talk about it in my spare time. Let me opt-in to Birdwatching, and opt-out of nuclear engineering so those weirdos at work don't find me. 3. Limit the Time for association. Maybe that means you only 'connect' people with their comments in the past year as the Public default. Once the user connects with your system and has control over their online persona, then the user can decide what to show and for how long.
Future use? 1. Push this tool as an anonymity improvement tool, which also encourages meaningful conversations with people you want to connect with. 2. Feature: "We notice you posted for the first time to /r/bsdm with your /u/PepsiOfficialMarketing. Was this a mistake?" 3. Feature: "Our AI has scanned your recent posts. We have connected your writing style across two accounts and this may compromise the seperation of accounts. The following phrase: 'Do what thou does MFK!' was used to connect these accounts by our AI."
(*Compared to current internet gasoline fire of ad-infested, algorithm targeted, Top-10 lists of Top Ten Lists.)
I left Twitter right before it became X, and focused on HN as a “breaking news and commentary” that is mercifully free of politics.
This tool fills a discovery gap.
In any case, interesting idea & project. Philosophically i'm not sure i like the idea of identifying experts - i'd much rather people's comments stand on their own instead of their clout, but nonetheless definitely interesting.
Would things like this create/increase incentives to game the metrics?
Anecdotally, total game time (in hours, usually) is used to convey experience in video games (WoW, CS:GO, PUBG). I've seen people create & run 3rd-party software to artificially inflate these types of metrics.
So, when something like HN has goals, and you create, say, something akin to a contest and link it to real-world benefits (e.g., people expect a recruiter or hiring to use that "who knows what" info), then many people will modify their behavior, probably to the detriment of original goals.
I looked myself up and "google" is proeminent. However I'm only posting anti google comments, nothing technical but on the privacy theme...
I also looked up 'yocto', which is something i know something about and i mentioned in posts a couple times, and the first user returned has some very interesting tags:
gur juvpu ohg guvax evtug qbrf
And the only really related tag i see is 'buildroot'. I guess it's just not a popular enough tag for the machine to have enough data.
Edit: and to join the choir of concerned voices:
1. It's not 'trusted'. It's at best 'popular'.
2. It reminds me of the main social networks and 'engagement'. Hope HN never becomes as predatory.
ROT13 for "the which but think right does". Not really any less cryptic even when decoded.
I guess mundane words in a ROT13 conversation people had once have attracted the keyword detector.
I suppose yocto does sound a bit like rot13...
I like it. I feel this tool is quite sophisticated and incredibly accurate.
Seems pretty accurate. I do love me some PHP and WordPress.
I guess I should apply for a job at ESA.
Why not add their Gmail signature as well?
"Trusted voices" on HN are often not experts, yet espouse views which are not what the larger body of experts would consider correct. In addition, actual expert voices are drowned out by whatever the popular position is. HN also provides its own cultural bias, in that everything spoken on HN has to follow a rigorous cultural sieve set down by the guidelines, such that a "negative" view, even if correct, is considered either wrong or distasteful and buried. This is exacerbated by banning the use of humor to disarm controversial or heated comments. And this is the comments that HN does get; many opinions are never entered as comments here, so there exists a large knowledge gap. Then there's the "taboo" subjects like race, gender, religion, politics, social justice, etc which get buried for fear of controversy, so you're definitely not gonna find any expert opinions on those, as the stories just aren't there for discussion.
The end result is that often experts go unheeded or even downvoted, popular shallow opinions get upvoted, and substantive commentary based on evidence and experience is frequently missing. The fact is that we have no idea who knows what, or what's true or right. We just have "popularity" according to the particular cultural quirks of this site. So you can definitely find out "who thinks what", and who is considered to be more trustworthy to a HNer, but it has absolutely nothing to do with objective truth or the body of real knowledge that exists outside the world of HN comments. This is an echo chamber, but it's not a chamber of experts. It just seems that way because occasionally you see a minor tech celebrity, and people talk with absolute authority regardless of if they have any.
You want to find out who knows what? Look at their diplomas and careers. If they've done 20 years in a single field, probably they're an expert. If they have a degree (or multiple) in a field, probably they're an expert. If they spent half their life working on a single hobby, they're at least very knowledgeable in that field. But you can't determine that just by looking at who's talking about what or how many completely subjective "points" they get for what they say. Determining real knowledge requires analysis of specific criteria, filtered to get a higher quality result.
This is a serious site, people!
I have a Ph.D. in neuroscience with exactly 20 years of experience. I talk about what I know and try to be curious about what I don't know. HN really helps me on both counts. How does it help you?
My name pops up.
Hell yeah. A++++ completely accurate would use again.
Please no. That sounds dystopian. We should not prefer having algorithms meddling with social interaction. We should not want things better designed to manipulate us.
I guess it’s inevitable at this point, but how does it feel to be among the first who dumb down real people to a set of caricature keywords based on an ml method of the day?
Better than at work where real people dumb down to a set of categories based on code, emails, age, gender, and projects?