Show HN: Semantic clusters and embeddings for 500k Hacker News comments
app.airtrain.ai
app.airtrain.ai
Any particular use-case to view the raw embedding?
How nice it would be to have an LLM trained on all of my previous writings and simply be able to click a button to indicate "reply to this person, please." I know I don't have enough training data from HN, and maybe not even from all of the sites I contribute, combined. It is still a nice thought, though.
But: Let's say I do acquire enough training data to have a local LLM do exactly what I describe. My volume of "replies" would certainly increase. Is that a good thing, on average? If the tool became ubiquitous, would it be a good thing for the average social media user? Or more pointedly, would it be a good thing for consumers of that social media? The cynic in me thinks "no" -- the effort required today surely weeds out _some_ idiots....
(Full disclosure, I nearly closed this window without clicking the "add comment" button.)
Would be a wonderfully fun app to try though.
Would it be a good thing? Do you think there are relies you should have posted that you did not?
I think there are probably times where a reply I deleted would have added value to the conversation. Not always, but perhaps enough.
> Would it be a good thing?
That is the issue. More != better, for all cases, as it says nothing about quality.
There is not enough training data for an LLM to emulate me, but what about prodigious writers? Could we realistically emulate Paul Graham or Joel Spolsky?
My idea was to be able to eventually use that data to train some AI with that data and to e.g. pick up on my different writing styles in a document editor, terminal etc.
[1] As per HN license: https://www.ycombinator.com/legal/
I know "me too" comments aren't generally acceptable here, but I think the wholesale pillaging of the web for LLMs etc is way out of hand.
I don't know precisely what changed that people decided that analysis is a bridge too far.
[1] https://huggingface.co/datasets/OpenPipe/hacker-news [2] https://github.com/HackerNews/API
It covers all User Content, which is comments and post titles (I think) at a minimum.
> Hacker News Information: If you create a Hacker News account (ID and profile), we do not collect any Personal Information unless you choose to provide your email address and/or information in the "about" field (“HN Information”). Your submissions to, and comments you make on, the Hacker News site are not Personal Information and are not "HN Information" as defined in this Privacy Policy.
I'm a little surprised people don't know the ownership story on HN. Didn't it raise questions when they realized they can't delete their posts without mother-may-I'ing the mods?
HN is pretty up-front that when you post here you are providing them content for free to more-or-less consume as they please.
> By uploading any User Content you hereby grant and will grant Y Combinator and its affiliated companies a nonexclusive, worldwide, royalty free, fully paid up, transferable, sublicensable, perpetual, irrevocable license to copy, display, upload, perform, distribute, store, modify and otherwise use your User Content for any Y Combinator-related purpose in any form, medium or technology now known or later developed.
My data is licensed only to Y Combinator and its affiliated companies (I would have prefer it to be licensed it only for news.ycombinator,com), not any other rando that can access it via a browser or an API
If they are YC affiliated, nothing I can do.
But I do feel saddened and personally betrayed; I thought the licence I gave to YN was just for news.ycombinator.com to store and show my comments, not for any other purposes.
Silly, silly me.
What this other company did with that (public) data seems to me to be a separate issue that you should take up with that company, just like the fact that your public comments (which you explicitly gave permission to HN to show) have been indexed by Google, Bing, and probably thousands of other spiders, bots, scrapers, etc.
I'm curious how you expected this to work. Like if you only give HN permission to store and show your comments on the public web, then somehow no other entity out there will be able to do anything with them?
Again, the fact my user data has been scrapped already doesn't mean it was scrapped legally. I'm ok with HN showing my comments, I'm not ok with anyone else than HN using my user data.
Will it though? I would imagine Meta would block them and then posture with a C&D or a frivolous lawsuit, but if they share the phones you gave them on the public internet, they're publicly consumable right? What law do you feel is broken there?
Or are you complaining here (ironically) to make a point, but you don't actually care that much?
HN cares a lot less; they're a tech comment site and don't actively discourage people using the dataset gleanable from the contents of the site for novel experimentation.
(Sidebar: I see "scrapped" coming up a lot in these conversations these days. Where is that neologism coming from? I'm familiar with people calling it "scraping" but it seems like the term has drifted for some reason?)
Given that the license is explicitly identified as both sublicensable and transferable and includes the right to distribute, I have a very hard time seeing how anyone could argue that the recipient of data that YC explicitly exposes through their "Official HN API" isn't allowed to use it.
If HN API exposes personal information publicly through their API then there is a problem.
And AFAICT the only way for HN to prevent user comments from being used by 3rd-party is preventing access to those comments, meaning a) sign-up will have to be more stringent and b) visitors will have to sign-in just to read (or scrape) comments.
You put things in a place with an expectation of a certain standard of use and then go after people hammer and tongs with a strict interpretation. Sometimes, the strict interpretation need not be valid. You can just shake them down.
I dont really care (for now) about this... but on the principle, I'm a bit fed up too by companies just crawling anything to train anymodel without any care about the datas, the people that produced them, and the consequences on people's life.
Maybe I could use the new European AI Act too (https://artificialintelligenceact.eu/fr/high-level-summary/) ... although I'm not sure because I didn't read it yet
It'll kill a lot of experiments (Mastodon immediately comes to mind; can't be pulling comments from other people's servers if those comments are attached to personal data like the commenter's username, right?).
Actually, Europe is so slow that a lot of experiments may take place. And there wont be any legislation if there no abuse...
It took a loooonnnng time for Europe to react to Facebook, Google & co abuse with users datas. Same for OpenAI using a awful lot of copyrighted material without giving anything back... So thank'em for Europe legislation :-)
I'd be very surprised if that were the case.
Personal data is (broadly) considered to be data that could be used to track or tie your behavior online together into a profile. The UK's ICO calls out usernames specifically as an example of such data. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-re....
(For those of us who have been around on the Internet long enough to remember the era where people intentionally chose handles to remain pseudonymous and separate from their IRL personas, this seems counter-intuitive and a little preposterous, but the GDPR doesn't care what "netizens" think about privacy; it's a broad attempt to impose a "non-native" concept of privacy over the preexisting net culture).
For example, you may agree to have your linkedIn profile name next to your HN username... maybe. But I'm not sure that you would agree to have your LinkedIn profile name next to your Tinder username.
And you sure don't want that to happen without your agreement and even without you knowing about it (but learning about it from a colleague for example).
That's why GDPR has some right to deletion or modification. And why some days, Europe may go after data brocker
(as a side note: not sure why my comments were downvoted. I didn't say that I would go after anybody - and surely not HN - I only said that uncontrolled use of any data without any anonymization and without consent might be the source of problems with regard to legislation decided BECAUSE too many shady business abused of it. You may not like it but then... well... downvote the abusers)
Am I not allowed to read your comments?
Am I not allowed to learn from reading your comments?