You can identify who is an expert related to each tool or any topic in general
You can identify who is an expert related to each tool or any topic in general
Why am I tense? Is it the sentences I formulate? Upvote vs. downvote patterns? RNG?
I pretty strongly believe that you should take that piece rather seriously, even if you don't previously consider it a big deal. People generally don't like when a stranger tells the world what their mood will be. Especially bad when someone believes it's quite wrong.
Also I love the opportunity for introspection. I'm trying to become a more whole, happy person who isn't defined whatsoever by his career. I would love to learn how some algorithms perceive me and why.
Highly recommend reading the "how it works" section (I assume you already have):
https://hnprofile.com/learn-more
That section has a few examples and details. I'll add further details below.
The system ONLY looks at comments sent to it. It, kind of looks like:
{ Author, Comment }
From there, everything else is derived. Nothing else, no profiles, no voting, nothing. That's it.The system identifies:
- Expertise
- Knowledge
- Interests
- Relevant content
- Promoter score (an improved net promoter score)
- Trends
- Mood
- Related Topics
- etc.
To your question, there are multiple algorithms at work determining mood, most are pretty standard models. These are ones you can find in NLTK, spaCy, CoreNLP, etc. There's also a couple home-baked ones. These models than "vote" for the mood, with an averaged score. Typically these scores are 3 - 7 models voting at any one time. The aggregate score is roughly 80% accurate (based on my tags), and it tags the topics, sentence, and overall comment all separately. Finally, the models are much more accurate at the extremes: Very Positive and Very Negative, but in the middle, it's much fuzzier. This is in part, just how language works, it's subjective.All NLP models are based on a combination of open datasets with the tagged sentiment, as well as my own manually created dataset, and finally, a dataset that was manually created via looking at following comments (for instance, "no need to be negative", is used to tag the prior comment.
The probable current mood is a prediction of your prior moods. I believe right now, it looks at the last time you discussed the given topic you searched (if you searched for a user, it is the most recent mood). If there is enough data, it'll also predict your mood, based on the time of day + topic being discussed.
Finally, I should add - there are definitely improvements to be made, and I'm currently working on mood. Primarily, the focus was on expertise, knowledge, and interests. That is what enables the "related content" section and is now very accurate (given a decent number of data samples). Mood, as stated, is ~80% accurate based on my tags. I think I can do better (probably >90%), especially with more labeled data for this particular dataset.
The goal of this is it can drop into any company and instantly can search for who and what is relevant to your company.
Last I checked I was "tense". Is it theoretically possible to explain why your models came to that conclusion? Is there a theoretical data view that could say "here's the comments that contributed to you being considered "tense" (or even better: specific parts of the comments).
From my perspective, without a way to concretely back up any conclusion made, I simply cannot consider the complex-sounding systems to be any better than RNG.
I alluded to this in my first comment, but I think it's imperative that if you're going to explicitly call out what mood you think someone was/is/will be, you need to back it up with "why" so that, at the very least, someone can say, "that's clearly an unfair representation of my mood and here's where the algorithms messed up."
I'm going to go on a bit of a limb and assume that you're just having some fun and exploring an interesting feature. And I think that's really quite cool. But I think we're getting into a software engineering ethics realm that I'm ill-equipped to dig in to. You've got a site that applies "moods" to individuals in a way that doesn't appear to be verifiable. And even if they are, and are accurate, is that even the right thing to do?
I don't want to ruin your fun or twist your arm into changing this. I'm just responding to my red-alert alarm that lives in my stomach. It's set off by the discovery that one of my few online accounts where I don't hide my identity has been assigned a mood label.
That being said, I'm 100% confident that Google, Apple, Amazon, and more know WAY more about you than I do. This system only looks at public data and only links back to usernames. Strangely enough, it can identify users across sites (even without knowing your identity) - say Reddit and Hacker News. It was originally built for https://projectpiglet.com/ to find company insiders, and make trades (which works ridiculously well btw).
There should be no expectation of privacy.
> You've got a site that applies "moods" to individuals in a way that doesn't appear to be verifiable. And even if they are, and are accurate, is that even the right thing to do?
The short answer is, I don't know. I don't want this to exist, but it does. I also know that I have already built systems for companies doing this and more. Albeit the system here is much more of a "turn-key" solution and its patent-pending for what it's worth (I'm also against patents and donate to the EFF, but feel I can potentially segment off this market). Still, lots of people are looking to build systems such as this. Personally, I'd rather be the one guiding the ship, because I do ask myself the same questions you're asking.
I'll be honest... I've used this system to show my friends that even when they change usernames I can still identify them. That particular system (which looks at speech patterns) is turned off here. That's in part because of those questions. However, I'm 100% confident, others are doing that right now, with this conversation. We know that because Snowden shared it with us.
I would be shocked and impressed, and it would serve as a great demo.
(I totally believe you. I'm just fascinated.)
Basically, I'm volunteering to be a lab rat.
(The algorithm identified not one, but both of my former alts.)
Edited to tone down the reaction, but that's amazing.
As I said in the previous comment, that IMO is going too far. Even though I'm 100% sure others are using similar systems, I don't feel comfortable in that business -_-
Kinda highlights, although the hnprofile.com demo may have it's current faults - there's a lot you can do. The cool part, is the platform (called Metacortex) is easy to build apps on top of.
If you’re ever in the Chicagoland area, happy to buy you a beer!
I think a larger demonstration could make for a very meaningful "Show HN" with hopefully a large impact on privacy consciousness.
I'm sure some of us expected this, but to have it demonstrated by a single person in a casual "Oh, it could do this as well, but I don't have it turned on for ethical reasons" manner is quite effective.
Maybe it will raise awareness that "big data" isn't just trying to correlate cat-pictures one posts on their communication medium to cat food advertising, but everything you write anywhere on the Internet, even on different media under different (or no) accounts, into permanent and guarded profiles with no recourse of opting-out of the machine, effectively rendering informal discourse over the Internet dead.
I'm not necessarily the one guiding the ship, but I'd like to be (by being the largest on the market), and I'm just working to launch this as a business (called Metacortex). Fact is, I've spent years building this out for my trading platform. Most of the research is my own, with some contributions from open source and academia as it relates to NLP. That's why I submitted a patent on it, I want to corner this particular (and seemingly highly effective) method.
That being said, I've worked on products with similar goals, and have done technical reviews of WAY creepier products. Luckily, those other products rarely work at all. IMO the technology isn't quite there yet (outside of some of what I demoed), however, it's uncomfortably close.
Snowden told us this was happening, years ago. I'm sure they've improved since.
Which (turns out one of my areas of interest is "which") really only gives you a monopoly on selling the method but not on internal usage by less scrupulous actors...at least that's my opinion, it would be near impossible to enforce a patent on non-public facing code used to track people in BigCorp.
I do like the idea though, it's not like anyone (reasonably) should believe they have any semblance of anonymity on the internets and this just goes to show how easy it is to follow you around on the webs. If an individual can pull this off then the sky's the limit for the TLAs with massive budgets and computing horsepower.
Ok, now I've waded in, I should do something to help. Reads the first few pages of your comments.. hmm Nah, you don't seem tense. All kinds of emotions. It's not you, it's HNProfile. Tries to think why it might say that.. There were a few with emotional words, frustration with situations..probably slightly more negative-emotion words than positive.. hmm I think maybe you just write better than most people, more vividly, mostly about serious topics. And are engaged with them. I don't know if it helps, but I say That's clearly an unfair representation of your mood. (I don't know you, wasn't paid by you)
I have an even better way to identify who is an expert on HN. Say "It's silly that X still isn't a thing." If X is a thing, you'll get ten people telling you why you're wrong within ten minutes. Example: https://news.ycombinator.com/item?id=17941772
You'll have to trade some karma for the answer though.
> dang
> Probable Mood: Tense
The eternal life of an admin I suppose!Sentiment analysis is a finicky thing. (Maybe categories like "polite" or "argumentative" would be easier to detect?)
"youtube, everything, video, times, data, issues, software, experience, companies, point"
Its not entirely wrong, but I don't know what "point", or "everything" means honestly
I should show this to my wife (:
Mood wasn't exactly the focus, so it's robustness is limited - i.e. lots of room for improvement. The mood detection (based on my manual tagging) is roughly 80% accurate, which IMO was "good enough". It's more accurate as you get to the edges (very positive or very negative), the middle is kinda fuzzy, but that's kind of the nature of language.
[edit] I do wonder if the analysis of topics is affected by both the frequency something shows up on HN[1] and not repeating the keyword in a reply post.
1) less examples and more likely to miss a topic a person is expert on
The probable mood was rather... entertaining.
The notion this can be used to find sockpuppets or throwaways does not surprise me. Without sufficient effort you still use certain words, in certain order, consistently. Such can even be used to check out what your native language is. So if you are e.g. a blackhat who wrote ransomware it makes sense to have someone who never wrote on the internet to paraphrase or deal with the communication, or perhaps use tools such as translators.
Also, if you use delay (e.g. delay 10) it is going to appear on HN after you wrote it meaning it could appear in the next hour despite being written in the previous.
Despite the above, I did find this an interesting project; I just very much doubt its accuracy and usefulness.
[1] https://hnprofile.com/author_profiles?utf8=%E2%9C%93&search=...
Happy. Happy. Happy...
I suppose if you desired you could consolidate your doxx point to Google and ISP's by running all your comments through a couple rounds of Google translate.
There is no privacy possible in the global village. Expectation of such is technically naive.