Similar Hacker News Users
swimwithoutgettingwet.com
swimwithoutgettingwet.com
S W G W einstein
newton
+------------------------+
| edw519 | liebniz
+------------------------+
turing
===========||=============
carnegie
Co-Commenting..SEMANTICS..
Word Choice tesla
godel
=====||===================
escher
Leaderboard.....KARMA.....
Diamonds in the Rough..... bach
edison
This site has no
affiliation with Hacker galois
News or Y Combinator
patio11http://news.ycombinator.com/item?id=701
And recently, too:
http://news.ycombinator.com/item?id=1036247
How does this particular tool work? It's based on the threads a given person comments on, who else comments on those threads, and how the topics and terminology of threads and comments relate to each other. Karma on the relevant threads is used as a subtle authority metric. Incorporating voting relationship histories would almost certainly make the tool better. Particularly at finding interesting (and not merely similar) stuff.
That said, it'll be interesting to hear what folks think.
http://www.swimwithoutgettingwet.com/hnusers/?user=antiismis...
I've done this before on different datasets and would love to cooperate with you on it...
- Defaults should be almost sufficient to use the app. (i.e. if it has some workflow, you should be able to pretty much hit "Next next next" and get it to work. Bonus points for making the number of nexts as low as possible.)
- Pick something that results in visually impressive output rather than a blank page. See Balsamiq for inspiration here -- they start you with a mockup in progress that demonstrates most of the highlights of the software.
- Ever seen Firefly? I really like how they use the word "shiny". Ideally, your defaults should show the shiny in your app. In Hollywood they have a saying: make sure your budget makes it onto the screen. In B2C apps, make sure the stuff you did all the work on makes it into the user experience most of your users will see.
- Assume your user is a novice at both your software and the problem domain until you have evidence otherwise. A lot of people ask the user "Hey, are you a novice?" That is one way to do things, but it makes your core workflow one stage longer and every stage costs you conversion. I prefer "Assume they are and give them a discrete 'skip ahead' button" or "Assume they are and watch them for evidence that they are not".
- If your app is supposed to make the user feel like they just killed an effing lion, then your default settings better have a lion bound and gagged sitting under a forty-ton weight suspended by a weak string which passes through an open pair of scissors next to a sign saying "Snip this."
- (Do this if nothing else.) Track actual usage of the app and modify your defaults based on actual usage. Bonus points if you can do it dynamically, if that makes sense for your app. For example, if you pick A as the default and 25% of your users go out of their way to change it to B, then that probably should have been B. (You can split test and see how many people would have changed it to A if it had defaulted to B.)
When I find people breaking these sorts of rules, it's usually because they're thinking of themselves or some non-customer stakeholder.
edited to add a corollary: until you've done the sort of testing patio11 advocates, you don't know who that customer is.
For example, say you're targeting teachers. Keeping a laser focus on what teachers want is important. However, I think you need to put extra focus on making their first five minutes absolutely amazing. (And their first 30 seconds. And their first 5 seconds.) My reason for this is simple: almost all apps are going to leak an amazing number of their customers between first and second use. I don't have my report in front of me at the moment, but I think something like 40% of BCC users never complete their first bingo card and never log in again. Essentially none of these people buy the software. On the other hand, roughly 2.4% of trialers convert, or roughly 4% of users who succeed in their first interaction with the app. Increasing my bottom line by 5% requires converting 5% more of that second group -- which, let me tell you, is hard freaking work -- or, in the alternative, improving the first run experience of three out of every 40 users who fail to complete Task #1.
It is really hard to optimize your entire application, user experience, value proposition, etc to get +5% conversion. However, polishing your first five minutes until it freaking shines is not nearly as difficult. Do you present people with a blank page currently? Spend an hour and put something on it. I will put money on that helping. Does it take a critical mass of friends/input/lions slain to get fun? Do the work for them. Fake it if necessary.
And, yeah, instrument everything. Can I plug Mixpanel here? plug Every time I think "You know I should really build more instrumentation into my app..." I remember "Oh, wait, it takes a twentieth of the time to just throw it on Mixpanel -- no visualization code or complicated controls to refine the data range required, praise be."
Start here: http://www.swimwithoutgettingwet.com/hnusers/?user=profquail...
The next two 'clicks' to the left (towards co-commenting) don't have 'spolsky' in my list: http://www.swimwithoutgettingwet.com/hnusers/?user=profquail...
But one more click to the left, and he reappears in the list: http://www.swimwithoutgettingwet.com/hnusers/?user=profquail...
I think that only really says something about HN, not about my comments :)
The only name I recognize is pg, and since I save my username and password, sometimes I forget my own username too.
edit: just realized that by default it's only returning the rockstars of the site (e.g. people high up on the leaderboard).
Not sure what that means though.
http://news.ycombinator.com/item?id=181868
:-)
Another route might be clustering ('technical', 'political', etc.)
I'm guessing either such a corpus was used here, or it's based on a cache of recent comments.
Neat little utility, thanks.
Most of the techniques for this sort of thing don't work that well for sparse datasets. So given a choice between showing bad results for users with relatively few comments, and not showing any results ... especially when users can search not just for themselves but also for folks they know and enjoy ... Also, scraping the full history of HN is not cool.
http://news.ycombinator.com/item?id=271066
I'd check the Internet Archive for the files & Google for the filenames.
For instance the top pick for tptacek is cperciva, which seems natural. Doesn't work the other way around though, so there's still some work to be done...
Put another way, it seems as though the algorithm considers tptacek more distinct (from all other HN users) than cperciva.
I take it that means I am the cheese.
works: http://news.ycombinator.com/user?id=SapphireSun
doesn't work: http://news.ycombinator.com/user?id=sapphiresun
One newer participant whose posts make me think "I wish I had posted that" doesn't show up on my list of associated participants. But I show up on his. Maybe that is because of the karma setting in the default operation of the search. Interesting.
If three things were to get added to this, what should they be?
- username has period afterward, for self-linking back to site
- various ratios next to each user name in a table, maybe linking to searchyc ie http://searchyc.com/user/riffer
2) For each of the matches could you show some of the data used to compute the matches... e.g., for semantic matches, show the top X (maybe five?) common words or phrases you matched on. For threads, show the parent or OP of the five most recent threads... something like that.
3) Really neat. I like this. It's quick too. How did you do it? Did you replicate the entire HN database? Third suggestion: post the source or just post an explanation of how it works.
osteele erikstarck zedshaw atarashi matthavener clemesha alrex021 cliff NateLawson jefffoster mrcharles fix3r
This was done with the karma effectively off (same as the karma slider in the middle). What do you think? Anybody there you think shouldn't be?
How about displaying the values used to order the result set, so you could compare weightings across users? reply
When I set the sliders 'semantics' all the way to the right ('word choice') and leaderboard all the way to the left, then check I get this:
- patio11
- mahmud
- nostrademons
- tptacek
- mahmud
- edw519
- swelljoe
- davidw
Shouldn't the relationships be symmetrical, so 'edw519' would get 'mahmud' as the first match and 'mahmud' would get 'edw519' ?edit: also, your 'match' is case sensitive, so 'riderofgiraffes' won't work but 'RiderOfGiraffes' does.
All of those guys showup in my default results, but I can't get any of them to pick me up. Clearly I don't post enough.
I've received some email from David (the guy that built it), he's going to fix this and the lowercase issue as soon as things quiet down a bit.
My matches using the default settings: pg, swombat, mahmud, wheels, SwellJoe, edw519, mattmaroon, gojomo, davidw, mixmax, unalone, tptacek
Did you get permission to scrape the data? (I tried once without asking with mediocre results: http://www.mattmazur.com/2008/08/the-wrong-way-to-get-notice...)
Slight note about the page formatting: My screen resolution is 800x480, and the text by the sliders wraps in a very confusing manner.
It looks like this to me
===============||===================
co-commenting......SEMANTICS.....word
choice
A little confusing at first, until I realized that it was wrapping. It's the same for the other slider.Great app though!
Are there any other details about the algorithm or how it works. I'm curious about what exactly the different weightings mean.
Are quoted sections filtered out? URLs?
Security risk blocked for your protection Reason: This Websense category is filtered: Potentially Damaging Content. URL: http://www.swimwithoutgettingwet.com/hnusers/
edw519
btilly
patio11
pg
tptacekId say thats a good thing. In general the whole list is people I would happen to be even somewhat similar to.