Analyzing HN readers' personal blogs
dannysalzman.com
dannysalzman.com
So HN readers are not necessarily contributors. And not all contributors would plug their blogs in a thread asking them to do so.
If you want to get an idea of the HN readership rather than of the HN contributors you may want to start off by scraping all the profile pages instead, it will give you a much larger set of sample data to work with.
Oddly enough, I'm in the middle of this project right now! There are over 600,000 users, and it turns out that many of them use their profiles to share links to things other than personal blogs. I've done some scripting to automate deciding if a link is a personal site or not, but the whole endeavor has been significantly less trivial than I had hoped.
Regardless, I plan on sharing results soon, so be on the lookout!
Could one of you downvoters please explain why you think people would post blog posts if they do not want to share them? I fail to see how that would be possible.
I called every bank in the country that was listed by the central bank, I knocked on their doors until I found one that offered students a card.
I even had to write a document because they didn't have a form for that case. I finally got my card after a week of trips to the bank, city hall, administration, providing documents that were unlisted, etc.
I computed that this wasted time multiplied by the number of people who would eventually go through the process and the fact almost nobody in the country had the card warranted detailing the steps. It would save time, and contribute to reduce the "underbanking". I wrote a blog post explaining how to do it with steps, as in provide the following documents, take these amount in these currencies. I attached the document/form I wrote so people could fill it and take it directly to the bank and save a trip. That asymmetry in information bothered me.
That post 2013 post is read hundreds of times per day. It had more than a thousand comments, although it only lists 700+. People asking me questions, then getting their own cards, then themselves answering other people's questions, then updating me on what has changed.
People I would meet in real life would tell me I looked familiar, and then they'd put it together and tell me they followed the post to get their card. Sometimes people referred me to my own post in case I wanted to get a card.
I received emails from people freelanced online who wanted to bring money back here. One person even sent me credentials to their online account with thousands of dollars in and asked me if I can find a way to transfer it here (I told that person not to do that again and she said she felt she could trust me).
Many times I'd receive a phone call from a friend who'd say they wanted to get a card, looked online, found a great post, then saw my name and laughed out loud because they knew me.
And of course, I met interesting people.
https://jacquesmattheij.com/if-you-have-nothing-to-hide/
This one too, even inspired a play!:
Ultimately, though, all the sharing of our blog URLs and this related discussion made me realize that I didn't really want an audience, so I killed off my domain.
https://twitter.com/shit_hn_says?lang=ca
I think the threads on imageboards mocking HN are generally funnier, though.
Exactly. On a pseudanonymous forum that’s as good as asking people to doxx themselves!
People can have a variety of personal concerns, from a nutcase stalker ex to "I work for BigCo and want to spout off online without getting fired" to "I happen to have some bizarrely unique name, so using my real name anywhere amounts to doxxing myself."
Lots of people on HN use a throwaway email just for HN and don't want general HN readership to know much about their lives.
People use a wide variety of approaches to having an online life while looking out for their own specific privacy concerns. Please note that most people with privacy concerns will not chime in to this discussion to explain to you why they make the specific choices they make as that would tend to be counterproductive and undermine their goals of protecting their privacy.
Where does this strange typo come from?
The term dox derives from the slang "dropping dox" which, according to Wired writer Mat Honan, was "an old-school revenge tactic that emerged from hacker culture in 1990s". Hackers operating outside the law in that era used the breach of an opponent's anonymity as a means to expose opponents to harassment or legal repercussions.[9]
>leet
one of these is not like the other
I run a popular website, but I remind people Apple is Evil.
The friction is just great enough that most people still stick with GA, since it's already pretty much everywhere.
Also, 382 sounds like a very small number, given the size of HN. I did try to find many blogs I read, but couldn't find one. So, crucially - was it a random sample? Or sample from top-liked, or from a particular month?
some findings (e.g. the prevalence of Wordpress) may depend on this procedure.
For example, here’s one from last year’s thread at this time: https://www.kickscondor.com/hrefhunt/6/
Very cool!
All of my other attributes were in the majority (Hugo, Google Analytics, etc...)
how can that be detected ? I'm curious.
That leaves nearly half of analyzed sites as unknowns! Speculation aside on what might be effective aside, the real answer wrt OP is basically "it wasn't".
A heuristic that might work would be to add a cachebuster query parameter to the page url (?cb=$RANDOM) and see how long it takes to respond. The idea is that the three most common setups are:
* static site served with apache/nginx/etc, which will just ignore the query param
* dynamic site, which will regenerate the page
* dynamic site behind a cache where the cache doesn't know that the query parameter isn't needed, and so the cachebuster will cause the page to be regenerated
<meta name="generator" content="Hugo 0.56.3">
<generator>Jekyll v{{ jekyll.version }}</generator>
But it's optional.
If it's a custom templates I think it's impossible.
There are 13 websites using Parse.ly, which starts at $500 per month! For a personal blog??
http://robotoverlordmanual.com/ https://medium.com/@andzwa https://medium.com/build-ideas https://medium.com/chrismarshallny https://medium.com/ing-blog/tech/home https://medium.com/@matthagy https://medium.com/modern-nlp https://medium.com/open-factory https://medium.com/@romanorac https://medium.com/@soatok https://medium.com/@zakjan https://sidstechcafe.com/ https://vladaionescu.com/
It’s artisanally hand crafted HTML with a little VanillaJS on a few pages. No static generator used. Also hosted on Netlify. Although I use BunnyCDN for large media. I post very infrequently.
Personally I found Matomo rather hard to use, I tried it for a while but decided that writing my own was easier than figuring out to get Matomo to do what I want (also: cheaper and easier). I think it's a good alternative for some use cases, but far from all. Related comment I left on Lobsters a few days ago: https://lobste.rs/s/cdrrty/why_you_should_stop_using_google#...
Bonus if you can segment the data by date as well so we can see trends over time.
A person that builds such a system would have access to some pretty useful data.
I had a chuckle. This is how so much data analysis happens in practice. Nothing like the command line for quickly cleaning some data.
Great work!
Surprising insights: many actually use Google AdSense, nginx over apache
I'd definitely be curious in: PageSpeedInsight (score, load time and size), post frequency/length past 12months, external link density
"time": {
"minTime": 0.7,
"maxTime": 53.5,
"meanTime": 5.5,
"medianTime": 3.1
}
"size": {
"minSize": 1,
"maxSize": 33484,
"meanSize": 1536,
"medianSize": 1565
}Based on my experience with the HN crowd, I'd predict that Gandi + Cloudflare would be a common one, with NameCheap closely behind.
Granted there were a lot of URLs in that post!
Or did I miss something?
Long story short, two completely separate backends each running on most reliable platform for the stack. And nginx is in front of it all. Ping the root domain and you’ll think you are on a big-standard Linux/nginx confit, even though it definitely is not.
So don’t assume no IIS!
Their stack is Rails+Preact.