Are Product Hunt's featured products still online today?
scrapingbee.com
scrapingbee.com
People pay for pain meds, product hunt featured products are colorful vitamins.
We are now a 20+ people team, 400+ B2B customers and $12M raised.
Really, really solid product. Checkly is a very core part of our infrastucture.
Why is that even a metric to speak about?
Shouldn't the metric that matters most be # of customers, how engaged they are with using the product and ultimately sales?
Bringing up how much you have raised seems like the priorities are misaligned.
But this is a forum managed by a VC so some people are certainly interested in this. I also thought it was interesting in the context of a side project launched on PH.
Feel free to check my Twitter (link in my bio) on how I think about and interact with customers.
A coin has a 50/50 chance of being heads, but just because the coin landed on tails last time doesn’t mean it’ll be a head the next time.
Here is the link I should have included originally https://en.m.wikipedia.org/wiki/Ecological_fallacy
No, you shouldn't read much more into that, but there factually is information behind the number.
Not to mention that convincing a vc, regardless of clout, is not about objective profitability or anything correlated with success.
It’s about convincing a person that you can make money, and people are in capable of being objective
I feel like this is flawed, especially considering 1/2 of the successful responses were 3XX. It's possible that they had just linked a short URL that was a redirect, but it's also possible that the product was shuttered and a redirect put in place to a replacement product, the company homepage, or even an acquiring company. I don't think there is an easy way to tell based just on the response code, and I'm not sure you could even programmatically determine it unless you had samples of what the pages looked like on launch day (maybe compare today vs the Internet Archive?).
I don't visit the site really every anymore, but I did take a week the other week, and like you, nothing really stood out to me.
I left my job end of that week and been doing it full time for a year now.
However, I agree with comments about launches in general. If you have a good network, you can launch rubbish and end up in the top 5.
The most interesting chart is one of the last: Proportion of Failures over time. As expected, more recent product links are less likely to 404 or 5xx.
Going back to 2014, almost 1/3 of the featured links give a 4xx or a 5xx response. That’s a lot!
More surprising, links as recent as 2020 show a 1/4 failure rate. Those projects basically launched on PH, then shut down shortly afterward.
Moreover, this analysis can’t actually account for products that have been shuttered but still have landing pages online. It’s ultra cheap to keep a placeholder “Sorry we’re closed” page online, so I imagine a lot of these projects are shutdown but counted as “success”.
Subjectively, this matches what I’ve gathered from watching PH. Getting a PH featured product listing seems to be a badge of honor, but PH users aren’t really interested in using 99% of the products and the submitters aren’t actually interested in building them past proof of concept. Recently, the bulk of postings seem to be advertisements for paid information products or pay-to-join communities.
But anyway,
I thought about taking a random sample of pages who returns a "200". Let's say 150, and manually tagging them to find if they're "dead" or not.
And then reuse the "dead or alive but a 200" ratio for all the pages but I was afraid that I'd need to tag much more than 150 pages to have a significant statistical result.
It’s obviously blog content designed to promote your product, hosted on the company’s product website. I don’t see how the FYI is unfair.
I added it because the content was valuable but HN can be finicky about blog posts from companies advertising their own products. Trying to get ahead of indignant dismissals.
There's so many blog posts posted here that could fall under "content marketing" umbrella if you want to be strict. I feel like there's no problem with that if the content is valuable and people like/upvote it. After all this is a platform that is doing marketing for YC where YC companies are supposed to post their content too.
That "warning" also stuck out to me as a bit unfair as I was even looking for how it hooks into ScrapingBee (as I was curious how these scraping-aaS platforms interface with custom code) and couldn't find anything.
from scipy.stats import binomtest
chance_of_dead_page = binomtest(landing_page_counter["dead"], landing_page_counter["total"]).proportion_ci(confidence_level=0.90)
print(f'Chance of a dead but existing landing page (90% Confidence Interval):{chance_of_dead_page.low * 100:.2f}% to {chance_of_dead_page.high * 100:.2f}%')I've worked in or adjacent to the content marketing world long enough to know that a CTA is not necessary for the post to be marketing/advertising. One of the major goals of content marketing it to establish the authority of the brand. You are well aware that the raison d'etre of that post is to spread awareness of and establish the authority of ScrapingBee.
It doesn't mean the post is not interesting, useful or valuable. But that post exists fundamentally for marketing/brand purposes.
Parents warning is completely fair, especially since they immediately point out the value of the post.
If you're paying attention you know it's content marketing. If you're oblivious, the marketing probably isn't working.
Either way, you probably don't need a warning.
Is that a lot? I would have been less surprised if it were 1/3rd of links still live.
Sure, but it's no different than any other blog post from a company. And framing it that way is quite disingenuous since the post pretty much only sticks to the topic and doesn't overtly promote their product.
Yeah we have all seen it is on scrapingBee, no need for a warning.
This would very well fit a "fail fast" attitude with testing MVPs, wouldn't it? At least that's what I would guess. Got a great start with PH but didn't move on from there, so the domain was not renewed...
Granted there are cases where that market IS aligned with your product, eg if you built a low cost site-builder or low cost social media publishing platform
My startup is still around, [1] and we posted on PH one or two other times when we launched new products. Even though we had some powerful hunters (thanks to our early presence on the site), I found it took too much time to be worthwhile for follow-on product releases. I'd be interested to know if others have had the same experience, or if they have tips for how to get a meaningful bump out of subsequent posts.
1. How the online ones are doing financially.
2. Which sectors are doing well - what are the trending tools.
Skimmed through PH APIs, don't think this is possible. Courtland's Indie Hackers (they have stripe verified revenue) maybe of help - a quick google resulted in this¹ result
There's also microconf report on SaaS's²
[1] https://www.indiehackers.com/post/indie-hackers-are-making-6...
We could also have analyzed the sitemap to check the last update date.
Those articles are really fun to write (I haven't written this one, I'm just the editor), but at some point you have to stop otherwise you end with a 20k words essay.
I think the parent article is interesting, thanks for your contributions. I am not saying that the same should have contained revenue, performance data - just that it would be interesting to see :)
It makes it easy to see based on when a product was featured whether it’s becoming more or less likely to fail after a given time period.
Disclaimer: I created that demo
> Thanks to our large proxy pool, you can bypass rate limiting website, lower the chance to get blocked and hide your bots!
> Scrapingbee helps us to retrieve information from sites that use very sophisticated mechanism to block unwanted traffic, we were struggling with those sites for some time now and I'm very glad that we found ScrapingBee.
(not a user, but I do some amount of scraping through other means)
[EDIT]
> One of his opinions about programming was that programmers work "bullshit jobs" for their employer and do cool open source stuff in their free time which is demonstrably false.
Further, I'm not even sure that's incorrect. It can both be true that most open source (that's actually used by anyone) is done by people who are paid to do it, and that most programmers have very little interesting or challenging to do at work unless they work on hobby projects—maybe open source—in their free time.
The overall letter as quoted in the book, and Graeber's commentary on it, actually makes some good points aside from all this. Things don't have to be perfect to be useful.
Which would have been fine except they also imposed terribly low rate limits with no ability to check them.
We eventually pulled the partnership since it was more work than value.
Allowing unfettered scraping and repurposing of data would have a chilling effect on all types of services. For example I wouldn't necessarily want a bot to scrape my comment history on HN, doxx me, and share my identity and comments with others.
Running a site thats had a bot get stuck in a loop and suddenly x10000 times the request rate, when they go wrong it’s super annoying for the website owner. We ultimately just banned the whole AWS ip ranges.
At least in the United States, sounds like the jury is still out on the legality: https://www.reuters.com/technology/us-supreme-court-revives-..., but my perspective was more from an ethics standpoint anyway.
That's not what this tool does though. It allows you to distribute your scraping to a layer of proxies. So, the only difference is whether there is an intent to do harm to the target or merely collect data... which could be a form of doing harm as well?
There's definetly an argument that dangerous tools should be regulated to varying degrees. If we're arguing regulations in this specific area you'd probably also be balancing it with regulations that sites can't close an account for reasonable rate automated access and that public research can have higher rates so long as they're not crippling.
I wouldn’t regulate this but If you’re introducing regulations, why not just require the source to deliver the data in a neatly packaged format? The necessity for scraping and the potential for DDOS and potentially nefarious behavior basically goes away.
I think that means the jury is still out, as you mentioned, but it's leaning towards scraping being legal as long as the data is publicly available. IANAL
Please verify your experience with the Google ip range.
https://developers.google.com/search/docs/advanced/crawling/...
A lot of crawlers spoof the Googlebot user agent so you wouldn't block them ;)
It’s not a web crawler. They are all web scrapers. And Alphabet/Google sells this data and makes profits from it.
It is not like it is trying to hide the fact that it is king web scraper.
Google has gotten in trouble from various publishers for this before. It is no secret there is a double standard in big tech.
Again if you are going to arrest a web scraper, then arrest the king of all web scrapers first to make it fair.
Data wants to be free. If it is publicly accessible then it is fair game.
Source ?
You’re a smart guy. Surely you must know how ridiculous that sounds on the face of it.
It is common sense.
The sky is blue.
Source: Look up at the sky.
Think about it: Google has every advantage by respecting robots.txt and nothing to win by ignoring it.
Eg.
1) If a media company doesn't want to get crawled: add it in robots.txt
Then they realize their visitors drops and they'll remove it again.
Ergo: publishers sue. Because they want the advantages, but without the scraping. Which doesn't seem logical to me, since they currently give Google explicit permission to scrape content.
2) if they would sometimes leak personal documents protected by robots.txt they could have a lot of lawsuits on their hands.
Robots.txt is a simple method to not get blamed.
Ignoring robots.txt could literally be a core business liability from my POV.
---
So please, source outside of gut feeling, as requested before, would be greatly appreciated.
Im not sure why robots.txt was even brought up.
So google respects this file? I say so what.
Im arguing that while Google has free reign to scrape whatever data it wants, we indie devs are subject to the cider house rules.
Sources can be found for just about any argument. So they are more or less useless.
There is nothing wrong with self evident truths or reasonable hypotheses. That is how the modern world was created.
A search engine that scrapes the web for data to make a good search engine. Who wouldve dreamed of it?
We are not privy to what happens behinds closed doors at Google. They only work for their shareholders. Not us or the public good.
Source that google does what it wants based on what it thinks the web should be. Google can change its mind on a whim https://www.searchenginejournal.com/google-robots-txt-noinde...
It's an exclusion standard, not an inclusion one.
https://en.m.wikipedia.org/wiki/Robots_exclusion_standard
For helping individual url discovery, you can use sitemap.xml.
In case you know how it works ( and i suppose so considering your account age), your comment is just weird tbh.
Robots.txt does not fit into this argument. Im not sure why it was brought up. Google doesn’t scrape urls listed there? Ok. And so? Am I to believe that just because Google says so?
Google scrapes what it wants. It does so for its shareholders. It could care less about web standards.
Source: Amp
Interesting to see the categories that had the best responses include no-code!
This is a really interesting post! I think there's a little survivorship bias. As Product Hunt grew 2015-2017, users posted old projects of theirs which were already popular and successful.