367 karma · joined March 6, 2018
It analyzes how much sugar is in the food based on nutrition facts data scraped from Walmart. It also shows relation between amount of sugar and rating.
My co-founder was searching for an apartment to purchase and found that all the tools he was using had weaknesses and none of them met his needs. We decided to create an aggregator for real estate offers. Initially, we intended to target users looking to buy an apartment, but we later pivoted to focus on real estate agencies and added features around this. After more than two years of part-time work on the project, we gave up as it was not sustainable for various reasons.
However, while working on that project, we had to develop a web scraping infrastructure to collect data reliably for our system. We realized that we could create a separate product based on this, which led to the development of Scraping Fish API for web scraping: https://scrapingfish.com. We understood the market and our potential users/customers well, as we were the users of the system before. It makes it easier to sell a product when you understand the market and your target audience.
You can read more about our journey on Indie Hackers: https://www.indiehackers.com/product/scraping-fish.
Let's say that we want to pursue this idea anyway. We would have to reach our to every website owner that our clients scrape and figure out how to share the revenue. You cannot just transfer your money to another company. You need to sign a contract and in many cases have it approved by legal and compliance teams.
We have some domains blacklisted and periodically monitor logs for websites which our customers scrape but so far not a single case of illegal scraping.
We've also decided to go with a small $2 purchase instead of free trial account with no credit card required. If someone contacts us with their use case and request for a free test account, we're happy to create one for them but it has to go through us. This way we're not a good choice for people willing to do illegal stuff with our API.
We started with a simple web scraping solution for real estate market that was up and running in just a couple of days. We used it to track prices of apartments in our area aggregated across multiple websites.
Then, as we saw value in this, we expanded data scraping to other cities and types of properties and released the product to external users. We had a few paying customers after a couple of months.
As we wanted to include more websites to collect data from, we run into significant problems of being blocked. In result, we started investigating how to overcome different mechanisms that websites use to prevent automated traffic from web scrapers.
It turns out that one of the the most important factors is to use good quality proxy which provides IP addresses shared with other real users and change them frequently. So, we started building our own proxy infrastructure powered by 4G proxies and implemented an API on top of it. And this is how we created Scraping Fish API for web scraping.
Now, we can offer a reliable solution for scraping even the most demanding websites like Instagram or Facebook.
Here is the full story of our product on IndieHackers: https://www.indiehackers.com/product/scraping-fish
It turns out that sugar is the main nutrient in half of the products, even if you exclude cookies and candy.