HNHacker News
TopNewBestAskShowJobs

mateuszbuda

367 karma · joined March 6, 2018

https://mateuszbuda.github.io https://scrapingfish.com https://narf.ai
submissionscomments
mateuszbuda··on Ask HN: Any hardware startups here?
Our front is software but on the backend we’ve built a mobile proxy pool and continue to work with hardware to improve it. More details on our blog: https://scrapingfish.com/blog/byo-mobile-proxy-for-web-scrap...
mateuszbuda··on The 'fuck you' pattern (2021)
That’s why people use mobile proxies which rotate IPs to scrape Instagram, Facebook, TikTok, etc: https://scrapingfish.com/blog/scraping-instagram
mateuszbuda··on Ask HN: Side project of more than $2k monthly revenue? what's your project?
Thanks! Internally we use more sophisticated system required by larger scale to support higher volume and concurrent connections from multiple clients, but for smaller scale limited to one person or a small team, everything described on our blog should be enough.
mateuszbuda··on Ask HN: Side project of more than $2k monthly revenue? what's your project?
Yes, but we have so many plans and actually many of them are unlimited. We also track data usage and can either add extra data package or a SIM card gets excluded from our proxy pool.
mateuszbuda··on Ask HN: Side project of more than $2k monthly revenue? what's your project?
Actually, it’s a feature :)
mateuszbuda··on Ask HN: Side project of more than $2k monthly revenue? what's your project?
Good question. Here we explain, with a few simplifications, how we source our IPs: https://scrapingfish.com/blog/byo-mobile-proxy-for-web-scrap...
mateuszbuda··on Ask HN: Side project of more than $2k monthly revenue? what's your project?
For https://scrapingfish.com/ it took us about 5-6 months from idea to $2k/month. It’s still not our main source of income. It was fun to build especially that it involved working with hardware.
mateuszbuda··on Experimental library for scraping websites using OpenAI's GPT API
In this particular case, GPT can help you mostly with parsing the website but not with the most challenging part of web scraping which is not getting blocked. In this case, you still need a proxy. The value from using web scraping APIs is access to a proxy pool via REST API.
mateuszbuda··on Sucralose is a negative modulator of T cell-mediated responses in mice
This may be difficult as for half of the products scraped from Walmart, sugar is the main ingredient: https://scrapingfish.com/blog/scraping-walmart
mateuszbuda··on Ask HN: What have you created that deserves a second chance on HN?
This is definitely possible. I’m not sure if we’re going to have time for this as we’re occupied by work on Scraping Fish but we shared the code for scraping nutrition facts data from Walmart on github: https://github.com/pawelkobojek/scrapingfish-blog-projects/t.... Feel free to take it and build such app/website on top of it.
mateuszbuda··on Ask HN: What have you created that deserves a second chance on HN?
Here is my post from last year which didn’t get much attention on HN: https://news.ycombinator.com/item?id=33507260

It analyzes how much sugar is in the food based on nutrition facts data scraped from Walmart. It also shows relation between amount of sugar and rating.

mateuszbuda··on Ask HN: Those making $500+/month on side projects in 2023 – Show and tell
It means that our IPs are ethically sourced. You can read more here: https://scrapingfish.com/how-ips-for-web-scraping-are-source...
mateuszbuda··on Ask HN: Those making $500+/month on side projects in 2023 – Show and tell
Scraping Fish - a web scraping API powered by custom-build, ethical, mobile proxy pool: https://scrapingfish.com/
mateuszbuda··on Ask HN: How to come up with ideas for micro side projects?
Here is my story, which may provide some inspiration.

My co-founder was searching for an apartment to purchase and found that all the tools he was using had weaknesses and none of them met his needs. We decided to create an aggregator for real estate offers. Initially, we intended to target users looking to buy an apartment, but we later pivoted to focus on real estate agencies and added features around this. After more than two years of part-time work on the project, we gave up as it was not sustainable for various reasons.

However, while working on that project, we had to develop a web scraping infrastructure to collect data reliably for our system. We realized that we could create a separate product based on this, which led to the development of Scraping Fish API for web scraping: https://scrapingfish.com. We understood the market and our potential users/customers well, as we were the users of the system before. It makes it easier to sell a product when you understand the market and your target audience.

You can read more about our journey on Indie Hackers: https://www.indiehackers.com/product/scraping-fish.

mateuszbuda··on Recent improvements to Safari
What is the best setup for Safari on mac to block ads, block YouTube ads, block trackers, block cookies popups and automatically deny cookies? I like Orion but it is simply broken on many websites, especially some forms or buttons do not work sometimes, and it is not compatible with keychain and doesn't suggest passwords. Because of this, my browser workflow is Orion for everyday browsing, for some websites I know I have to use Safari, for some I have to use Chrome, if a website doesn't work with Safari I first try Firefox and then Chrome. It seems to me like there's a room for improvement here :/
mateuszbuda··on Building an Internet Scale Meme Search Engine
Great project! With your DIY attitude, if you need to build your own infrastructure for web scraping, here is a tutorial for mobile proxy setup which might be helpful: https://scrapingfish.com/blog/byo-mobile-proxy-for-web-scrap...
mateuszbuda··on Going full time on my SaaS after 13 years
I don't want to spread negativity here but such success stories from indie hackers or side projects always remind me that most of them don't make any money: https://scrapingfish.com/blog/indie-hackers-revenue
mateuszbuda··on I am done. I give up
You’re right that your not alone in this. More than 50% of products built by indie hackers don’t make any money at all: https://scrapingfish.com/blog/indie-hackers-revenue
mateuszbuda··on Ask HN: Those with money-making side projects,how did you come up with the idea?
Interesting point. As far as I know, it's not that easy. We have a client who contacted the owners of the website he needs to scrape and offered to pay them for access to their data but they were not interested. I know that they are not their competitor and they're not doing anything that would harm their revenue in any way. I assume that they figured it was not worth the legal, organisational and other formal hassle to share their data.

Let's say that we want to pursue this idea anyway. We would have to reach our to every website owner that our clients scrape and figure out how to share the revenue. You cannot just transfer your money to another company. You need to sign a contract and in many cases have it approved by legal and compliance teams.

mateuszbuda··on Ask HN: Those with money-making side projects,how did you come up with the idea?
On our blog, we give away how to create a home-scale version of our 4G proxy: https://scrapingfish.com/blog/byo-mobile-proxy-for-web-scrap...

We have some domains blacklisted and periodically monitor logs for websites which our customers scrape but so far not a single case of illegal scraping.

We've also decided to go with a small $2 purchase instead of free trial account with no credit card required. If someone contacts us with their use case and request for a free test account, we're happy to create one for them but it has to go through us. This way we're not a good choice for people willing to do illegal stuff with our API.

mateuszbuda··on Ask HN: Those with money-making side projects,how did you come up with the idea?
A web scraping API: https://scrapingfish.com

We started with a simple web scraping solution for real estate market that was up and running in just a couple of days. We used it to track prices of apartments in our area aggregated across multiple websites.

Then, as we saw value in this, we expanded data scraping to other cities and types of properties and released the product to external users. We had a few paying customers after a couple of months.

As we wanted to include more websites to collect data from, we run into significant problems of being blocked. In result, we started investigating how to overcome different mechanisms that websites use to prevent automated traffic from web scrapers.

It turns out that one of the the most important factors is to use good quality proxy which provides IP addresses shared with other real users and change them frequently. So, we started building our own proxy infrastructure powered by 4G proxies and implemented an API on top of it. And this is how we created Scraping Fish API for web scraping.

Now, we can offer a reliable solution for scraping even the most demanding websites like Instagram or Facebook.

Here is the full story of our product on IndieHackers: https://www.indiehackers.com/product/scraping-fish

mateuszbuda··on The price of ‘sugar free’: are sweeteners as harmless as we thought?
On a related topic, here is an article which presents interesting findings based on nutrition data scraped from Walmart food products: https://scrapingfish.com/blog/scraping-walmart

It turns out that sugar is the main nutrient in half of the products, even if you exclude cookies and candy.

mateuszbuda··on Why Web Scraping Is Vital to Democracy (2020)
Web scraping is considered shady business by many people mostly because of how IP address are obtained by proxy providers. More details on this here: https://scrapingfish.com/how-ips-for-web-scraping-are-source...
mateuszbuda··on Build Your Own Mobile Proxy for Web Scraping
Scraping Fish uses puppeteer/playwright (headless browsers) connected to our mobile proxy pool under the hood.
mateuszbuda··on Build Your Own Mobile Proxy for Web Scraping
The distribution of IP addresses depends to a large extent on 4G provider. Some of them reuse IPs and you can get the same IP after you change network mode 4G > 3G > 4G. The same applies to IP addresses allocation space. You can request IP change as often as you want. Resetting network mode takes 5-10 seconds. There is not rate limiting. From time to time reconnection fails but it has nothing to do with rate limiting. Sticky IP is an issue for some 4G providers but, again, it has nothing to do with rate limiting
mateuszbuda··on Ask HN: Can I see your scripts?
I have a script for concurrent web scraping: https://github.com/mateuszbuda/webscraping-benchmark It takes a file with urls and scrapes the content. For more demanding websites it can use web scraping API that handles rotating proxies. I add some logic to process the output as needed.
mateuszbuda··on Ask HN: What are the best tools for web scraping in 2022?
If you're looking for web scraping API having more friendly pricing, especially with premium proxies and JS rendering, check out https://scrapingfish.com
mateuszbuda··on Analyzing Indie Hacker Products with Verified Revenue
This is fishy.
mateuszbuda··on Taking action against scraping for hire
In general I agree that harvesting public data is moral. I think that in these particular cases it's: 1) extracting data from profiles that opted for not being public (only available to logged in users) and 2) reposting scraped data (publicly?) as belonging to the guy who scraped it without users consent.
mateuszbuda··on Web Scraping with Python
Maybe they use their product to generate upvotes O.O
← PreviousPage 2 of 3Next →