Show HN: Link.fish – API to extract data from websites as JSON
link.fish
link.fish
A couple comments in general:
1. Personally I think its better to be great at extracting one kind of data instead of average at many types. It makes sales and growth efforts easier. Pick one of those things (products, recipes, social, etc.) and just focus on that and get great at it.
2. I don't think you need the credit <-> request abstraction. Anyone using an API knows what a request is (I hope).
Now, a few comments regarding products specifically:
1. I got 500 errors on a couple random product URLs.
2. On an Amazon product that's on sale, I got back the original price but not the sale price.
3. If you truly want to be GREAT at scraping products, the 2 things most people in this space can't do are: (a) extract ALL high-res images for a product, and (b) extract a product's options and variant data (colors, sizes, etc. and availability for each combination)
Personally I think there are a ton of opportunities in this space. This is a good start and I wish you the best.
"There is only one exception if the page should be rendered with a full browser (not headless). In this case, 5 credits get charged."
About the product.
1. It logs all requests which had issues with the more descriptive cause. Always go through all of them and fix the issues. The more people use it the more stuff breakes and the product can be improved. So I guess will get way better in the next days ;-) 2. Will also check and fix. 3. Will definitely look into that!
If you run into more issues or have more comments would love to hear them here or at api@link.fish . Thanks again!
Are they "caching" responses or that offer is tailored to your user/cookie?
Totally agree, the problem I see is that when you become big enough to be noticeable, websites start growing ban hammers at you for flagrant disregard of TOS. Working around that stuff becomes an art in of itself.
Page.REST supports extracting contents using CSS selectors. oEmbed and OpenGraph tags. Also, do you plan to support extracting from client-side rendered pages (a la React)?
BTW, I'm interested how you decided the pricing? I went with a $5 one time fee as most people use such tools for ad-hoc purposes.
General question to readers: What do you think of the Schema.org format? Is it easy to consume? (from a language library perspective)
Is actually already supported when a special parameter is set(on the API-Test-Box on the landing page it is not set).
About the pricing. Was a longer process. Mainly involved what other similar services charge and much more important a price which makes the service viable in the long term.
PS: I'd appreciate if you can become an early adopter of the service. That helps me to scale fast :)
Major differences I can see (OP feel free to correct if I'm wrong):
Link.fish
* doesn't provide a web crawler
* relies heavily on microdata, schema.org, RDFa, etc
* relies on manual parsers for sites that don't have microdata embedded
* doesn't full-render pages by default (Diffbot renders every page, so it can use computer vision to automatically extract the data)
* doesn't support proxies
* doesn't support entity tagging
Probably plenty more, but that's what jumps out to me at first blush.
--
Since I see other people have mentioned price as a concern, we're always willing to help out bootstrapped startups. Just shoot me an email: dru@diffbot.com
I'm just wondering, since a lot of people are using fairly advanced cloud-hosting solutions with, I assume, tools offered by their respective hosting place to fight spam, is web-scraping a lot different from what it used to be about a decade ago? What steps do you guys take to prevent being identified as a bad actor by the place that you are scraping?
And on the other end, if you have a data-rich website, what are your feelings toward aggressive scrapers?
For example, the exact same request will work from your home connection where it doesn't work from EC2.
A lot of small things... but basically if you load from an actual browser (headless) and cycle IPs, it's pretty hard for a site to pinpoint you as a bot vs a user.
1. Concentrate on a specific segment (like real estate) 2. Consider a browser extension (helps mitigate problem of too many requests coming from one central server)
I have long planned on building an open source real estate website scraper but just haven't found the time to do it.
I've got as far as something that works much better and is much more flexible than things like UoW OIE or just using Stanford named entity recognition.
Is this a thing others need or would find useful?
Then I tried it with the first link I got from news.google.com which was nytimes. No article text included.
Maybe I'm misunderstanding the purpose of this? Or was that just a string of bad luck?
I could see why you would want to figure out how much value the MVP is creating, and $$$ is an honest way to do that.
It sounds like two things are happening with the MVP: - emphasis on more complicated sites with more data (higher propensity to pay) - this functionality is actually possible but a user needs to take the time to set it up via a the GUI.
Feels like a pretty good tradeoff to me.
If I submit a request as a low-priority user, should I expect a response in 1 second? 1 minute? 1 hour? Something else? And how consistent is the amount of time I should expect to wait?
The question is how that will affect someone's real-world use.
Error: 500 Message: { "status": 500, "message": "Internal Server Error" }
> Additionally, do we have a growing collection of custom parsers for websites and website independent parsers for specific data.
I don't think you need the "do".
* shameless plug * Our little startup, Feedity - https://feedity.com, also helps create custom RSS feeds for any webpage.
It did actually just extract the data of the "hero" item. The thing is that it gets offered by multiple companies for different prices. So all the prices are valid and none is right or wrong. So really depends what you want. If you want simply "a" price, you can take the first. If you want the cheapest one you would have to itterate over them to find it.