Exactly. gotta read the fine print :) We think this could get more interesting with time though. It was really cool when we got our first couple days of data and were able to see the basic principles of economics at work by looking at the anti-correlation of supply with price!
Assuming you mean that it only supports GET requests against the existing data. it's amazing though that something as simple as querying a table of olympics data wasn't possible without this (without a whole bunch of manual work). if we did support transactions what would you have in mind?
Yeah, it seems the facts are the most important part. I mean the data should be openly available, right? and that lowering the barrier to use by programmers and data scientists should improve the quality of how people interact with and experience the olympics. We're hoping it will help devs build super interesting visualizations and/or apps (that otherwise weren't possible).
No. You can extract any content you want from the page (incl. text, images, links) just not meta elements or invisible ones (e.g. <title> or anything else that usually shows up in <head> for that matter)
It's unfortunately not possible right now, but we know people want it, so we're working on it. Probably have to tuck it into our advanced mode feature as it's not something we'll ever be able to solve with the point and click UI.
Although we do put the onus on the users to pay attention to robots.txt we realize the reality is that some amount of them won't necessarily respect it. That's one of the primary reasons behind the way we designed our crawler the way we did -- requiring people to actually specify the links it will visit (as opposed to spidering around sites following all links). Our hope at least is this requires people to put a little thought into the data they want (and where they want it from) before hitting a site.
Thanks for the suggestion. You're right, videos aren't the best way to comprehensively describe how to use this. We're working to put together some deeper documentation (and improving the API editor) so you'll have more control over how things come out.
Feedback from webmasters is really helpful for us. We want to make sure we're making data available via API responsibly, so would love to hear your suggestions/ thoughts as we define a scalable solution.
Thanks. We support regex matching now. Try dragging to select text, if there's a relevant regex pattern and kimono will find it (there's an example inthe blog post). You can preview (and soon, you'll be able to edit) the CSS and Regex also in advanced mode.
Kimono can handle many pages with malformed/ old HTML. Of course, it's still beta and there are still pages that break it, so we're improving it as we go with the help of early adopters like everyone on this thread. Of course, our goal is an ideal state it works everywhere perfectly :)
It would be great to automate this eventually. For now, we're trying to make it really easy to set up and rebuild the scraper. If it goes down, you'll see it in the status on your user dashboard. We're also implementing alerts, so you can opt to get an email notification if a scrape fails
Thanks, for the suggestion. We're rolling out advanced mode soon, which will allow you to edit the CSS selectors and RegEx operating on the page's HTML to define the selected data elements