I only use an iFrame to crawl and scrape content
airovic.com
airovic.com
The article only mentions malformed URLs and browser run-time errors. Were there any other "edge case limitations" that became intolerable.
that's because they have a proper iframe-embed url: https://www.youtube.com/embed/quyj70RogxI (instead of https://www.youtube.com/watch?v=quyj70RogxI )
Yes of course injecting an iframe into a third-party site with devtools isn't going to replace Selenium. But it's a clever little hack in a pinch. No need to get upset.
The site is working now, it was retrieving some localhost scripts.
I was just trying to get some feedback and to check if that document was interesting, because I was spending a lot of time on it.
For what I see, it seems Selenium does exactly the same, but I would choose this iframe solution (small-medium projects) anyway.
It's a super small tool that do the job.
Please, let me know if you can fully use airovic.com
Or is this meant to run on the dev console on the target website? In which case, the iFrame and the Airovic website doesn't make sense (the electron app mentioned does sure, but it doesn't exist)
But the airovic.com tool uses and external script, in the same server that provides http tunneling from same server, so I can modify headers and display everything on the iframe because it is under same domain.
Thanks for your comment!
That's at least what it looked like from his examples.
Just typing in and pressing the button is much easier than automating the task, so that's why the iframe is something useful. You can interact with the content (without code).
https://news.ycombinator.com/favorites?id=pcr910303
In the case login was required, using tmux it's easy to automate login to HN with a text-only browser such as links and saving the desired text/html, etc. to a file. Takes me about 500 characters of script to non-interactively log in, grab some text and log out
Okay, I'll stop speaking now and revealing the fact I started my career as a data guy at a giant corporation instead of a software engineer.
I’m very confused about this submission, and even more confused about how it managed to almost top the front page.
Edit: Having read the code samples, it seems the code snippets are supposed be run from the same origin in the dev console. A quick and dirty way to interactively scrape without navigation, I guess? Still not sure what the “all together: Airovic.com” is supposed to mean, and definitely more limited than puppeteer.
Edit2: To be fair to the author, they did say
> You cannot bypass their protections without using a HTTP Tunneling component.
Which I didn’t see until just now. This is a pretty big caveat though, should probably be more upfront...
I think it's kinda clever.
It’s still more limited than puppeteer though.
I just submitted the article to recieve some feedback. I was working a lot on the tool and the article, but needed to check if I was into something.
I fixed the errors already. it is woking.
https://developer.mozilla.org/en-US/docs/Web/HTML/Element/if...
This was useful for a brief period when I ran a news aggregator that used iFrames to display content from other news websites. Adding the sandbox attribute prevented scripts, ads, modals, etc.
For the purpose of scraping, unless you're always on the same domain(or running a proxy to add CORS), I don't see how an iFrame is better than either a web extension or a backend script using Puppeteer.
Because otherwise - since you use the dev tools to inject the iframe - you don't really need the iframe. You can just run it as a "snippet" in Chromium or from the multi-line-code-editor in Firefox.
Both have the problem that it all has to be a single file. It would be much nicer if one could import modules.
Isn't this a solved problem in javascript land? Just use a compiler/minifier and your module oriented js code is in a single file as a build artifact.
Is there some reason es modules wouldn't work here? Just a snippet that inserts a tag of type=module
$iframe.contents().find('.result-row').each(function(){
data.push({
title: $(this).find('.result-title').text(),
img: $(this).find('img').attr("src"),
price: $(this).find('.result-price:first').text()
});
// And everything starts running when you set first iframe's target url
$iframe.prop("src", "https://newyork.craigslist.org/d/apts-housing-for-rent/search/apa");
Looks like he wants output something like title:
img:
price:
I tried reproducing this example without using Javascript, instead using curl and sed. The output is image:
title:
price
I did not try to move "title:" above "image:" though I bet this could be done using the hold space.
Nor did I format this as JSON though that would be easy to do. n=0;while true;do test $n -le 3000||break;
curl https://newyork.craigslist.org/d/apts-housing-for-rent/search/apa?s=$n|sed -n '
/result-title hdrlnk/{s/.*\">/title: /;s/<.*//;/^title: /p;};
/./{/result-meta/,/\/span/{/result-price/s/.*\">/price: /;s/<.*//;/price/p;};};
/data-ids=\"/{s|1:[^,\">]*|https://images.craigslist.org/&_600x450.jpg|g;s/,/, /g;
s/1://g;s/>//;s/.*data-ids=/image: /;/^image: /p;}'
n=$((n+120));doneSadly scratchpad is going away soon. Fortunately the console now has a multiline mode, unfortunately it's not as convenient for this use.
Normally if you click a link with jQuery, you lose the current context after the next page loads.
By controlling it inside an iframe it's more convenient
It’s exactly why we’re currently pushing for the ability to disable developer tools, we want it added to Chrome and other browsers. I should be able to, as a web site owner, not allow any kind of developer tool usage.
Users do not own our product and have no right to go poking around like this!
Users own their computers and their browser (user agent) is for serving them. Not you.
You have no right to be telling a users computer exactly what to do. Do that on your own servers.
The state of tracking and telemetry is insane enough already with chrome gearing up to cut the legs out from ad blocking.
Plus, even if you're lucky enough to have your wishes with the browser, it doesn't affect anyone serious anyway. They will scrape outside of chrome, as they already do and have always done.
I feel just the same about websites and apps poking around in my stuff.
I mean, that ship sailed long ago and your energy is best invested in something else.