Monitor web page changes with Go
silviosimunic.com
silviosimunic.com
- Setting a User Agent (req.Header.Set(Key, Value))
- Setting Timeouts (https://blog.cloudflare.com/the-complete-guide-to-golang-net...)
- Avoiding ioutil.ReadAll() and checking body length (or using https://golang.org/pkg/io/#LimitReader)
- Caching DNS lookups (https://github.com/viki-org/dnscache)
an http Response.Body is a io.ReadCloser so you can pass it directly with `html.Parse(response.Body)`
But don't forget to defer io.Close the body.
A perusal of the awesome-go list [1] (allegedly, a "curated list", which may merit a debate) names a few:
https://github.com/parnurzeal/gorequest
https://github.com/smallnest/goreq
https://github.com/franela/goreq
https://github.com/sethgrid/pester
https://github.com/mozillazg/request
https://github.com/levigross/grequests
https://github.com/go-resty/resty
I would add my two cents:
- Calculate checksum and compare it with the latest entry (storing single 32/40-bytes long checksum takes lower amount of memory).
- Respect various HTTP headers like Expires, Cache-control, Accept-Encoding (accepting compressed content may relieve strain on bandwidth), Content-length, Cookie-length etc.
- If you were downloading a web-page and sending it via e-mail (for later usage?) you may gzip it - it may reduce costs of attachments or S3.
- Attach header "User-Agent" and "From" (with an e-mail address to contact developer) - if you don't mind.
- As you wrote, avoiding ioutil.ReadAll should be a good practice - Datadog published a blogpost which is worth reading - https://www.datadoghq.com/blog/crossing-streams-love-letter-...
- Use logrotate or implement it on your own with i.e. Mutex. Eventually you use fluentd, Kafka, Pub/Sub pipeline or some SaaS offer specialized in storing logs.
Most (or all) of Go packages with "request" in the name that I know doesn't cover aforementioned scenarios automatically - it is up to developer.
> Avoiding ioutil.ReadAll
AND
> checking body length
or do you mean,
> Avoid ioutil.ReadAll()
> and [instead you should be] checking the body length
?
Do you have an example of that, or an example of LimitReader?
Or just pass them to the request functions. And why define the global LOOP_EVERY_SECONDS in terms of seconds, rather than a time.Duration?
This code also doesn't handle errors at all. And it can exit any time from within checkChanges. And it compares data by converting the []byte to a string. Use bytes.Equal instead.
Finally, wouldn't this be a bit more useful if it printed a diff of the old-vs-new?
I like the idea. If I had to do this quickly, I'd be tempted to use cron, diff and wget as a script. I can't see any handling of HTTP response codes [0] so lots of things could happen that need handling. Does Go have equivalent to the fabulous python ^Requests^ by Kenneth Reitz? [1] Page diffs is a nice feature I might add to a page size tool I hacked up. [2]
Reference
[0] https://tools.ietf.org/html/rfc2616#section-6
I gave up and simply used Aignes Website-Watcher.[1] If you're still itching to program your own webpage change detector, I recommend reading through their 16 years of release history[2] to get an idea of the various bugs they had to solve. The scope of edge cases can be overwhelming if you want to write something very robust.
Every production site should have content (at least for any page that gets more than 5% total traffic), uptime / load time, application performance, server health, and security monitoring.
Can't count the number of times content monitoring bailed me out... a CMS change went live to early, or someone released the wrong branch to prod, or a DNS subscription I didn't control expired... I usually just string-check the Title or H1 on the page but would be great to have more advanced tools.
(1) log into the main era commons site with a username and pw
(2) then a new page opens up, on which you have to click on a particular link to get a list of your grants
(3) then click on the grant of interest
(4) then click on the competition date
(5) then the magic page opens up. It's this page I want to check for changes
I've done the requisite google searching and found lots of generics suggestions about cookies, etc, but I don't have the skills, obviously.
Is there a tool whereby I can click "record" and go through my various steps, and then the tool will keep track of the programmatic aspects of this surf-and-click dance, such that I can execute that script later (e.g. as a cron job)?
Thanks!
I think though I'll try Python/Mechanize first.
Also I'll try not being so neurotic about it and just wait like a human being for the email to be sent to me :)
For Go I don't know any.
When I was working on a project like this another issue that I found regularly is that there are regions of the page(counters, advertising, etc.) that change all the time too but they are part of the content. For this I would make several requests to the page and find the regions that would change in every request and disregard them when doing the diff.
I may not care if they change a template but would care if they change a price or a warrant canary etc?
Similarly with a library like Beautiful Soup you can traverse the HTML a lot like a jQuery selector.
[1] Yes, I'm familiar with the history of its creation, I just wish it weren't so
Your parent clearly mentions it's a tangent, and likely not a big deal. I'll go a bit further and admit that the choice of logo or mascot does indicate something about the project, whether it be a sense of style or respect for tradition or history. It's by no means the most important criteria, but it's a datum nonetheless.
Language creator hairstyles are also important signals, by the way.
In this case Go's logo has much less to do with the language than a SaaS app's landing page, but as another comment below shares, it still shows some level of influence in the overall project.
I agree that there are flaws with this as zlagen said (where the page is different in some other parts of HTML).
I've made this according to my need, and that was to check a page for only one number change, where everything else wasn't changing. This is meant to demonstrate general idea behind making something like this.
nemo1618, I agree that code should handle errors and you should definitely do that. For printing out old vs new version, as I said, for my needs, that wasn't needed since I just wanted to get emailed once one number on page changes.