Editing my blog's HTTP headers with Cloudflare Workers
jvns.ca
jvns.ca
> clearing my Cloudflare cache lots of times (this would temporarily fix the problem, but it would just crop up again later)
> upgrading to a new ‘realm’ on my webhost, in the hopes that there was a bad Apache server or something that I could move away from
> Making sure <!DOCTYPE html> was at the beginning of all my HTML in case that helped browsers figure it out it was HTML (it didn’t)
> Switching away from nearlyfreespeech’s “free beta bandwidth” program
> emailing Cloudflare’s support to see if they knew anything about this
> making a lot of curl requests to my webhost directly to see if I could reproduce it (I couldn’t)
The last point in this list makes it seem like the issue could probably be with Cloudflare. Considering that she spends an additional $5 per month for the Cloudflare Workers solution now, a similar thing I would've tried would be to turn off Cloudflare DNS/caching on the site for a few weeks (or longer) and observe (this can be done by turning off caching on the Cloudflare console or by changing the DNS servers on NearlyFreeSpeech.net from Cloudflare to what NFSN provides). This would add a minor cost every month, though a lot lesser than $5, IMO. If the problem shows up, then it's certainly something on NFSN. If not, then it's certainly Cloudflare.
When thinking about cost, don't forget to factor in developer time. As a rule of thumb, your time as a developer is worth $1 per minute.
(Disclosure: I'm the tech lead of Cloudflare Workers.)
FWIW, with Cloudflare Workers you can implement logic that is (almost) strictly superior to Vary, and doesn't have these performance challenges. Instead of varying on a header, you can write code to vary on arbitrary properties of the request. For example, `Vary: User-Agent` would normally partition the cache not only for every browser, but for every minor version of every browser and every OS version, ruining your cache hit rate. With a Worker, you can parse the User-Agent and decide the buckets in your own way in code, then compute a custom cache key (or just add a query parameter to the URL) based on that. (Workers run "in front of" cache, so you get to modify the request before cache lookup occurs.)
The only down side compared to the header is that your code has to know what to vary on at request time, whereas the `Vary` header can be determined at response time. But if you're writing code specific to some application or web site, usually this is not that hard to determine based on URL alone.
(Disclosure: I'm the tech lead of Cloudflare Workers.)
I'm curious. Why can't the headers listed in the `Vary` header just be made part of the cache key?
> You either have to do a linear scan of all entries for that URL, or you need to maintain some sort of fancy index
Thanks.
I had to debug some pretty nasty issues caused by missing Vary. I would certainly do without it if it were possible, but it's not an option.
Vary is strictly required for anything that renders content per browser (User-Agent), compression (Content-Type Content-Encoding) or CORS policies (Origin).
It seems Workers is their standard answer for more flexibility now and it works well (if you're ok with the pricing) since you can create your own cache key easily by combining and hashing the different headers you're interested in and just turning them into a querystring param in the origin request.
My gut says its CloudFlare, because what I learned from CloudBleed is that they've engineered a really complex system to get good performance from a bad architecture, and complex code is prone to failure.
https://en.wikipedia.org/wiki/Cloudbleed
It might also be on the hosting end, but hosting setups are somewhat "commoditized" now, and probably less tuned for performance, so I would expect fewer exotic, unreproducible bugs.
Personally I just use Dreamhost, and my blog has been on HN many times, including the #1 spot, and it's worked fine. You don't need a lot of technology to serve static content these days. Computers are fast.
I couldn't (easily) find documentation of the webhosts' webserver setup, but I did find that Apache 2.4 deprecated DefaultType and will return pages without a Content-Type header if there are no matching rules. It seems possible that some portion on the original hosts may not be configured the same as everything else, leading to this problem.
Detailed access logs (if available) might help show where the problem request hit and help track it down?
It's probably worth poking at the host's customer support too. They seem pretty competent, and worst case, they say they're not going to look at it, because there's not enough information.
I have been hosting my minimal static website on GH Pages for a long time (and I've used it for testing many static websites) for free.
And I've been browsing blogs/websites hosted on it with no issues.
It's probably her host using a wonky server.
To use Lambda@Edge, as far as I can tell, you need to pay for:
- Lambda@Edge base cost: $0.60/M
- Lambda@Edge CPU/RAM cost: minimum $0.31/M (every request is rounded up to at least 128MB+50ms)
- CloudFront requests: minimum $1.00/M (assuming HTTPS in US+Canada)
- CloudFront egress bandwidth: minimum $1.70/M (assuming average 20KB responses in US+Canada)
- Probably other things, too?
So... $3.61 per million... and probably more due to bandwidth and region.
Cloudflare Workers charges $0.50 per million requests, with a $5 monthly minimum (covering your first 10M requests). There are no other costs: You can use it on top of Cloudflare's free plan, which gives you unlimited bandwidth. And... it performs better: https://blog.cloudflare.com/serverless-performance-compariso...
(Disclosure: Again, I'm the tech lead of Cloudflare Workers, so take my comments on competitors with appropriate grains of salt...)
{ printf 'HTTP/1.1 200 OK\r\nContent-Length: %d\r\n\r\n' "$(wc -c < a.html)"; cat a.html; } | nc -l 8080
Yet Chrome, Firefox and Safari automatically detect the payload is HTML and renders it properly without any Content-Type header, unlike with her pages. What am I missing?But if you're just doing author-created content, it should be ok security-wise.
There is no way to contact Cloudflare without the account you just lost access to.
It is very much something I wish more companies would do, or at least let users opt into. "Never under any circumstances reset my password no matter how much I might beg you to." Only tangentially related to support as a whole but more directly related to "being able to recover an account".
You might be happy with that, but most people aren't. The "contract" is: so long as I can prove who I am to a reasonable degree, give me access to my account.
The only contact method provided that doesn't require a login is to the Sales department. I searched and searched and searched and searched. I've clicked all over that website 3 or 4 times and never found a way to contact Support.
> But we'll go through a lot of hoops to make you prove that you own that account.
That's great. I'd love to provide whatever is necessary to prove the account is mine.
If one of these happens to work, I still have to hope they'll respond. Given the fact that they don't have any way to contact support on their website, it appears they aren't that fond of providing support.
I sort of got the impression that this was just a paid promotion for CloudFlare workers. If it wasn’t, maybe they could do you a solid and help you identify the actual issue. :)