HNHacker News
TopNewBestAskShowJobs

Nick-Craver

1,412 karma · joined October 16, 2013

Developer & Site Reliability Engineer for Stack Overflow

[ my public key: https://keybase.io/nickcraver; my proof: https://keybase.io/nickcraver/sigs/K_wyLPMPj5zqbqYfYcdiL1P3hcFTnAovLGJB7UR6e1g ]

submissionscomments
Nick-Craver··on Do you really need Redis? How to get away with just PostgreSQL
For pub/sub yes that's correct. For full info though: Redis later added streams (in 5.x) for the don't-wan't-to-miss case: https://redis.io/topics/streams-intro
Nick-Craver··on I think it’s time to re-wire the house
Clarification there: I have a US-8 behind each TV, taking 802.11af loads to power. For the TV with an AP or something else PoE behind it, those switches can take in 802.11at and output 802.11af on port 8. The TVs themselves aren’t PoE...I’m crazy, just not that crazy :)
Nick-Craver··on Why is Stack Overflow trying to start audio?
I’m very interested and very serious. Email sent.
Nick-Craver··on Why is Stack Overflow trying to start audio?
I don’t know. I am so very much trying to find out and push to make things better.
Nick-Craver··on Why is Stack Overflow trying to start audio?
I just wanted to chime in from Stack Overflow here and let people know: we are aware of the issue. And we're NOT okay with it. We're trying to sort out how to kill the audio behavior now. It's not very straightforward to find where it's coming from, but we are working on it. We've also reached out to Google for their assistance in tracking it down. If anyone can offer advice, we'll more than happily take it.

- Nick Craver, Architecture Lead at Stack Overflow

Nick-Craver··on HTTPS on Stack Overflow: The End of a Long Road
12,095,709 questions have an answer, 7,506,004 of those have an accepted answer, and 1,813,270 aren't yet answered.

I'd say your 1:20 ratio is just a little bit off :)

Nick-Craver··on HTTPS on Stack Overflow: The End of a Long Road
We could - but the network side isn't the problem. There's a lot of logging, user banning, etc. pieces that need IPv6 love first. We just haven't had the time yet.

There are network bits we'd have to evaluate heavily as well, e.g. firewall rules - basically the very limited benefits don't make it a priority, yet. When things change there, we'll do it.

Nick-Craver··on HTTPS on Stack Overflow: The End of a Long Road
Split horizon would point you at the same data center, rather than the writeable one. So that's more of a .local than a .internal. We discussed this, but ultimately the AD version we're on (pre-2016 Geo-DNS) it's not actually supported the way you'd need, and it's a nightmare to debug.

We'd consider it for a .local, when the support it properly there in 2016. Even subnet prioritization is busted internally, so that's a bit of an issue. Evidently no one tried to use a wildcard with dual records on 2 subnets before (we prioritize the /16, which is a data center) and it's totally busted. Microsoft has simply said this isn't supported and won't be fixed. A records work, unless they're a wildcard. So specifically, the <star>.stackexchange.com record which we mirror internally at <star>.stackexchange.com.internal for that IP set is particularly problematic.

TL;DR: Microsoft AD DNS is busted and they have no intention of fixing it. It's not worth it to try and work around it.

Nick-Craver··on HTTPS on Stack Overflow: The End of a Long Road
Well, yes and no - it depends on the length. Let's take 3 common examples. Here's GitHub's relevant headers (that we don't have):

Content-Security-Policy:default-src 'none'; base-uri 'self'; block-all-mixed-content; child-src render.githubusercontent.com; connect-src 'self' uploads.github.com status.github.com collector.githubapp.com api.github.com www.google-analytics.com github-cloud.s3.amazonaws.com github-production-repository-file-5c1aeb.s3.amazonaws.com github-production-user-asset-79cafe.s3.amazonaws.com wss://live.github.com; font-src assets-cdn.github.com; form-action 'self' github.com gist.github.com; frame-ancestors 'none'; img-src 'self' data: assets-cdn.github.com identicons.github.com collector.githubapp.com github-cloud.s3.amazonaws.com *.githubusercontent.com; media-src 'none'; script-src assets-cdn.github.com; style-src 'unsafe-inline' assets-cdn.github.com

Public-Key-Pins:max-age=5184000; pin-sha256="WoiWRyIOVNa9ihaBciRSC7XHjliYS9VwUGOIud4PB18="; pin-sha256="RRM1dGqnDFsCJXBTHky16vi1obOlCgFFn/yOhI/y+ho="; pin-sha256="k2v657xBsOVe1PQRwOsHsw3bsGT2VzIqz5K+59sNQws="; pin-sha256="K87oWBWM9UZfyddvDfoxL+8lpNyoUB2ptGtn0fv6G2Q="; pin-sha256="IQBnNBEiFuhj+8x6X8XLgh01V9Ic5/V3IRQLNFFc7v4="; pin-sha256="iie1VXtL7HzAMF+/PVPR9xzT80kQxdZeJ+zduCB3uj0="; pin-sha256="LvRiGEjRqfzurezaWuj8Wie2gyHMrW5Q06LspMnox7A="; includeSubDomains

Those are 1220 bytes. I'm not sure what they'll compress down to, but it's still non-trivial and not near 0 (anyone want to run the numbers?).

The same pair of headers are 969 bytes for facebook.com and 2,772 for gmail.com.

I don't know what ours would be - since we're open-ended on the image domain side it's a bit apples-to-oranges compared to the big players.

When you take into account that you can only send 10 packets down the first response (in almost all cases today) due to TCP congestion window specifications (google: CWND), they get more expensive as a percentage of what you can send. It may be that you can't send enough of the page to render, or the browser isn't getting to a critical stylesheet link until the second wave of packets after the ACK. This can greatly affect load times.

Does HPACK affect this? Yeah absolutely, but I disagree on "negligible". It depends, and if something critical gets pushed to that 11th packet as a result, you can drastically increase actual page render time for users.

If it helps, I did a blog post with some details about this a while back: https://nickcraver.com/blog/2015/03/24/optimization-consider...

Nick-Craver··on HTTPS on Stack Overflow: The End of a Long Road
Well if it wasn't for someone buying <star>.com back in the day, we probably could have them. Oh and then buying <star>.<star>.com after browsers banned that one, which led to RFC 6125 rule clarifications and restrictions.
Nick-Craver··on HTTPS on Stack Overflow: The End of a Long Road
Yep - we're aware. I thought about putting in our Content-Security-Policy-Report-Only findings about what all would break, but the post was already a tad long. It's quite a long list of crazy things people do.

As the headers go, here's my current thoughts on each:

- Content-Security-Policy: we're considering it, Report-Only is live on superuser.com today.

- Public-Key-Pins: we are very unlikely to deploy this. Whenever we have to change our certificates it makes life extremely dangerous for little benefit.

- X-XSS-Protection: considering it, but a lot of cross-network many-domain considerations here that most other people don't have or have as many of.

- X-Content-Type-Options: we'll likely deploy this later, there was a quirk with SVG which has passed now.

- Referrer-Policy: probably will not deploy this. We're an open book.

Nick-Craver··on How We Make Money at Stack Overflow
If curious, I did a post on that a while back. I'm settling into a new house and will pick this series back up soon.

https://nickcraver.com/blog/2016/02/17/stack-overflow-the-ar...

Nick-Craver··on Salary transparency at Stack Overflow
Stack Overflow employee here. This is just my view, but: I wouldn't want to work for Google r Facebook over Stack Overflow. Happiness is part of the compensation and is has real value in choosing where to work.

Would I make more dollars working for Google? Sure. But I couldn't work remote. And I can't make an impact like I can here. Right now I can rapidly push out changes to help millions of people and interact with them directly, every day. That's a really rare thing and something that you just can't do or do on the same level once you're part of a much more massive machine.

Does Google make an impact? Of course, you'd be crazy to argue they don't make a massive one. But I'd argue employees 1-50 were able to do a lot more to change the world with their hours in a week than employees 60,000-60,050 are able to. I want to improve people's lives, every day. I just can't do that or be as close to the result if I did it working somewhere like Google or Microsoft. There's value in these things, to me.

Nick-Craver··on Salary transparency at Stack Overflow
In the US at least (I can't speak for elsewhere), preventing this is illegal. Companies discourage it through various ways and it's practically an embedded culture thing at this point throughout the country. But: you cannot legally prevent it.

The National Labor Relations Act of 1935 is what you're looking for here - it provides employee protections for such discussions.

Nick-Craver··on Stack Overflow Outage Postmortem
Correct. I'll be adding this functionality into Opserver so that we can override the health check in HAProxy when we know better.
Nick-Craver··on Stack Overflow Outage Postmortem
Correct. 10 minutes was from checkin to all servers built out, including a dev and staging tier.
Nick-Craver··on Stack Overflow Outage Postmortem
We have an internal tool where we can dump stack traces almost instantly by attaching to a running process. We can't open source it due to using Microsoft lab code which was never licensed itself. However, clrmd (https://github.com/Microsoft/clrmd and https://www.nuget.org/packages/Microsoft.Diagnostics.Runtime) means an open source version is hopefully on the horizon. As an example, here's the result of me piddling in 2 hours with LinqPad: https://gist.github.com/NickCraver/d6292c0c7f93767686e8c5c89... Process attachment is in the clrmd roadmap, but isn't stable yet.
Nick-Craver··on Stack Overflow Outage Postmortem
Yep.
Nick-Craver··on Stack Overflow: How We Do Deployment
Yep...for the on-premise reasons listed in the article. Once upon a time a lot of projects were on Mercurial, hosted by Kiln. The Stack Overflow repo specifically has always been on an internal Mercurial and then Git server. Originally this was for speed, now it's for speed and reliability/dependency reduction.
Nick-Craver··on Stack Overflow: How We Do Deployment
I don't believe your assessment is correct. I very specifically said built-in. This remains true. If curious, we're on CentOS 7 specifically. I didn't say there aren't any options, only that there aren't any built-in. What you described as alternatives are totally true, but they still aren't built-in. It's a manual/puppet/chef/etc. config everyone has to do.

As for the applications - we have little direct input to TeamCity of Gitlab (the problem children here). And even if we did, I think we agree: the application level shouldn't cache anyway.

That being said, we're looking at `dnscache` as one of a few solutions here. But the point remains: we have to do it.

Nick-Craver··on Stack Overflow: How We Do Deployment
> With the foreign key table, performance would suffer, but probably not enough to matter for most use cases.

Citation needed :) That's going to really depend.

I'm not for or against NoSQL (or any platform). Use what's best for you and your app!

In our case, NoSQL makes for a bad database approach. We do many cross-sectional queries that cover many tables (or documents in that world). For example, a Post document doesn't make a ton of sense, we're looking at questions, answers, comments, users, and other bits across many questions all the time. The same is true of users, showing their activity for things would be very, very complicated. In our case, we're simply very relational, so an RDBMS fits the bill best.

Nick-Craver··on Stack Overflow: How We Do Deployment
It only deploys to our development/CI environment automatically. Deploying out to the production tier is a button press still.

So yes, it will build to dev, but we're using this in situations where we're very confident the changes are correct already. I'd argue blind pushes are the problem otherwise. If the developer is not very certain: they can open a merge/pull request or just hop on a hangout to do a review.

Nick-Craver··on Stack Overflow: How We Do Deployment
If you did different tables, that's even more complicated by making every query dynamic. It also makes backups, etc. far more complicated as well. Multiple databases is simply the simplest solution for multiple things that need a database with the same schema :)
Nick-Craver··on Stack Overflow: How We Do Deployment
If you mean pinbot - that's literally all it does. It takes a message and pins it, knocking the old one off the pins.

The build messages build...that's also literally all it does. It simply puts handy notices in the chatroom. Why wouldn't you want that integration? Everyone going to look at the build screen and polling it to see what's up is a far less efficient system. A push style notification, no matter the medium, causes far less overhead.

I doubt we'll ever build from chat directly for anything production at least, simply because those are 2 different user and authentication systems in play. It's too risky, IMO.

Nick-Craver··on Stack Overflow: How We Do Deployment
Not that I think it's an invalid argument, but almost all of our outages have been either a database issue or (far more often) CloudFlare's inability to reach us (the origin) across the internet. Deploying code rapidly very, very rarely causes an issue.

To us, the speed of deployment and overhead savings we get 24/7 is also absolutely worth those very rare issues.

Nick-Craver··on Stack Overflow: How We Do Deployment
Webhooks didn't used to work well for many builds off a single repo, but I think this changed very recently in TeamCity. Thanks for the reminder - I'll take another look this week at adding web hooks. We'd still want the poll in case of any hook failures.

At the moment, Gitlab knows nothing about our builds - and we'd want to keep it simple in that regard. If we can generically configure a hook to hit TeamCity to alert of any repo updates though, that's tractable...I need to see if that's possible now.

Nick-Craver··on Stack Overflow: How We Do Deployment
Yeah! What's wi...wait, what?
Nick-Craver··on Stack Overflow: How We Do Deployment
It's not for localhost, it's for the server name. While Gitlab and Teamcity normally are on the same box, they can operate on different boxes or in different data centers. It's looking up a DNS name which happens to point at the same box...does that explain it more clearly?
Nick-Craver··on Stack Overflow: How We Do Deployment
We do, but in the specific case of Active Directory, we want to fail over and auth against another data center if the primary is offline. This means for our domain, the local (to the /16) domain controllers are returned first and then the others. The problem is BIND locally doesn't preserve this order and applications are suddenly authenticating across the planet.

DNS devolution isn't a good idea here, since the external domain is a wildcard. We'll be paying for that mistake from long ago until (if ever) we change the internal domain name.

This is a pretty recent problem we're just now getting to because the DNS volume has been a back-burner issue - we'll look into permanent solutions for all Linux services after the CDN testing completes. Recommendations on the Linux DNS caching are much appreciated - we'll review each. It's something that just hasn't been an issue in the past so not experts on that particular area. I am surprised caching hasn't landed natively in most of the major distros yet though.

Nick-Craver··on Stack Overflow: How We Do Deployment
If we want a code review on anything risky, we may push a branch or we may just post the commit in chat for review before we build out. Which is chosen depends on how big or blocking the change may be.

We ask for code reviews all the time, we simply don't mandate them - I think that's the main difference.

Page 1 of 3Next →