2,421 karma · joined February 21, 2007
jonolson at google dot com
Opinions are my own, not those of my employer.
Antarctica was the only place that came to mind, but I wasn't sure about the costs of getting a shipping container in / out.
In particular if you have ~100PB of data to move and you are in a location that can be reached by a 45-foot long shipping container on a truck with access to the ~350KW of power it takes to run Snowmobile, you're clearly not somewhere _completely_ inaccessible. Given that 100PB of data will cost you $400k/mo in storage costs on AWS Glacier (before the discounts that you'll obviously negotiate for), even relatively remote locations become "accessible" for 100 Gbps+ fiber or microwave connectivity, and _very_ remote locations for tens of Gbps. The Snowmobile itself is only 1 Tb/s, so if you think you're going to save time this work, consider how long it takes to move it to you and then move it to an AWS facility versus the time it takes to fill it.
I don't get where this is the right call for _any_ customers, even as a one-off. I'd love it if someone from AWS could tell me where my math is off on this.
The docs are always a work in progress -- anything you can recall trying to do with GCP that was particularly inscrutable (can't fix everything in one go, but one fix at a time is better than nothing :)
I'm generally fond of our console, particularly the tools that show both the API and CLI invocations for most UI activities. On the other hand, I work on core virtualization, so the things I want to express tend to be very simple ("make me a giant VM", "make me a gianter VM", etc. :)
Any particular complaints?
(I work on Compute Engine)
Hi Matt!
So sorry about the dishes, but perhaps your dish washing efficiency will improve? If it's directly impactful to your productivity now for a single individual's daily calories it suggests you've got a lot to learn here! Perhaps a G2G course once normal business returns?
In terms of the monitor I can't help you, my most productive years were on an 11" MacBook air with a single tmux session. A 13" laptop sounds heavy.
(Disclaimer: I work at Google)
They check ID entering the liquor section (and it is a human that checks), but otherwise it is bag it and just walk out.
I'm a bit surprised by this -- it transitions from dark gray on light gray when it's off to _bright_ white on _bright_ blue when it's on. This would seem to mirror many physical things that have status lights which are not illuminated when disabled and are illuminated when enabled. What's the source of initial confusion for you?
(I work on virtualization in GCE)
Fun fact: none of these VMs had rebooted in that time, or they wouldn't have crashed.
Anyway, back in 2014 or so I dropped a bunch of transmit packet completions. In most cases I also double completed packets which was immediately fatal. Kernels get mad about that sort of thing.
Turns out, not all of the affected VMs died. Some of them lived on with head indices forever unequal to tail indices (until they rebooted).
In 2018 a developer realized there was a potential bug in waiting for VMs entering a quiescent state -- a truly idle networking stack had retired all Tx packets that it had admitted. Having unequal indices was impossible under correct operating conditions. They fixed the glitch.
This change rolled out gradually.
Gradually, the kernel panics appeared.
The change rolled back, halting the impact, but then the analysis began. What had we broken?
Another fun fact: Linux often includes an uptime in dmesg logs.
Slowly a pattern appeared. The dmesg logs included unusually large numbers for uptimes. Plotting these, there was a clear cliff in terms of a minimum uptime. Historical deployment logs showed a noteworthy release at that date, years past. Noteworthy in that it was rolled back for my bug, years prior.
On the plus side, I realized this was almost certainly my years prior fuckup slightly sooner than anyone else, so at least I got to call myself out :)
That's an extreme example, but automotive suicides that kill other passengers, drivers, or pedestrians fall into the same category. Consider also deaths from accidents involving drunk driving or fatigue -- thousands of motorists take to the roads every day modified in one manner or another that reduces their driving aptitude.
Also, while it may be correct to say that computers don't "fear death", there's no reason that "risk to self" can't be part of the criteria for decision making by an autonomous system.
We have precedent for how we qualify and evaluate things for safety: test them across a variety of conditions, accumulate driver-miles or operator-hours and incident frequencies. Then, using that data establish a bar for what constitutes an acceptable level of risk given the utility something provides. If we wanted to ensure nobody ever died in a car accident, we would ensure there were no cars, but collectively we've made a different choice.
After all, every driver on the road today is an incomprehensible black box where not only do we not know the parameters, we don't even know the function they're parameterizing. Every instance functions differently, and our testing procedures have woefully low coverage.
Today (literally) I wanted to launch Xilinx hw_server as a daemon that my peers could restart if it broke itself (it's... prone to doing so). While I could write an init script that knew about PID files, creating a systemd unit was _vastly_ easier.
Yeah, it's not Unix, but UNIX is also sort of terrible for anything that's not a one-shot, no?
The net result is a Bazel (and Blaze) that are less burdened by the baggage of legacy, but the cost is a faster treadmill to keep pace with changes.
Employment often imposes lots of restrictions on the actions we can take. I'm fortunate enough to have some choice in the set of restrictions I have to live with, so for me this particular restriction is just part of the deal.
If you read the thread you'll note that this policy has been around for a long time (at least the ~7 years I've been with Google). The author also notes that, in terms of open source, things have become dramatically _more_ permissive over time.
Or, depending on how you feel about sending a message to nil simply returning nil, Objective-C?
Swift's variant is much more modern, though, as it has the ability to inline assert.
https://docs.swift.org/swift-book/LanguageGuide/OptionalChai...
Citing a good practice in someone's code as "yes, please do more of this" alongside "don't do this please" is not, in my opinion, fluff.
> I don't want to have to go through comments that do not add any value to the exercise of finding issues, and I have never seen people leave such comments in 20 years.
So only pay attention to unresolved comments?
Not really. Positive comments help teach engineers which things they've done that conform to local best practices (and why) without them having to meticulously dig those up (assuming they're even documented). A lack of positive comments leaves engineers to learn them only by running afoul of them. Effectively it provides direction only when some threshold of badness is crossed, while leaving positive comments on good code (especially for new team members or junior engineers) provides a beacon pointing away from the badness threshold entirely.
Put differently: commenting on good code makes for swifter and less eventful code reviews by steering engineers away from the bad practices that make code challenging to review in the first place.
Also, it's just nicer to spend forty hours a week with people who demonstrably appreciate their peers' good work.
The "hardware acceleration" that helps QEMU emulate a computer operates in root VMX and non-root VMX (when actually executing guest instructions) modes. When operating in root VMX mode as a supervisor, it has approximately the highest privileges in the system (ignoring SMM), as it is operating in host ring 0. QEMU runs in host ring 3, making it a typical user mode process. If you manage to compromise QEMU (and only QEMU), you have only escaped into host user mode. In that case you haven't compromised the hypervisor, merely the virtual machine monitor.
If, on the other hand, you manage to compromise KVM (or vhost) you could reasonably claim to have actually compromised the hypervisor.
Type 2 hypervisors make this generally much uglier than type 1, but it's reasonable to say that the hypervisor in a type 2 deployment is limited to the kernel component rather than the user-mode VMM.
Developing for a small but demanding set of enterprise customers with concrete (but sometimes arcane) requirements is a very different problem from developing for consumer markets.
The latter is lower friction, in my opinion.
This is especially common for "personal" projects that people end up working on or using on the clock.
I don't know if this was the case with Cue specifically (this is the first I've heard of the project, but less boilerplate in my JSON has some appeal).