1,737 karma · joined April 11, 2011
Concretely, the system is regularly taking checkpoints (which include model weights and optimizer state) and so if the spots disappear (as they do), the system has enough information to resume from where things were last checkpointed when resources become available again.
Full disclosure: I'm a founder of the project.
It's absolutely useful as a debugging tool to build up some semantic understanding of what the query means, and I encourage every database user to learn to use EXPLAIN, but relying on a mental execution model is borderline dangerous.
These things can happen in parallel but let’s also assume no more than 32 simultaneous TCP connections per host through a Tor proxy.
So we’re looking at ~75k1005/32 seconds = 14 days to run through all of them. You may not need to distribute this but there are situations (e.g. I want a fresh index daily) where it is warranted.
If you go by publication count alone, I suspect that IBM Research is still near the top. They have a large research organization that made some fundamental contributions in computer science in everything from speech recognition to databases.
That said, I'd expect something similar to happen with a well-written C program. Would equivalent abstractions in C++1{1,4,7} be "costly"?
Meanwhile, on Lambda, you can actually run 1600 60s jobs (or 27 CPU-hours) in 3 minutes. This is inclusive of setup time, job submission, stragglers, etc. [1]
Of course, if you've got sustained load, it's cheaper to go with spot instances, but the "occasionally I need a buttload of compute," model is well-served by Lambda.
Assuming 70 hours at $0.000000834/100ms [1]
The whole job costs $2.10.
The ricecooker makes rice and keeps it warm all day. The food processor makes quick work of chopping veggies.
You could pack the cod/chicken in vacuum bags and freeze it, and then pop it in the water bath when it's time to eat. The sous-vide means you're not worried about over-cooking/attending to the process so you can do other things while your food cooks.
Costco will sell you high quality cod and chicken at ~$10/lb (probably less in store) and steak around $20/lb. You could similarly buy rice/potatoes/fish oil there in bulk.
Admittedly you're still spending $50+ on food/day, but that's not out of the question for someone earning $150k/year.
That said, it is probably much easier when you're earning $150m/year and can afford a private chef.
The bulk of the work done in this code (in terms of FLOPS and, likely, wall-clock time) is going to be in BLAS-3 operations in the feed-forward and back-prop steps. That is, almost all of the work is done using Matrix-Matrix multiplies and in-place arithmetic/transcendental functions.
CUBLAS[1] will allow you to run these types of operations on your GPU at highly accelerated rates, without much more effort than replacing your BLAS library with a new binary. Additionally, if you want finer granularity control over what gets done on the GPU, there are other libraries[2] which provides a direct interface to CUBLAS.
[1] https://developer.nvidia.com/cublas [2] https://github.com/JuliaGPU/CUBLAS.jl
When you talk scale-out of any large system, or manycore whatever, Amdahl's law is absolutely fundamental, and when you hear someone say "perfect horizontal scalability," Amdahl's law governs when that will no longer be true.
People who do theory on distributed algorithms use it even if they don't realize it.
In the Spanner design the maximum throughput is governed by the accuracy of GPS and atomic clocks and, ultimately, the speed of light. You can only make 7 round-trips per second between a data center in New York and on in Singapore unless you're willing to bore a hole through the center of the earth.
Order is expensive and, more fundamentally, coordination is expensive. At the center of traditional OLTP workloads is a promise given by the system to the developers that their transactions will happen in some order.
Peter Bailis and Joe Hellerstein are two prominent academics working on addressing this problem, and at the core of their research is giving up pieces of this promise.
However, Philip trivializes the "10,000 hours" spent learning this stuff as merely learning how to cope with this interface. This is so far off it hurts. Those countless hours spent messing with free software is how I learned to use other people's work, compile it, read it, fix bugs in it, and learn how other people think and write software. This is an incredibly important educational experience and there's sometimes no substitute for time and patience.