798 karma · joined May 19, 2010
The article walks through how resources are provisioned, how environments are created, how GPU jobs are scheduled, and what abstractions we use to keep the system flexible while hiding most of the complexity of the underlying cluster. It also includes some of the design decisions we made along the way and a few of the tradeoffs we ran into.
Since we built the system, I’m happy to answer questions about the architecture, decisions, limitations, or areas we are still iterating on.
We get all developers at our company System 76 linux laptops as their primary dev machine (Lemur Pros with 40 GB of ram)
Hobbyists often use our free tier (30 hrs of GPU per month) or pay for more compute ($0.69/hour) to do a variety of GPU workloads.
We do prohibit things like DDOS attacks and cryptocurrency mining.
https://towardsdatascience.com/supercharging-hyperparameter-...
https://towardsdatascience.com/random-forest-on-gpus-2000x-f...
Disclaimer: We produced those benchmark. I'm a founder of Saturn (https://www.saturncloud.io/) and we focus on providing Databricks-like capabilities with Jupyter + Dask + Prefect, so I definitely have strong feelings in this area.
if isinstance(x, ndarray):
...
elif isinstance(x, other_array):
...
In the most ideal case having the standard means that scientific libraries can support all conforming implementations by default. Then sklearn would automatically support cupy/numpy/dask/jax/mxnet/pytorch/tensorflow arrays. Multiply that by all the scientific libraries and the effect is pretty profound.Saturn Cloud is DataBricks for Dask. We're building an integrated data science platform leveraging Jupyter, Prefect and Dask. Our code base is in Python, Vue/TypeScript, and lots of k8s. We also have some roles that are 50% allocated for contributing to the scientific python ecosystem (mostly Dask, but also sklearn/pandas + friends). We only have 6 engineers right now, so you'll make a big impact.
Right now we're hiring people who can build Python web backends and are somewhat comfortable building UIs. Understanding/empathy of data scientists is a plus, as is understanding of devops/aws/azure/k8s.
The team is fully remote, but we keep to similar time zones. For most regions, our salaries are quite competitive.
Our interview process: Phone Screen, a few 30 minute chats with team members (to give you a sense of what we're like, and help you figure out if you want to work here), 2 hour take home project, 2 hour pair programming session. The take home project is time boxed so that we don't burn your time.
Please apply here:
https://www.saturncloud.io/s/careers/careers-list/?gh_jid=44...
1. Cold email/linkedin message then speak to 100 potential customers. Attempt presales. You don't actually have to get presales, but that process should give you a good indication for willingness to pay / desired functionality, how people are currently working on the problem, and what language they use to describe their pain (quoting that back to customers is good marketing copy)
2. Build a landing page, run facebook / google ads, find out how much it costs to acquire an email address. Your real CAC will be higher of course, but this is a good starting point.
I'm a founder of https://www.saturncloud.io/. I started with #2, and found it cost 70 cents in advertising revenue to get a click, ~7$ for an email address, and ~$100-$200 to get someone to signup. When we started we were attempting a low touch sales model.
Since then we've started focusing on enterprise sales, because the willingness to pay is much higher. We did #1 when moving towards enterprise sales, but should have done that in the beginning.
https://www.nytimes.com/interactive/2018/10/21/business/preg...
For example, last October, a 58-year-old woman died of cardiac arrest on the warehouse floor after complaining to colleagues that she felt sick, according to a police report and current and former XPO employees. In Facebook posts at the time and in recent interviews, employees said supervisors told them to keep working as the woman lay dead.
Everyone should read Traction by Garbriel Weinberg. It's for startup marketing but job searching is basically marketing yourself in a world that has one dominant channel (recruiters). The efficacy of different channels is time varying as they get saturated- You have to find the one that works for you at any given time
JupyterHub is more flexible - for example, you could deploy JupyterHub to one beefy server and have Jupyter deployed for many users, which could all read data from a shared filesystem. that kind of thing is not easy to do with SageMaker since everything runs on a separate ec2 instance.
I can't comment on EMR notebooks.