114 karma · joined March 1, 2022
I also blog sometimes at https://april.dev/
Quiz 2 is confusingly worded but is, iiuc, referring to intranode GPU connections rather than internode networking.
Additionally, our applications consume our cloud configuration (eg something that launches pods on heterogenous GPUs needs to know which clusters support which GPUs, our colo cluster has H100s but our Google cluster has A100s etc.) Writing in the same language in the same monorepo makes it very easy to share that state.
I'm not certain but I think part of this is that XLA (for example) is a mountain of chip-specific optimizations between your code and the actual operations. So comparing your throughput between GPU and TPU is not just flops-to-flops.
I never got to interact with him directly but he was a leader worth following. This is a sad day.
Myself and a friend have been working on this for quite a while. It's frustrating because it has come so far (showing trail data worldwide interactively on a tiny budget is challenging), and yet it is so far from being something super useful like AllTrails due to low data quality and a lack of relevant features (photos, reviews, etc etc.)
jobs: https://atomic.ai/careers#opportunities, company: https://atomic.ai
Atomic AI is fusing cutting-edge machine learning and structural biology to unlock RNA drug discovery.
We're looking for sharp and thoughtful engineers excited to help us design therapeutics to treat untreatable disease. There are two positions currently open.
* Senior Software Engineer, Data - design and build solutions with our in-house experimental biology team to analyze, store, and serve experimental data and enable new analysis workflows (https://boards.greenhouse.io/atomai/jobs/4726839004)
* (junior - mid-level) Software Engineer, Infrastructure - integrate and build broad-spanning systems to create a holistic research platform (https://boards.greenhouse.io/atomai/jobs/4531035004)
I'm the hiring manager for both, please apply via the links above or email me at april@atomic.ai if you have any questions!
So I spent more than 8 years as a SWE at Google, and now work here with both experimental biologists and machine learning scientists. And yes, a lot of the concerns mentioned in this thread are also things I have had anxiety about.
Most obvious to me, being a software engineer at Google felt like being the center of the universe. Coming here, the focus is the scientific research. And yes, the scientists all managed to complete their PhDs so they don't necessarily need me to unblock them every second of their day. But contrary to my expectations, this has been remarkably freeing. I think one particularly important part of our company that makes this work is that, even on the science side, we're multidisciplinary (at a high level, emphasizing both experimental biology and ML.) And so engineering feeling like another arm of that multi-discipline nature is fairly... natural.
The reason I feel it's freeing, and the reason I enjoy working here, is also the greatest challenge. Because the scientists are focused on the science, because they respect me and trust me to figure it out, and because they aren't constantly blocked by me, my job is mostly about dreaming extremely expansively about what I can do to reduce toil and make the scientists more productive. Of course they have feedback and input, but how I use my time and what I build is ultimately my decision because I am the engineer. And I have been able to do some things I am very proud of, like rolling out Bazel and Kubernetes and finding ways to seamlessly bring them into the cloud (we're even multi-cloud now without them even noticing!) On the other hand, it's very challenging because when you work on a product, say Google Photos, as a SWE, you always have some direct tether to the product ("what should we build next? ahhhh, well I guess we could just embed stable difficusion and a million people would immediately play with it".) At Atomic, my tether is very ambiguous. If I do my job successfully, they'll be able to do research more quickly (? effectively?), and eventually we'll be able to produce a therapeutic that hopefully changes the world. Identifying what I can do today to speed up that far outcome in the future is very challenging, but it is a far more interesting challenge than gluing some pre-existing software into my UI or running A/B tests to turn a red button blue.
If, like me, you enjoy being given ownership over incredibly ambiguous problems, please do reach out!
This role focuses on directly partnering with the biologists: https://boards.greenhouse.io/atomai/jobs/4726839004
This role is expansive cloud infra: https://boards.greenhouse.io/atomai/jobs/4531035004
And this role is directly partnering with the ML scientists: https://boards.greenhouse.io/atomai/jobs/4191285004
I also miss being able to write files with /ttl=72h/ anywhere in the path and knowing they will be cleaned up automatically without me doing any more work. In contrast S3 just rots. Need to find an intern to write an SQS pipeline to schedule file deletions :)
I'm a strong apologist for Bazel, but I can absolutely see why projects avoid it. The more you deviate from what Google does the more paper cuts you'll get, but the less work you'll have to do. Bazel is great but it's certainly not a clear win in the general case.
But something fun to notice with Felt is that the rendering library (LeafletJS) is using transform3d to do 2d translations (instead of just using translate), so you may wonder "why?" At least in Chrome (and it seems in Safari) if you use transform3d the browser is more likely to keep the div in its own layer. This will reduce a lot of the paint/compositing time, and make the frame rate dramatically better. Of course in this case the micro-optimizations are irrelevant in comparison to Felt's JS performance problems (which on my machine appears to be due to projecting 2000 points from lat/lng space to pixel space every frame.) Choosing to project points every frame is super confusing because the whole point of the translate/transform3d optimization is to avoid having to recalculate pixel space during the latency-sensitive pan interaction. Odd.
The Python interpreter is also quite simple. There's several ways you can do it, but the simplest thing to imagine is if you make a launcher.py script that just invokes Bash as a foreground subprocess. The pstree is kind of funky (bash -> python -> bash), but inside that shell PYTHONPATH will be set approximately correctly. There are reasons to prefer an approach that works with sourcing (eg so you can set PS1), but it's a little harder to describe. You can do some acrobatics to make runfiles (mostly) work, and my recollection is that PATH mostly works but that may require some more work. We do the same thing for Jupyterlab and IPython.
* I have an archive Starlark function that I use for both this and containers. It sets up a folder structure similar to <target>.runfiles with everything symlinked to the actual location, then it tars the whole thing following symlinks. It has a parameter to include files that start with external/ or not.
* This archive function is used by my Bazel container rules, so I simply made a runner.py target that depends on every possible external Python dependency and made a Docker image with it.
* I then made a Bazel rule that uses the archive function to archive a given executable without external/ and uploads it to a shared location.
* At runtime runner.py is given the location as an argument, downloads it, extracts it, and then execv's it.
The major downside I've experienced is that any time you're trying to do something in a less-than-Bazel way (for example relying on binaries built outside of Bazel) things can get really hairy. My containers often need various things from apt repositories, so I had to give up on rules_docker and made my own rules for Podman. I think you need someone who understands aspects and rules before adopting it, or else the sharp edges of Bazel will keep cutting you until you drop it.
Atomic AI, Inc. (https://atomic.ai) is a well-funded, early-stage biotech company transforming the rational design of molecules and medicines through the cutting-edge fusion of artificial intelligence and structural biology. Atomic’s unique R&D platform, based on our research featured on the cover of Science (https://doi.org/10.1126/science.abe5650), provides new strategies to treat or cure previously undruggable diseases by targeting RNA structure. We are an interdisciplinary team working across computational and experimental biology and believe that our strongest asset is our people.
I am hiring a generalist Senior Software Engineer interested in truly full-stack work: from AWS resources, Kubernetes, building classical and supporting ML pipelines, all the way to client-side frontend development (though if you're in bio or ML check out the roles posted at https://boards.greenhouse.io/atomai! We haven't yet posted this SWE role.) This Senior Software Engineer role is fairly unique, we're in the early stages of trying to put together a modern cloud environment to support researchers and scientists. One of my current problems is figuring out how to build a secure environment while also giving folks the flexibility to iterate quickly on their code, maybe you know the solution :)
If you enjoy taking end-to-end ownership of problems in ambiguous areas, thinking about requirements, designing solutions, and implementing them, please reach out! april@atomic.ai