While it's not an API, Anthropic's Agent SDK does require MCP to use custom tools.
1,931 karma · joined April 2, 2019
Email is my firstname@lastname.com
Formerly at Anthropic, co-founder of Toolchest (a YC startup), and a bio startup
While it's not an API, Anthropic's Agent SDK does require MCP to use custom tools.
I think this is starting to change. Next-generation sequencers and other imaging devices are causing more wet labs to produce massive amounts of data – which is increasing the number of companies hiring for bioinformatics roles.
I don't think anyone has really figured out commercialization in the space yet – us included. The community is still rooted strongly in academia, so commercializing requires a delicate balance between profitability and openness.
I imagine it's what building dev tools was like a couple decades ago. It's fun to see the field grow and evolve.
Candidly, we haven't figured out how to do international remote work well. Hopefully we will in the future!
FlowDeploy builds dev tools for bioinformatics, and we're looking for a product-minded software engineer and a bioinformatics engineer. I think a former or future founder would do well in this role.
Curiosity matters more than domain-specific experience in bioinformatics, although some bioinformatics context is helpful: understanding if what you've built solves a problem requires talking to users and understanding them.
You would be working in a few key areas of our product:
- Improving our integration with bioinformatics pipelining languages like Nextflow and Snakemake.
- Building our core API. This is currently written with Express/Node.js with a Postgres database.
- Building the UI for launching, monitoring, and sharing bioinformatics pipelines and data. This is currently written in React with Typescript.
- Improving our pipeline execution. This is mostly in AWS Batch.
- Improving our data handling. Most raw data is stored in S3, with metadata in a Postgres database.
We're a very small team, and we plan to stay small until we have strong product-market fit. We're funded by Y Combinator, have revenue from the FlowDeploy product, and can keep going for years without raising additional funding.
Interested? Apply through YC's Work at a Startup:
- Product Engineer: https://www.ycombinator.com/companies/flowdeploy/jobs/KrwNpl...
- Bioinformatics Engineer: https://www.ycombinator.com/companies/flowdeploy/jobs/I9F9sI...
You can reach me directly at "noah" at this domain.
Some features that bridge the gap:
1. Command-line tools are often used in steps of a bioinformatics pipeline. The workflow managers expect this and make them easier to use (e.g. https://github.com/snakemake/snakemake-wrappers).
2. Using file I/O to explicitly construct a DAG is built-in, which seems easier to understand for researchers than constructing DAGs from functions.
3. Built-in support for executing on a cluster through something like SLURM.
4. Running "hacky" shell or R scripts in steps of the pipeline is well-supported. As an aside, it's surprising how often a mis-implemented subprocess.run() or os.system() call causes issues.
5. There's a strong community building open-source bioinformatics pipelines for each workflow manager (e.g. nf-core, warp, snakemake workflows).
Airflow – and the other less domain-specific workflow managers – are arguably better for people who have a stronger software engineering basis. For someone who moved wet lab to dry lab and is learning to code on the side, I think the bioinformatics workflow managers lower the barrier to entry.
Snakemake is used mostly by researchers who write code, not software engineers. Their alternative is writing scrips in bash, Python, or R; Snakemake is an easy-to-learn way to convert their scripts into a reproducible pipeline that others can use. It's popular in bioinformatics.
Snakemake also can execute remotely on a shared cluster or cloud computing. It has built-in support for common executors like SLURM, AWS, and TES[1].
Snakemake isn't perfect, but it helps researchers jump from "scripts that only work on their laptop" to "reproducible pipelines using containers" that easily run on clusters and cloud computing. Running these pipelines is still pretty quirky[2], but is better than the alternative of unmaintained and untested scripts.
There are other workflow managers further down the path of a domain-specific language, like Nextflow, WDL, or CWL. Nextflow is a dialect of Java/Groovy that is notoriously difficult to learn for researchers. Snakemake, in comparison, is built on Python and has a less steep learning curve and fewer quirks.
There are other Python based workflow managers like Prefect, Metaflow, Dagster, and Redun. They're great for software engineers, but don't bridge the gap as well with researchers-who-write-code.
[1] TES is an open standard for workflow task execution that's usable with most bioinformatics workflow managers, like HTML for browsers.
[2] I'm trying to fix this (flowdeploy.com), as are others (e.g. nf-tower). I think the quirkiness will fade over time as tooling gets better.
Seymour Hersh's stories have faced similar backlash every time they're released, including counter-statements by the US government. He has put out more dubious pieces as of late – he could be right or wrong about this – but I'd rather be exposed to his ideas than have them censored.
Anecdotally, I found the Seymour Hersh story intellectually gratifying, and was forewarned of the murkiness of its contents by the HN comments. I think all functioned pretty well on the HN side.
In the background, a new instance is spun up for every request. That means this has a high cold-start time, so is best for longer functions. I use it for computational biology (which contains some ML) batch processing.
You can also switch between cloud environments with the "provider" argument, but it's limited to just a couple options right now. We use it to cycle between AWS and our on-prem servers depending on capacity.
That said, it's possible find work that's respected and pays well. Most of that kind of work is happening in the context of startups or freelancing. My favorite example of this is Robert Edgar: he's a freelance computational biologist with over 100k citations who has made a living for the past 20 years by selling licenses to his bioinformatics software (https://scholar.google.com/citations?user=RzVMRc0AAAAJ).
To find those kinds of jobs, I'd try YC's Work at a Startup, Flagship Pioneering's portfolio companies, and emailing founders of companies that have a bioinformatics component (my email is in my profile!).
I think the issues with the field are because it's a new and growing space. We do need better tooling, respect for engineering, and established best practices, but that seems to have been the case in the past for other domains that moved from research to industry – including software engineering itself.
Responses to your points below. Azure does have some better HPC infrastructure than AWS, so maybe some of my response to your comment will be wrong. I'm happy to talk about this more if you want, you can reach me using the email address in my profile.
Spot/pre-emptible instances vs. the prices in this post: large CPU/RAM instances have availability issues, especially for spot instances. I spent a lot of time trying to exploit spot instance pricing, and for standard (run command-line tool that reads in file and writes out files) bioinformatics programs, spot instances haven't made sense averaging their performance over a long period after factoring in restarts and their corresponding data transfer vs. other options for decreasing instance cost (reserved instances, negotiating, etc).
Sending S3 links: yeah, AWS makes sending data easier! Although if the destination is not within same same cloud provider (or region), you get hit with a surprising large charge for sufficiently large files.
Input size vs. output size and storing the results: generally, I agree. Cloud storage costs aren't unreasonable for S3, but I want to note how significantly the pipeline can differ. A new Illumina NovaSeq sequencer (about the size of a copy machine) with dual S4 flow cells produces 6Tb every couple days. Some pipelines are definitely inefficient, but others have more raw data. Storing that data in an infrequent access or archive tier decreases the restore speeds and increases the restore cost. If you have 100 TB of data, that increases the cost of re-running data – especially if it's in cold storage or archived.
Improved algorithms and re-running large sets of data: sure, it's a trade-off between cost and the bandwidth of a queue than you can run. For some use cases, the cloud does make sense.
Hardware improvement cycles in bioinformatics: what software are they using in bioinformatics that uses a GPU, AlphaFold? From what I've seen, most computational genomics still happens on a CPU, although fields like computational chemistry use more GPUs.
Infrastructure components and easily installable components: yeah, this is a definite value-add of cloud services, and the off-AWS/GCP/Azure analogues aren't as good yet.
Cost of networking equipment vs. by-the-hour in a cloud: yeah, if you want results quickly and occasionally, this makes sense.
Overall, this post is about most of scientific computing, not all. For this to work, you need a smoothable queue of jobs. Most computational science (by % of compute) run in this context, in universities, larger/growing co's, and government research institutions. If you want instant scalability, the math is different.
Also, if all of scientific computing switched to AWS in order to exploit spot instance pricing, I don't think those market dynamics would stay the same.
As an aside, I've had trouble created large clusters of high memory instances in us-east-2. They might have increased capacity recently, though.
That said, I'll repeat something that I commented somewhere else: most of scientific computing (by % of compute) happens in a context that still doesn't make sense in AWS. There's often a physical machine within the organization that's creating data (e.g. a DNA sequencer, particle accelerator, etc), and a well-maintained HPC cluster that analyzes that data.
Spot instances are still pretty expensive for a steady queue (2x of Hetzer monthly costs, for reference), and you still have to pay AWS data transfer egress costs – which are at least 30x more expensive than a colo or on-prem, if you're saturating a 1 Gbps link. Data transfer to optimize for spot instance pricing becomes prohibitive when your job has 100 TB of raw data.
For of our more popular AWS instance types I use – a c5a.24xlarge, used for comparison in the post – the cheapest spot price over the past month in us-east-1 was $1.69. That's still $1233.70/mo: above on-prem, colo, or Hetzner pricing. Data transfer is still extremely expensive.
That said, for bursty loads that can't be smoothed with a queue, spot instances (or just normal EC2 instances) do make sense! I use them all the time for my computational biology company.
For example, the Blue Waters supercomputer at UIUC was originally expected to last five years, although they kept it in service for nine; it was considered a success: https://www.ncsa.illinois.edu/historic-blue-waters-supercomp...
What type of HPC do you work in? Maybe I'm over-indexing on computational biology.
Spot instances are still pretty expensive for a steady queue (2x of Hetzer monthly costs, for reference), and you still have to pay AWS data transfer egress costs – which are at least 30x more expensive than a colo or on-prem, if you're saturating a 1 Gbps link.
This post was born from frustration at AWS for their pricing and offerings after trying to get people to switch to AWS in scientific computing for years :)
That said, most of scientific computing (by % of total compute) happens in a different context. There's often a physical machine within the organization that's creating data (e.g. a DNA sequencer, particle accelerator, etc), and a well-maintained HPC cluster that analyzes that data. The researchers have already waited months for their data, so another couple weeks in a queue doesn't impact their cycle.
For that context, AWS doesn't really make sense. I do think there's room for a cloud provider that's geared towards an HPC use-case, and doesn't have the app-inspired limits (e.g data transfer) like AWS, GCP, and Azure.
Most scientific computing still happens on supercomputers in slower moving academic or big co settings. That's the group for whom cloud computing – or at least running everything on the cloud – doesn't make sense.
Using the estimate from the article, spot instances are still over 8x more expensive than on-prem for scientific computing.
Once things are containerized, it's still annoying to cycle between testing code locally on a laptop vs. running in the cloud. I hit this recently with Stable Diffusion on my laptop vs. running on an AWS GPU – I spent most of my time shifting execution between the two.
Lug is a semi-sane interface for running commands in a container and deploying briefly to the cloud: it reroutes Python system calls (like subprocess.run) to the Docker container. It pickles the function and its dependencies, so it's a single argument to switch between local and remote execution.
Quick start: https://toolchest-python-client.readthedocs.io/en/latest/
Airtable for signing up and receiving an API key: https://airtable.com/shrKzQNuDHrGkEAI2
Toolchest is an API for running data analysis tools easily (i.e. copy and paste a few lines of code), without managing the infrastructure. We're starting with computational genomics tools, but tools in other spaces can be added. Please drop me a message if you have a use case in mind! For example, I've thought about making hashcat powered by Tesla V100 GPUs accessible via our API.
All feedback is welcome! If you're curious about how it works, feel free to check out our docs: https://toolchest-python-client.readthedocs.io/en/latest/use...
Our company specializes in fast, reliable microbiome analysis using cutting-edge genomics and informatics. Our unit ensures quick, accurate, and reproducible data analysis in a secure environment for ever-increasing dynamic workloads. By facilitating fast and easy data access for our clients, we expedite advances in scientific knowledge.
You would be working in a few key areas of our code base:
- Building internal tools to help automate and record laboratory processes
- Building the API that unifies different portions of our pipeline, internal tools, and customer portal. This is currently written in Flask.
- Optimizing the pipeline for speed and reproducibility. The pipeline is comprised of Python, R, Nextflow, optimized executables and tools, with pytest in CircleCI and Jenkins for testing.
- Ensuring data security. Genomic data is personal data, so security is a top concern with regulatory requirements. Data is stored in S3, with some metadata in PostgreSQL and MongoDB databases.
Interested? We accept two forms of applications:
1. Standard HR application at https://careers-corebiome.icims.com/jobs/1106/software-engin...
2. Via POST request to https://corebiome-api.com, with a JSON body containing the string fields "first_name", "last_name", "email", "phone", "role", and a "urls" list containing any relevant links (e.g. LinkedIn, resume)
Credit to Phil Freo for the POST request idea.
Feel free to reach out at nlebovic@corebiome.com!