HNHacker News
TopNewBestAskShowJobs

lebovic

1,931 karma · joined April 2, 2019

Noah Lebovic

Email is my firstname@lastname.com

Formerly at Anthropic, co-founder of Toolchest (a YC startup), and a bio startup

submissionscomments
lebovic··on Claude Advanced Tool Use
Yep, the Anthropic API supported tool use well before an MCP-related construct was added to the API (MCP connector in May of this year).

While it's not an API, Anthropic's Agent SDK does require MCP to use custom tools.

lebovic··on Claude 3.7 Sonnet and Claude Code
If you can reproduce the issue with the other API key, I'd also love to debug this! Feel free to share the curl -vv output (excluding the key) with the Anthropic email address in my profile
lebovic··on Ask HN: Who is hiring? (April 2024)
Will do!
lebovic··on Ask HN: Who is hiring? (April 2024)
This pattern is common. Anecdotally, I think the majority of people trained in bioinformatics end up working full-time in standard software engineering.

I think this is starting to change. Next-generation sequencers and other imaging devices are causing more wet labs to produce massive amounts of data – which is increasing the number of companies hiring for bioinformatics roles.

lebovic··on Ask HN: Who is hiring? (April 2024)
Thanks! Ironically, I was hired for my first job in bioinformatics by one of the QIIME authors. Unfortunately, that didn't make it any easier.

I don't think anyone has really figured out commercialization in the space yet – us included. The community is still rooted strongly in academia, so commercializing requires a delicate balance between profitability and openness.

I imagine it's what building dev tools was like a couple decades ago. It's fun to see the field grow and evolve.

lebovic··on Ask HN: Who is hiring? (April 2024)
Argh, this could be a perfect fit for you. I'm disappointed that we won't be able to make it work for these roles.

Candidly, we haven't figured out how to do international remote work well. Hopefully we will in the future!

lebovic··on Ask HN: Who is hiring? (April 2024)
FlowDeploy | Product Engineer, Bioinformatics Engineer| Full-time | REMOTE or ONSITE | https://flowdeploy.com

FlowDeploy builds dev tools for bioinformatics, and we're looking for a product-minded software engineer and a bioinformatics engineer. I think a former or future founder would do well in this role.

Curiosity matters more than domain-specific experience in bioinformatics, although some bioinformatics context is helpful: understanding if what you've built solves a problem requires talking to users and understanding them.

You would be working in a few key areas of our product:

- Improving our integration with bioinformatics pipelining languages like Nextflow and Snakemake.

- Building our core API. This is currently written with Express/Node.js with a Postgres database.

- Building the UI for launching, monitoring, and sharing bioinformatics pipelines and data. This is currently written in React with Typescript.

- Improving our pipeline execution. This is mostly in AWS Batch.

- Improving our data handling. Most raw data is stored in S3, with metadata in a Postgres database.

We're a very small team, and we plan to stay small until we have strong product-market fit. We're funded by Y Combinator, have revenue from the FlowDeploy product, and can keep going for years without raising additional funding.

Interested? Apply through YC's Work at a Startup:

- Product Engineer: https://www.ycombinator.com/companies/flowdeploy/jobs/KrwNpl...

- Bioinformatics Engineer: https://www.ycombinator.com/companies/flowdeploy/jobs/I9F9sI...

You can reach me directly at "noah" at this domain.

lebovic··on Snakemake – A framework for reproducible data analysis
The bioinformatics workflow managers are designed around the quirkiness of bioinformatics, and they remove a lot of boilerplate. That makes them easier to grok for someone who doesn't have a strong programming background, at the cost of some flexibility.

Some features that bridge the gap:

1. Command-line tools are often used in steps of a bioinformatics pipeline. The workflow managers expect this and make them easier to use (e.g. https://github.com/snakemake/snakemake-wrappers).

2. Using file I/O to explicitly construct a DAG is built-in, which seems easier to understand for researchers than constructing DAGs from functions.

3. Built-in support for executing on a cluster through something like SLURM.

4. Running "hacky" shell or R scripts in steps of the pipeline is well-supported. As an aside, it's surprising how often a mis-implemented subprocess.run() or os.system() call causes issues.

5. There's a strong community building open-source bioinformatics pipelines for each workflow manager (e.g. nf-core, warp, snakemake workflows).

Airflow – and the other less domain-specific workflow managers – are arguably better for people who have a stronger software engineering basis. For someone who moved wet lab to dry lab and is learning to code on the side, I think the bioinformatics workflow managers lower the barrier to entry.

lebovic··on Snakemake – A framework for reproducible data analysis
I work with Snakemake for computational biology. I see a lot of confusion as to why Snakemake exists when workflow management tools like Airflow exist, which mirrors my sentiment when moving from normal software to bio software.

Snakemake is used mostly by researchers who write code, not software engineers. Their alternative is writing scrips in bash, Python, or R; Snakemake is an easy-to-learn way to convert their scripts into a reproducible pipeline that others can use. It's popular in bioinformatics.

Snakemake also can execute remotely on a shared cluster or cloud computing. It has built-in support for common executors like SLURM, AWS, and TES[1].

Snakemake isn't perfect, but it helps researchers jump from "scripts that only work on their laptop" to "reproducible pipelines using containers" that easily run on clusters and cloud computing. Running these pipelines is still pretty quirky[2], but is better than the alternative of unmaintained and untested scripts.

There are other workflow managers further down the path of a domain-specific language, like Nextflow, WDL, or CWL. Nextflow is a dialect of Java/Groovy that is notoriously difficult to learn for researchers. Snakemake, in comparison, is built on Python and has a less steep learning curve and fewer quirks.

There are other Python based workflow managers like Prefect, Metaflow, Dagster, and Redun. They're great for software engineers, but don't bridge the gap as well with researchers-who-write-code.

[1] TES is an open standard for workflow task execution that's usable with most bioinformatics workflow managers, like HTML for browsers.

[2] I'm trying to fix this (flowdeploy.com), as are others (e.g. nf-tower). I think the quirkiness will fade over time as tooling gets better.

lebovic··on Blowing Holes in Seymour Hersh's Pipe Dream
If community moderation worked perfectly, there would be no reason to moderate. dang was clear and consistent in his moderation, even though he received the most backlash I've seen him face in the thread [1].

Seymour Hersh's stories have faced similar backlash every time they're released, including counter-statements by the US government. He has put out more dubious pieces as of late – he could be right or wrong about this – but I'd rather be exposed to his ideas than have them censored.

Anecdotally, I found the Seymour Hersh story intellectually gratifying, and was forewarned of the murkiness of its contents by the HN comments. I think all functioned pretty well on the HN side.

[1] https://news.ycombinator.com/item?id=34712496

lebovic··on Show HN: Swap between local and cloud execution with a Python decorator
Hey HN! I wanted a way to quickly switch between local and cloud execution for prototyping, so we built Lug. Notably, Lug doesn't require you to define your dependencies; it tries to extract and rebuild the same environment automatically. That means the decorator can literally just start as "@lug.hybrid(cloud=True)".

In the background, a new instance is spun up for every request. That means this has a high cold-start time, so is best for longer functions. I use it for computational biology (which contains some ML) batch processing.

You can also switch between cloud environments with the "provider" argument, but it's limited to just a couple options right now. We use it to cycle between AWS and our on-prem servers depending on capacity.

lebovic··on AWS has a low elasticity ceiling for big servers
Seems plausible. It would be awesome if AWS had a publicly accessible metric to plan around this, do you know of anything like that?
lebovic··on AWS has a low elasticity ceiling for big servers
Author here, while us-east-1 has disproportionately bad uptime that's not our biggest concern (two nines is actually okay for our use-case!). Our concern is lack of high memory instance availability, for which us-east-1 was better than us-east-2 until recently – we even switched us-east-2 -> us-east-1 for our primary deploy.
lebovic··on Consider working on genomics
I'm a software engineer who works on genomics. I see a lot of negativity in this thread, which mirrors my experience: in most places, you'll be paid like a researcher with the respect of a lab assistant – unless you have a PhD and a postdoc.

That said, it's possible find work that's respected and pays well. Most of that kind of work is happening in the context of startups or freelancing. My favorite example of this is Robert Edgar: he's a freelance computational biologist with over 100k citations who has made a living for the past 20 years by selling licenses to his bioinformatics software (https://scholar.google.com/citations?user=RzVMRc0AAAAJ).

To find those kinds of jobs, I'd try YC's Work at a Startup, Flagship Pioneering's portfolio companies, and emailing founders of companies that have a bioinformatics component (my email is in my profile!).

I think the issues with the field are because it's a new and growing space. We do need better tooling, respect for engineering, and established best practices, but that seems to have been the case in the past for other domains that moved from research to industry – including software engineering itself.

lebovic··on AWS doesn't make sense for scientific computing
Author here. It's clear that you genuinely believe my intent here is to be deceptive, so I think your comment deserves a thoughtful response. For context, I work on a computational biology infrastructure company that only uses cloud computing; my incentive is for more scientific computing to be on the cloud.

Responses to your points below. Azure does have some better HPC infrastructure than AWS, so maybe some of my response to your comment will be wrong. I'm happy to talk about this more if you want, you can reach me using the email address in my profile.

Spot/pre-emptible instances vs. the prices in this post: large CPU/RAM instances have availability issues, especially for spot instances. I spent a lot of time trying to exploit spot instance pricing, and for standard (run command-line tool that reads in file and writes out files) bioinformatics programs, spot instances haven't made sense averaging their performance over a long period after factoring in restarts and their corresponding data transfer vs. other options for decreasing instance cost (reserved instances, negotiating, etc).

Sending S3 links: yeah, AWS makes sending data easier! Although if the destination is not within same same cloud provider (or region), you get hit with a surprising large charge for sufficiently large files.

Input size vs. output size and storing the results: generally, I agree. Cloud storage costs aren't unreasonable for S3, but I want to note how significantly the pipeline can differ. A new Illumina NovaSeq sequencer (about the size of a copy machine) with dual S4 flow cells produces 6Tb every couple days. Some pipelines are definitely inefficient, but others have more raw data. Storing that data in an infrequent access or archive tier decreases the restore speeds and increases the restore cost. If you have 100 TB of data, that increases the cost of re-running data – especially if it's in cold storage or archived.

Improved algorithms and re-running large sets of data: sure, it's a trade-off between cost and the bandwidth of a queue than you can run. For some use cases, the cloud does make sense.

Hardware improvement cycles in bioinformatics: what software are they using in bioinformatics that uses a GPU, AlphaFold? From what I've seen, most computational genomics still happens on a CPU, although fields like computational chemistry use more GPUs.

Infrastructure components and easily installable components: yeah, this is a definite value-add of cloud services, and the off-AWS/GCP/Azure analogues aren't as good yet.

Cost of networking equipment vs. by-the-hour in a cloud: yeah, if you want results quickly and occasionally, this makes sense.

Overall, this post is about most of scientific computing, not all. For this to work, you need a smoothable queue of jobs. Most computational science (by % of compute) run in this context, in universities, larger/growing co's, and government research institutions. If you want instant scalability, the math is different.

lebovic··on AWS doesn't make sense for scientific computing
True! It sometimes drops even more, which definitely makes spot instances attractive. The r5.16xlarges had a ~80% discount and a <5% termination rate for while.

Also, if all of scientific computing switched to AWS in order to exploit spot instance pricing, I don't think those market dynamics would stay the same.

As an aside, I've had trouble created large clusters of high memory instances in us-east-2. They might have increased capacity recently, though.

lebovic··on AWS doesn't make sense for scientific computing
Author here. I agree that pricing is highly negotiable for any large cloud provider, and there are even (capped) egress fee waivers that you can negotiate as a part of your contract. There's also a place for using AWS; I used it for a smaller DNA sequencing facility, and I use it for my computational biology startup.

That said, I'll repeat something that I commented somewhere else: most of scientific computing (by % of compute) happens in a context that still doesn't make sense in AWS. There's often a physical machine within the organization that's creating data (e.g. a DNA sequencer, particle accelerator, etc), and a well-maintained HPC cluster that analyzes that data.

Spot instances are still pretty expensive for a steady queue (2x of Hetzer monthly costs, for reference), and you still have to pay AWS data transfer egress costs – which are at least 30x more expensive than a colo or on-prem, if you're saturating a 1 Gbps link. Data transfer to optimize for spot instance pricing becomes prohibitive when your job has 100 TB of raw data.

lebovic··on AWS doesn't make sense for scientific computing
Author here! Spot instance pricing is better than on-demand, but it doesn't include data transfer, and it's still more expensive than on-prem/Hetzner/etc. Data transfer costs exceed the cost of the instance itself if you're transferring many TB off AWS.

For of our more popular AWS instance types I use – a c5a.24xlarge, used for comparison in the post – the cheapest spot price over the past month in us-east-1 was $1.69. That's still $1233.70/mo: above on-prem, colo, or Hetzner pricing. Data transfer is still extremely expensive.

That said, for bursty loads that can't be smoothed with a queue, spot instances (or just normal EC2 instances) do make sense! I use them all the time for my computational biology company.

lebovic··on AWS doesn't make sense for scientific computing
Yep, that's right! Toolchest focuses on compute, deploying and optimizing popular scientific computing packages.
lebovic··on AWS doesn't make sense for scientific computing
In HPC land, most hardware is amortized over five years and then replaced! If you keep your service in life for five years at high utilization, you're doing great.

For example, the Blue Waters supercomputer at UIUC was originally expected to last five years, although they kept it in service for nine; it was considered a success: https://www.ncsa.illinois.edu/historic-blue-waters-supercomp...

lebovic··on AWS doesn't make sense for scientific computing
Toolchest actually runs scientific computing on AWS! I'm just frustrated by what we can build, because most scientific compute can't effectively shift to AWS
lebovic··on AWS doesn't make sense for scientific computing
Author here! I think running an HPC service that has a steady queue in AWS can be more than 3x as expensive.

What type of HPC do you work in? Maybe I'm over-indexing on computational biology.

lebovic··on AWS doesn't make sense for scientific computing
Author here! I worked for the computing infrastructure for a DNA sequencing facility, and I run a computational biology infrastructure company (trytoolchest.com, YC W22). Both are built on AWS, so I do think AWS in scientific computing has its use-cases – mostly in places where you can't saturate a queue or you want a fast cycle time.

Spot instances are still pretty expensive for a steady queue (2x of Hetzer monthly costs, for reference), and you still have to pay AWS data transfer egress costs – which are at least 30x more expensive than a colo or on-prem, if you're saturating a 1 Gbps link.

This post was born from frustration at AWS for their pricing and offerings after trying to get people to switch to AWS in scientific computing for years :)

lebovic··on AWS doesn't make sense for scientific computing
Author here. I agree with your points! I use AWS for a computational biology company I'm working on. A lot of scientific computing can spin up and down within a couple hours on AWS and benefits from fast turnaround. Most academic HPCs (by # of clusters) are slower than a mega-cluster on AWS, not well utilized, and have a lot of bureaucratic process.

That said, most of scientific computing (by % of total compute) happens in a different context. There's often a physical machine within the organization that's creating data (e.g. a DNA sequencer, particle accelerator, etc), and a well-maintained HPC cluster that analyzes that data. The researchers have already waited months for their data, so another couple weeks in a queue doesn't impact their cycle.

For that context, AWS doesn't really make sense. I do think there's room for a cloud provider that's geared towards an HPC use-case, and doesn't have the app-inspired limits (e.g data transfer) like AWS, GCP, and Azure.

lebovic··on AWS doesn't make sense for scientific computing
For fast-moving researchers who are blocked by a queue, cloud computing still makes sense. I guess I wasn't clear enough in the last section about how I still use AWS for startup-scale computational biology. My scientific computing startup (trytoolchest.com) is 100% built on top of AWS.

Most scientific computing still happens on supercomputers in slower moving academic or big co settings. That's the group for whom cloud computing – or at least running everything on the cloud – doesn't make sense.

lebovic··on AWS doesn't make sense for scientific computing
Even spot instances on AWS are still over 2x more expensive per month than Hetzner. The cheapest c5a.24xlarge spot instance right now is $1.5546/hr in us-east-1c. That's $1134.86/mo, excluding data transfer costs. If you transfer out 10 TB over the course of a month, that's another $921.60/mo – or now 4x more expensive than Hetzner.

Using the estimate from the article, spot instances are still over 8x more expensive than on-prem for scientific computing.

lebovic··on Show HN: Lug – run Python locally or in the cloud, paired with a container
I worked as a software engineer in bio and am building a company adjacent to this. There's a subset of popular software – especially in computational science and machine learning – that doesn't package well with pip or conda. For those tools, containers are actually a good solution.

Once things are containerized, it's still annoying to cycle between testing code locally on a laptop vs. running in the cloud. I hit this recently with Stable Diffusion on my laptop vs. running on an AWS GPU – I spent most of my time shifting execution between the two.

Lug is a semi-sane interface for running commands in a container and deploying briefly to the cloud: it reroutes Python system calls (like subprocess.run) to the Docker container. It pickles the function and its dependencies, so it's a single argument to switch between local and remote execution.

lebovic··on Show HN: An API for running computationally intensive tools
Thanks! Right now we only have a couple tools live, but you can run kraken2 by following the quick start and signing up for an API key. The API key comes with 10GB of free data analysis. Cost is not yet finalized, but will be on the order of ~$1/GB for a tool like kraken2.

Quick start: https://toolchest-python-client.readthedocs.io/en/latest/

Airtable for signing up and receiving an API key: https://airtable.com/shrKzQNuDHrGkEAI2

lebovic··on Show HN: An API for running computationally intensive tools
While implementing and scaling data analysis pipelines at a biotech startup, I spent most of my time getting new tools running efficiently and scaling them. Implementing something like Kraken2 for genomic analysis (https://github.com/DerrickWood/kraken2) on our infrastructure took weeks and was hard to scale. I expected a library for running these tools on managed infrastructure via an API to exist – like Twilio for sending text messages or Stripe for processing payments – but I couldn't find any.

Toolchest is an API for running data analysis tools easily (i.e. copy and paste a few lines of code), without managing the infrastructure. We're starting with computational genomics tools, but tools in other spaces can be added. Please drop me a message if you have a use case in mind! For example, I've thought about making hashcat powered by Tesla V100 GPUs accessible via our API.

All feedback is welcome! If you're curious about how it works, feel free to check out our docs: https://toolchest-python-client.readthedocs.io/en/latest/use...

lebovic··on Ask HN: Who is hiring? (July 2019)
CoreBiome | Full-stack Software Engineers | Minneapolis/St. Paul, MN | Full-Time | ONSITE or REMOTE | https://corebiome.com

Our company specializes in fast, reliable microbiome analysis using cutting-edge genomics and informatics. Our unit ensures quick, accurate, and reproducible data analysis in a secure environment for ever-increasing dynamic workloads. By facilitating fast and easy data access for our clients, we expedite advances in scientific knowledge.

You would be working in a few key areas of our code base:

- Building internal tools to help automate and record laboratory processes

- Building the API that unifies different portions of our pipeline, internal tools, and customer portal. This is currently written in Flask.

- Optimizing the pipeline for speed and reproducibility. The pipeline is comprised of Python, R, Nextflow, optimized executables and tools, with pytest in CircleCI and Jenkins for testing.

- Ensuring data security. Genomic data is personal data, so security is a top concern with regulatory requirements. Data is stored in S3, with some metadata in PostgreSQL and MongoDB databases.

Interested? We accept two forms of applications:

1. Standard HR application at https://careers-corebiome.icims.com/jobs/1106/software-engin...

2. Via POST request to https://corebiome-api.com, with a JSON body containing the string fields "first_name", "last_name", "email", "phone", "role", and a "urls" list containing any relevant links (e.g. LinkedIn, resume)

Credit to Phil Freo for the POST request idea.

Feel free to reach out at nlebovic@corebiome.com!

← PreviousPage 3 of 4Next →