AWS doesn't make sense for scientific computing
noahlebovic.com
noahlebovic.com
Put like for like in a well managed data center against negotiated and planned cloud services, and the former may still win, but it won't be dramatically cheaper, and figured over depreciable lifetime and including opportunity cost, may cost more. It takes work to figure out which is true.
If there was a comedy tour for IT/Programmer types, I'd pay to see you guys in it.
Best thing about your stuff is that it's literally all funny precisely because it's all true.
Well I was in the middle of that. When the Director decided to show off the new security doors. So he closed the room up. Then found out that new security doors didn’t work. I find out as I’m coming back to turn AC back on. Room will get hot really fast.
We get office Security to unlock door. He says he doesn’t have authority. His supervisor will be by later in the day.
Completely deadpan, and in front of several VPs of a forth one 50.
I turn to guy to my right who lived nearby. “Go home and get your chainsaw”
We were quickly let in. Also got fast approval to install proper cooling.
Fire extinguisher nearby, smart temp sensors, but still...
I have extinguishers all over the house, but hadn't considered a temperature sensor set to send alerts.
Do you have any recommendations?
Hell, even in DC you can look at temperatures and see in front of which server technican was standing just by those sensors.
Second cheapest would be USB-to-1wire module + some DS18B20 1-sire sensors. Easy hobby job to make. They also come with unique ID which means if you put it in TSDB by that ID it doesn't matter where you plug those sensors.
Hell, some weeks ago my oven decided the lower heater's line is now connected to ground and blows fuses...
Our 8 racks in DC only had single event of something blowing (power supply) and aside of smell and fuse blowing nothing really happened
Servers are essentially metal boxes with a bit of glass-reinforced epoxy and some plastic inside so there is limited amount of stuff that can burn. UPS is probably bigger problem
For example the OVH datacenter fire was more of "stuff around servers was flammable" (they had wooden ceilings for some reason...) rather than "just" servers.
I deal with Scientists that think AWS is some sort of a massively expensive enterprise thing. I can be, but not for the use case they're going to be embarking on. Our budget is $7M spanning 4 years.
Compared to using dedicated instances with way cheaper bandwidth, storage and compute power, it might as well be.
Cloud makes sense when you have to scale up/down very quickly, or you'd be losing money fast. But most don't suffer from this problem.
Love how you’ve fuzzed the root cause to make it seem like the “dollar store strip” is the problem and not that it was plugged into an overloaded outlet or run inside a closet at significantly elevated temperatures, leading to plastic melting and wires shorting.
Always helps to keep the “magic” a secret so the rubes have to keep us wizards employed, right?
My PhD involved quite a lot of high performance computing simulations. We kept running into problems where on warm days, all our jobs would get killed pretty consistently. Our IT guy noticed a pattern where the temps would be perfectly normal, and then suddenly out of nowhere go through the roof, triggering a thermal shutdown of the racks.
In the end our IT guy camped out in the server room on a hot day to watch what was happening. The Astrophysics' cabinet were directly infront of ours, and they had jerry rigged their cabinet door so that when it got too hot, it would swing open and their hot air would be blown all over the neighbouring cabinets...
That said, many scientists are operating on premise hardware like this: some servers in a shared rack and an el-cheapo storage solutions with an ssh access for people working in the lab. And it works just fine for them.
Cloud services focus for running business computing in a cloud, emphasizing recurring revenue. Most research labs are much more comfortable with spending the hardware portion of a grant upfront and not worrying about some student who, instead of working on some fluid dynamics problem found a script to re-train a stable diffusion and left it running over winter break. My 2c.
Until it doesn't because there's a fire or huge power surge or whatever.
That's the point -- there's a lot of risk they're not taking into account, and by focusing on the "it works just fine for them", you're cherry picking the ones that didn't suffer disaster.
When your own hardware rack goes down. You know the problem, how much it costs to fix it, and when it will come back up; usually within a few hours (or minutes) of it going down.
Do things catch fire, yes. But I think you’re over-estimating how often. In my entire life, I’ve had a single SATA connector catch fire and it just melted plastic before going out.
With AWS it's extremely easy to keep an up-to-date database backup in a different region.
And it's great that you haven't personally encountered disaster, but of course once again that's cherry-picking. And it's not just a component overheating, it's the whole closet on fire, it's a broken ceiling sprinkler system going off, it's a hurricane, it's whatever.
For the rest, there’s insurance. Most calculations done in a research setting are dependent upon that research surviving. If there’s a fire and the whole building goes down, those calculations are probably worthless now too.
Hell, most companies probably can’t survive their own building/factory burning down.
It is just as extremely easy on Hetzner or on premises
For home gamers like myself, it's has become a no brainer with advances in tunneling, docker, and cheap prices on Ebay.
An outage, or even permanent loss of hardware, might not be a big problem if you're running easily repeatable computations on data of which you have multiple copies. At worst, you might have to copy some data from an external hard drive and redo a few weeks' worth of computations.
I see it the other way: experimental scientists operate with unreliable systems all the time: fickle systems, soldered one-time setups, shared lab space, etc. Computing is just one more thing that is not 100% reliable (but way more reliable than some other equipment), and usb data sticks serve as a good enough data backup.
I'd further say that you're probably over-estimating how valuable mitigating that risk is to anyone, although there are a few limited set of customers that genuinely do care.
There are few places I can think of that would benefit more by avoiding cloud costs than scientific computing...
They often have limited budgets that are driven by grants, not derived by providing online services (computer going down does not impact bottom line).
They have real computation needs that mean hardware is unlikely to sit idle.
There is no compelling reason to "scale" in the way that a company might need to in order to handle additional unexpected load from customers or hit marketing campaigns.
Basically... the only meaningful offering from the cloud is likely preventing data loss, and this can be done fairly well with a simple backup strategy.
Again - they aren't a business where losing a few hours/days of customer data is potentially business ending.
---
And to be blunt - I can make the same risk avoidance claims about a lot of things that would simply get me laughed out of the room.
"The lead researcher shouldn't be allowed in a car because it might crash!"
"The lab work must be done in a bomb shelter in case of war or tornados!"
"No one on the team can eat red meat because it increases the risk of heart attack!"
and on and on and on... Simply saying "There's risk" is not sufficient - you must still make a compelling argument that the cost of avoiding that risk is justified, and you're not doing that.
Consider the following, I have never considered applying meltdown or spectre mitigations if it makes my code run slower because I plain don't care, assuming anyone even peeks at what my simulations doing, whoopdeedo, I don't care. I won't do that on my laptop I use to buy shit off amazon with, but the workstation I have control of? I don't care. I DO care if my simulation will take 10 days instead of a week.
My use case isn't yours because my needs aren't yours. Not everything maps across domains.
Some sort of 80/20 principle is at works here. Most of the costs in professional cloud solutions comes from making the infrastructure 99.99% reliable instead of 99% reliable. It is totally worth it if you have millions of customers that expect a certain level of reliability, but a complete overkill if the worst case scenario from a system failure is some graduate student having to redo a few days worth of computations (which probably had to be redone several times anyway because of some bug in the code or something).
I found this article: https://journals.sagepub.com/doi/10.1177/1073110520917030
Their budget for the network gear was a couple hundred bucks and some old garbage consumer grade network gear. For something that spit out 10s of GB a second (at least) across a ton of network connections (they didn't seem to know what would even happen when they ran it), and was so bursty all but the highest end of gear could handle it.
Can confirm sometimes scientists aren't really up on the overall costs. Then they dump it "this isn't working" on their university IT team to absorb the costs / manpower costs.
The article points out that this is mostly not necessary for scientific computing.
Web apps quickly become finely tuned factory machines, executing a million times a day and being duplicated thousands of times.
Scientific computing projects are often more like workshops. Getting charged by the second while you're sitting at a console trying to figure out what this giant blob you were sent even is is unpleasant. The solution you create is most likely to be run exactly once. If it is a big hit, it may be run a dozen times.
Trying to run scientific workloads on the cloud is like trying to put a human shoe on a horse. It might be possible but it's clearly not designed for that purpose.
Classic big workstations are way more capable than people think - but at the same time it's hard to justify buying one machine per user unless your department is swimming in money. Also, academic budgets tend to come in fixed chunks, and university IT departments may not have your particular group as a priority - so often it's just better to invest once in a standalone server tower that you can set up to do exactly what you need it to than try to get IT to support your needs or the accounting department to pay recurring AWS bills.
Properly provisioned linux machines need no maintenance. You drive them until there is a hardware failure.
At the other end of the spectrum there's labs [0] where the scientists need to carefully study and develop increasing skills at operating the electronics the way IT wants it done, and even worse of a distraction when there's a moving target keeping up with the IT electronics approach changing faster than the progress most labs make in their own scientific field.
What labs need more of is your kind of IT operator who can bring that option (your extreme end of the spectrum in favor of lab workers) within reach for when it is the most appropriate choice.
When labs fail to retain such adequate talent, they can rule out that option going forward, and that's one less tool in the toolbox.
[0] Including many which have good records of breakthrough progress before becoming computerized to begin with.
Running a modern AMD-based server that has 48 cores, at least 192 GB of RAM, and no included disk space costs:
~$2670.36/mo for a c5a.24xlarge AWS on-demand instance
~$1014.7/mo for a c5a.24xlarge AWS reserved instance on a three-year term, paid upfront
~$558.65/mo on OVH Cloud[1]
~$512.92/mo on Hetzner[2]
~$200/mo on your own infrastructure as a large institution[3]
Footnote [3] explains this cost estimate as:"Assumes an AMD EPYC 7552 run at 100% load in Boston with high electricity prices of $0.23/kWh, for $33.24/mo in raw power. Hardware is amortized over five years, for an average monthly price of $67.08/mo. We assume that your large institution already has 24/7 security and public internet bandwidth, but multiply base hardware and power costs by 2x to account for other hardware, cooling, physical space, and a half-a-$120k-sysadmin amortized across 100 servers."
For a large university it probably makes sense to have and manage their own compute infrastructure (cheap post-doc labor, ftw!) but for smaller outfits, AWS can make a lot of sense for scientific computing (said as someone who uses AWS for scientific computing), especially if you have fluctuating loads.
What works best IMO (and what we do) is have a minimum-to-moderate amount of compute resources in house that can satisfy the processing jobs most commonly run (and where you haven't had to overinvest in hardware), and then switch to AWS/other for heavier loads that run for a finite period.
Another problem with in-house hardware is that you spent all that money on Nvidia V100's a few years ago and now there's the A100 that blows it away, but you can't just switch and take advantage of it without another huge capital investment.
Power (especially if there is some kind of significant scientific facility on premise), space (especially in reused buildings), manpower (undergrads, grad students, post docs, professional post graduates), running old/reused hardware, etc...
You can get away with those at large research universities. Some of that you can get away with at national lab sorts of places (not going to find as much free/cheap labor, surplus hardware). If you start going down in scale/prestige, etc... none of that holds true.
Running a bunch of hardware from the surplus store in a closet somewhere with Lasko fans taped to the door is cheap. To some extent, the university system encourages such subsidies.
In any case, once you get to actually building a datacenter, if you have to factor in power, if you have a 4 hardware refresh cycle, professional staffing, etc... unless you are in one of those low CoL college towns - cloud is probably more more than 1.5 to 3x more expensive for compute (spot, etc...). Storage on prem is much cheaper - erasure coded storage systems are cheap to buy and run, and everybody wants their own high performance file system.
One continuing cloud obstacle though - researchers don't want to spend their time figuring out how to get their code friendly to preemptible VMs - which is the cost effective way to run on cloud.
Another real issue with sticking to on-prem HPC is talent acquisition and staff development. When you don't care about those things so much, it's easy to say it's cheap to run on-prem, but often the pay is crap for the required expertise, and ignoring cloud doesn't help your staff either.
Less than a year later a story made national headlines that a professor with direct access to classified material had mysteriously disappeared & never disclosed his close ties with their home country. So not only should you be worried about back doors and side doors; but also watch what's going through your front door as well!
All that being said, I have never worked at a place (or in a dept) whose threat profile made APT a real thing.
You're fortunate if APTs don't consider you a worthy target. They are no joke, and in most cases are playing a long game, more interested in penetration, persistent presence, and quiet theft of information, than in doing anything you'd notice - until they aren't. We had one who burned an asset that had been cultivated on our network for a couple of years to make a hard press play for information about a very particular high profile patient immediately after that patient had been seen (the fact that the patient was seen was public information). But with over 20,000 servers and 300,000 total nodes on the network - some of which cannot be fully patched because they run software someone has to have access to but which won't run on the newest versions of OSs or databases - you still don't know what else they've burrowed into.
A big tech company supplier to our organization had a very telling incident that ended up being detected on our side of the connection, where a brief mistake by network admins opened a channel through their layers of protection for their buid pipeline. In the minutes it was open, an APT detected the access (likely because they already owned something internal), and inserted code into their certified OS build pipeline, which we ended up with in our institution.
Externally-hosted HPC means every single compute job is seen as something that directly costs money. This negatively affects the quality of scientific output (research playfulness / creativity / focus on the research / etc).
Seen by whom? Dummies it seems. Maybe the dummies are the problem, not the computing or accounting models.
As far as HPC goes specifically, we could get some of the financial numbers to make sense in the cloud (cpu intensive jobs), but couldn't make it work for others (data intensive jobs shipping PB's around on the reg)
That and HPC has a lot of grant funding. It can be quite advantageous for an org to use an on prem data center almost like a slush fund. Can keep key projects running that would otherwise die when they are having a rough funding year.
A lot of scientific computing doesn't need a persistent data center, since you are running a ton of simulations that only take a week or so, and scientific computing centers at big universities are a big expense that isn't always well-utilized. Also, when they are full, jobs can wait weeks to run.
These computing centers have fairly high overhead, too, although some of that is absorbed by the university/nonprofit who runs them. It is entirely possible that this dynamic, where universities pay some of the cost out of your grant overhead, makes these computing centers synthetically cheaper for researchers when they are actually more expensive.
One other issue here is that scientific computing really benefits from ultra-low-latency infiniband networks, and the cloud providers offer something more similar to a virtualized RoCE system, which is a lot slower. That means accounting for cloud servers potentially being slower core-for-core.
That said, most of scientific computing (by % of total compute) happens in a different context. There's often a physical machine within the organization that's creating data (e.g. a DNA sequencer, particle accelerator, etc), and a well-maintained HPC cluster that analyzes that data. The researchers have already waited months for their data, so another couple weeks in a queue doesn't impact their cycle.
For that context, AWS doesn't really make sense. I do think there's room for a cloud provider that's geared towards an HPC use-case, and doesn't have the app-inspired limits (e.g data transfer) like AWS, GCP, and Azure.
The work won some award at SC20[3] (fka Supercomputing conference). I had considered submitting for the Gordon Bell prize, which had been specifically requesting covid work, though I thought the stuff we had done wasn't terribly sexy. We were getting ~250-500x better performance than single CPU runs.
Looking back over these, I gotta chuckle, as this (press releases) is pretty much the only time I'm called "Dr.". :D
Back to the OPs points, they are right. In most cases, cloud doesn't make sense for traditional HPC workloads. There are some special cases where it does, those tend to be large ephemeral analysis pipelines, as in bioinformatics and related fields. But for hardcore distributed (mostly MPI) code, running for a long time on a set of nodes interconnected with low latency networks, dedicated local nodes are the better economic deal.
During my stint at Cray, I was trying (quite hard) to get supercomputers, real classical ones, into cloud providers, or become a supercomputing cloud provider ourselves. The Met Office system is in Azure, is a Cray Shasta, but that was more of a special case. I couldn't get enough support for this.
Such is life. I've moved on. Still doing HPC, but more throughput maximized.
[1] https://www.uah.edu/science/departments/math/news/14954-uah-...
[2] A whole marketing writeup was done here https://www.hpe.com/us/en/newsroom/journey-to-accelerate-dru... . I tried very hard to correct the errors in the writeups. Sadly I wasn't successful.
Reminds me of a quip I made back in my SGI-Cray (1st time) days. A Cray supercomputer (back then) was a bunch of static ram that was sold, along with a free computer ... Not really true, but it gave a sense of the costs involved.
This said, Azure had (last I checked) real Mellanox networking kit for RDMA access. At Cray we placed a cluster in Azure for an end user (who shall rename nameless), and used several of Mellanox's largest switch frames for 100G Infiniband across > 1k nodes, each with many V100 GPUs. Unit would have been mid single digits on the top500 list that year.
AWS is doing their own thing network wise. Not nearly as good from a performance (latency or bandwidth) as the Mellanox kit. I don't know if Google Cloud is doing anything beyond TCP.
You can do bare metal at most/all of these. You can do some version of NVMe/local disk at all of them. Some/most let you spin up a parallel file system (network charges, so beware), either their own Lustre flavor, or one of BeeGFS, Weka, etc.
The author is making a brilliant argument for getting a secondhand workstation and shoving under their desk.
If you are doing multi machine batch style processing, then you won't be using ondemand, you'd use the spot pricing. The missing argument in that part is storage costs. Managing a high speed, highly available synchronous file system that can do a sustained 50gb/sec is hard bloody work (no S3 isnt a good fit, too much management overhead)
Don't get me wrong AWS _is_ expensive if you are using a machine for more than a month or two.
however if you are doing highly parallel stuff, Batch and lustre on demand is pretty ace.
If you are doing a multi-year project, then real steel is where its at. Assuming you have factored in hosting, storage and admin costs.
Actually, you are right. Consumer SSDs I've seen only do about 1.5GB/s sustained.
> Iceberg is a high-performance format for huge analytic tables.
How would that help speedup S3? Genuine question?
Its a huge raid-0, so long as your entire team understands that, you'll be ok. Its a lot better than in 2008, but now that AWS have a managed service, I'd just use that. (my heart is always in GPFS land...)
You could invest in an HPC - but I think the human cost of maintaining one especially if you’re in a high cost of living area (e.g. Bay Area, NYC, etc.) is going to be pretty high. Admin cost, UPS, cable wiring, heat/cooling etc. can all be pretty expensive. Maintenance of these can be pretty pricey too.
Are there any companies that remotely manage data centers and rent out bare metal infra?
Interactive analysis of large datasets (e.g. genome & exome sequencing studies with 100s of 1000s of samples) is well suited to low-latency, server-less, & horizontally scalable systems (like Dremel/BigQuery, or Hail [1], which we build and is inspired by Dremel, among other systems). The load profile is unpredictable because after a scientist runs an analysis they need an unpredictable amount of time to think about their next step.
As for productionized workflows, if we redesign the tools used within these workflows to directly read and write data to cloud storage as well as to tolerate VM-preemption, then we can exploit the ~1/5 cost of preemptible/spot instances.
One last point: for the subset of scientific computing I highlighted above, speed is key. I want the scientist to stay in a flow state, receiving feedback from their experiments as fast as possible, ideally within 300 ms. The only way to achieve that on huge datasets is through rapid and substantial scale-out followed by equally rapid and substantial scale-in (to control cost).
[1] https://hail.is
Moreover, I think we have differing definitions of “experiment”. In the context of a sequencing study, I think an “experiment” can be as simple as answering the hypothesis: does the missingness of a genotype correlate with any sample metadata (e.g. sequencing platform). You might try to test that hypothesis by looking at a PC1-PC2 plot with points colored by sequencing platform where the PCA is conducted on the 0/1 indicator matrix of missingness.
In the dry lab, that is what I mean by experiment. By that definition, a scientist does many experiments a day. Particularly for sequencing studies, these experiments are data-intensive, I need to run a simple computation on a lot of data to confirm the hypothesis.
Then the problems of data management, transfer and egress are huge. Again the "no idea what we are doing" factor comes into play. If you have a really good idea up front what is going to happen you can plan out a strategy that minimises costs. But if you have no idea at all - genuinely, because this is science and we are doing new things - then you could end up blowing huge amounts of money on unnecessary egress and storage costs. And at the small end we can be talking about experiments run on a shoestring where a few thousand dollars is a big deal.
The way I see it, we need everything - powerful individual workstations / laptops for direct analysis, then a tier of fixed HPC style compute for this kind of work that is poorly matched to cloud, and then for specific projects where it makes sense (massive scaleout, exotic hardware needs - GPUs, FPGAs etc) you embrace cloud resources for that.
Building on AWS has provided scale, security, and redundancy at a substantially lower cost than doing any on-prem solution (except for a shitty one strung together with lowendbox machines).
The combined AWS bill for the three startups is less than the cost of an F5, even on a non-inflation adjusted basis.
The cloud doesn't mean that you can be totally clueless. I've had experience in HA/scalability/redundancy/deployment/development/networking/etc. It means that if you do know what you're doing you can deliver a scalable HA solution at a ridiculously lower price point than a DIY solution using bare iron and colo.
1 Month, for sure. What about 1 year? Also did those companies required to provide any training or hiring to achieve that? Because you also need to add that to the cost comparison
If you are comparing one month bill agains 1 time purchase (which if is correctly chosen should not happen but once every 10 years at the earliest) for sure it will be cheaper. When it comes down to scalability, development and deployment, you should check your tech stack rather than your infrastructure. Kubernetees and containerization should easily take care of those with on premise hardware while also reducing complexity + you will no longer have to worry for off the chart network transit fees
On the other hand, if transitioning to a bursty cloud model means you can do your full run in hours instead of weeks, that has real impact on how many iterations you can do and often does appreciably affect velocity.
Most scientific computing still happens on supercomputers in slower moving academic or big co settings. That's the group for whom cloud computing – or at least running everything on the cloud – doesn't make sense.
All of this allows you to run code that you can't run in AWS unless it fits on one computer only. It's also way more expensive than clusters of commodity hardware.
For problems that are trivially parallelizable without much communication between nodes - I don't think that most universities can actually operate those cheaper than renting them from cloud computing services. A lot of these calculations don't take the staff to operate data centers, the cost of the building itself or the opportunity cost of using lots of space for this purpose vs something else into account. Economics of scale also kick in here. It's way cheaper per computer for AWS to admin a data center because they do this for orders of magnitude bigger data centers than your typical university.
Nah.
What is a legitimate reason for this restriction?
So my guess is it's an overly broad patch for that sort of thing.
There are other regulations to keep people from installing Steam on their ML workstations (which also cover machines below the threshold).
It's entirely about one grant-giving entity not wanting to pay for a piece of capital equipment that will have use beyond the project they're funding. It's a federal regulation, and it comes up far more commonly with lab equipment than it ever does with computers.
There are other grant mechanisms for large capital expenditures.
The problem is the thresholds haven't shifted in a long time, so you can easily trigger it with a nice workstation. But then, the budget for a modular NIH R01 was set in 1999, so thats hardly a unique problem.
For business, outsourcing compute cost center eliminates both cost and risk for a big win each quarter.
Scientists never say, Gee it isn't the holiday season, guess we better scale things back.
Instead they will always tend to push whatever compute limit there is, it is kinda in the job description.
As for the grant argument, that is letting the tool shape the hand.
business-science is not science, we will pay now or pay later.
Storage is a huge issue for us. We have a petabyte of local storage from big name vendor that's bursting at the seams, and expensive to upgrade. A lot of our users leave big files laying around for a large time. Every few months we have to hound everyone to delete old stuff.
The other thing that you get with the cloud is there's way more accountability for who's using how much resources. Right now we just let people have access and roam free. Cloud HPC is 5-10x more in cost and the beancounters would shut shit down real quick if the actual costs were divvied up.
We also still have a legacy datacenter so in a similar vein, it's hard to say how much not having to deal with physical hardware/networking/power/bandwidth would be worth. Our work is maybe 1% of what that team does.
* If you want to run really big jobs e.g. with multiple multi-GPU nodes, this might not even be possible depending on your institution or your access. Most research-intensive Universities have a cluster but they’re not normally big machines. For regional and national machines, you usually have to bid for access for specific projects, and you might not be successful.
* You have control of exactly what hardware and OS you want on your nodes. Often you’re using an out of date RHEL version and despite spack and easybuild gaining ground, all too often you’re given a compiler and some old versions of libraries and that’s it.
* For many computationally intensive studies, your data transfer actually isn’t that large. For e.g. you can often do the post-processing on-node and then only get aggregate statistics about simulation runs out.
I guess the point is also that scientists often don't realize computer costs money, when the computers are already bought
(The expensive part of these experiments is simulating billions of collisions and how the thousands of outgoing particles propagate though a detector the size of a small building. Simulating a single event takes around a minute on a modern CPU, and the experiments will simulate billions of events in a few months. If AWS is charging 5 cents a minute it works out to tens of millions easy.)
> "I buy all resources on AWS - it's painfull because I have to contact AWS almost monthly for accidental over-billing, but we don't have a solution".
All of this made me really sceptical, since coming from a big University in Germany, we have unlimited HPC resources for free. I have 16 VMs, the biggest one 125 GB Memory, I can set those up or move around how I want. No space limitations - in need 10 TB of space for 3 month? Open a service ticket, 3 hours later it's available. Ports need to be opened worldwide to the web? No problem. Need a Jupyter Hub Cluster on Kubernetes? Here you go. This has really improved my work (quality, performance, and convenience) so much.
I was once coordinator of a research project where we had 30k EUR left and didn't know what to do with it. I contacted our HPC and asked if they want the money - answer: "30k really isn't worth the effort, we don't know what to do with it atm."
Another aspect of scientific computing on the commercial cloud that’s a pain if you work in academia is procurement or paying for the cloud. Academic groups are much more comfortable with the grant model. They often operate on shoe-string budgets and are simply not comfortable entering a credit card number. You can also get commercial cloud grants, but they often lack long-term, multiyear continuity.
Where cloud/aws doesn’t make sense is storage, especially if you need egress, and if you actually need IB
And why do people only know Hetzner, OVH and Linode as alternatives to the big cloud providers?
There are so many good and inexpensive server hosting providers, some with decades of experience.
However, one of the most important differences is the lack of unrelated web services related components that pose a major distraction/headache to users that don't have a DevOps background (which AWS obviously caters to). AWS can be incredibly complicated. Simple tasks are encumbered by a whole host of unrelated options/capabilities and the learning curve is very steep. A platform that is specifically designed to serve the scientific computing audience can be much more streamlined and user-friendly for this audience.
Disclosure: I work on Paperspace.
Lambda A100s - $1.10 / hr Paperspace A100s - $3.09 / hr Genesis A100s - no A100s but their 3090 (1/2 the speed of 100) is - $1.30 / hr for half the speed
My first thought was related to colocation services. From what I understand, a lot of people avoid on-premise/in-house solutions because they don't want to deal with server rooms, redundant power, redundant networks, etc.
So people go to the cloud and pay horrendous prices there.
Why not take a middle path? Build your own custom server with your perferred hardware and put in a colocation
However, I know not all HPC/scientific computing works that way and some workloads are much more continuous.
Also, AWS is notoriously easy to undercut with on-prem hardware, especially if your budget is large and your uptime requirements aren't - you'll save a few hundred thousand a year alone by not having to hire expert engineers for on-call duty and extreme reliability.
Using the estimate from the article, spot instances are still over 8x more expensive than on-prem for scientific computing.
Personally, when I estimate the total cost of ownership of scientific cloud computing versus on prem (for extremely large-scale science with both significant server, storage, and bandwidth requirements) the cloud ends up winning for a number of reasons. I've seen a lot of academics who disagree but then I find out they use their grad students to manage their clusters.
When people claim it is 10x more expensive to use public cloud, they have no earthly idea what it actually costs to run a HPC service, a data centre, or do any of the associated maintenance.
When the claim is 3x more expensive in the cloud, they do know those things but are making a bad faith comparison because their job involves running an on-premises cluster and they are scared of losing their toys.
When the claim is 0-50% more to run in the cloud, someone is doing the math properly and aiming for a fair comparison.
When the claim is that cloud is cheaper than on-prem, you are probably talking to a cloud vendor account manager whose colleagues are wincing at the fact that they just torched their credibility.
What type of HPC do you work in? Maybe I'm over-indexing on computational biology.
It can categorically be stated that for a year's worth of CPU compute, local will always be less than Amazon. Of course, putting percentages on it doesn't work - there are just too many variables.
There are many admins out there who have no idea what an Alpha is who'll swear that if you're not buying Dell or HP hardware at a premium with expensive support contracts, you're doing things wrong and you're not a real admin. Visit Reddit's /r/sysadmin if you want to see the kind of people I'm talking about.
The point is that if people insist on the most expensive, least efficient type of servers such as Dell Xeons with ridiculous service contracts, the savings over Amazon won't be large.
It's a cumulative problem, because trying to cool and house less efficient hardware requires more power and that hardware ultimately has less tolerance for non-datacenter cooling.
Rethink things. You can have AMD Threadripper / EPYC systems in larger rooms that require less overall cooling, that have better temperature tolerance, that're more reliable in aggregate, which cost less and for which you can easily keep around spare parts which would give better turnaround and availability than support contracts from Dell / HP. Suddenly your compute costs are halved, because of pricing, efficiency, overall power, real estate considerations...
So percentages don't work, but the bottom line is that when you're doing lots of compute, over time it's always cheaper locally, even if you do things the "traditional" expensive and inefficient way, so arguing percentages with so many variables doesn't make any sense - it's still cheaper, no matter what.
That sounds very much like an argument for a cloud. Instead of waiting months to do your processing, you spin up what you need, then tear it down when you are done.
* cheap
* fast
* available
https://en.wikipedia.org/wiki/Project_management_triangleCould whatever actions have to be performed on the data also be performed in AWS?
Also while briefly looking into this I found that AWS has an egress waiver for researchers and educational instiutions: https://aws.amazon.com/blogs/publicsector/data-egress-waiver...
The other is for reproducibility - typically you need to preserve lots of in-between steps for peer review and proving that you aren't making things up. Some intermediary data is wiped out, but usually only if it can be quickly and easily regenerated.
As for leaving data in AWS, data is often (not always) revisited repeatedly for years after the fact. If new questions are raised about the results it's often much easier to check the output than rerun the analysis. And cloud storage is not cheap. But yes it sometimes makes sense to egress only summary statistics and discard the raw data.
You may use AWS yourself, but your collaborators have made their own choices for their own reasons. The data goes where it's needed, often crossing institutional boundaries.
For that, whether it makes sense or not, a lot of universities are moving to AWS, and the infrastructure cost of AWS for what would be a pretty modest server are still considerably less than the cost of complying with the policies and regulations involved in that.
Recently at my institution I asked about housing it on premise, and the answer was that IT supports AWS, and if I wanted to do something else, supporting that - as well as the responsibility for a breach - would rest entirely on my shoulders. Not doing that.
For of our more popular AWS instance types I use – a c5a.24xlarge, used for comparison in the post – the cheapest spot price over the past month in us-east-1 was $1.69. That's still $1233.70/mo: above on-prem, colo, or Hetzner pricing. Data transfer is still extremely expensive.
That said, for bursty loads that can't be smoothed with a queue, spot instances (or just normal EC2 instances) do make sense! I use them all the time for my computational biology company.
Also, if all of scientific computing switched to AWS in order to exploit spot instance pricing, I don't think those market dynamics would stay the same.
As an aside, I've had trouble created large clusters of high memory instances in us-east-2. They might have increased capacity recently, though.
A lot of scientific computing isn't happening continuously, and a lot of it is one time experiment or maybe couple of times after which you would have to tear down and reassign.
Another fun fact people forget is our ability to predict future is still pretty poor. Not only that, we are biased towards thinking we can predict it when in fact this is complete bullshit.
You have to buy and set up infrastructure before you can use it and then you have to be ready to use it. What if you are not ready? What if you will not need as much resources? What if you stop needing it earlier than you thought? When you borrow it from AWS you have flexibility to start using it when you are ready and drop it immediately when you no longer need it. Which has value on its own.
At the company I work for we found out and basically banned signing long term contracts for discounts. We found that, on average, we pay many times more for unused services than whatever we gained through discounts. Also when you pay for the resources there is incentive to improve efficiency. When you have basically prepaid for everything that incentive is very small and is basically limited to making sure you have to stay within limits.
Makes sense if the jobs are all low urgency.
We have a similar problem in trading so we have a composite solution with non-cloud simulation hardware and additional AWS hardware. That's because we have the high utilization solution combined with high urgency.
PXE booted lightweight compute nodes with a robust API, including operator portal, user portal, and cli.
Keep an eye out for the work we are doing with Triton Linux + K8s. Very lightweight Triton Linux compute node + baremetal k8s deployments on Triton.
What is the stack written in? Looks like a lot of javascript and makefiles from the github side of things but idk if that's the whole kit and caboodle.
If you are doing a big distributed-memory numerical simulation, on the other hand, you probably want infiniband I guess.
AWS seems like an OK fit for the former, maybe not great for the latter...
Our EM facility generates 10 TB of raw data per day, and once you start computing it, that increases by 30%-50% depending on what you do with it. Plus, moving between network storage and local scratch for computational steps basically never ends and keeps multiple 10 Gbe links saturated 100% of the time.
There are 2 major costs that are overlooked by the in-house crowd, which are operational maintenance cost (an increasingly rare and expensive skillset), and also the cost of downtime -- how much does it cost you when your team of data scientists are blocked because of a failed OS update etc. That being said, hiring competent people to maintain AWS properly isn't cheap either -- and it is quite easy to start running up very wasteful AWS bills on things you don't need.
As always there's a tradeoff -- the key is to choose a path and to execute it well.
- the people doing the research
- the institution's IT services group
- the administrator who writes the checks
And in my experience, "actual knowledge of what must be done and what it will or could cost" can vary greatly across these three groups; frequently in very unintuitive ways.
- Deploy a thousand servers with GPUs in 10 minutes, churn over a giant dataset, then turn them all off again. Nobody ever has to wait for access to the supercomputer.
- Automatically back up everything into cold storage over time with a lifecycle policy.
- Avoid the massive overhead of maintaining HPC clusters, labs, data centers, additional staff and training, capex, load estimation, months/years of advance planning to be ready to start computing.
- Automation via APIs to enable very quick adaptation with little coding.
- An entire universe of services which ramp up your capabilities to analyze data and apply ML without needing to build anything yourself.
- A marketplace of B2B and B2C solutions to quickly deploy new tools within your account.
- Share data with other organizations easily.
AWS costs are also "retail costs". There are massive savings to be had quite easily.
I don't control my AWS account. I don't even have an AWS account in my professional life.
I tell my IT department what I want. They tell the AWS people in central IT what they want. It's set up. At some point I get an email with login information.
I email them again to turn it off.
Do I hate this system? Yes. Is it the system I have to work with? Also yes.
"AWS as implemented by any large institution" is considerably less agile than AWS itself.
If you’ve ever worked at a big company or university (any place where you spend at scale), you’ll know you rarely pay sticker price. Software licensing is particularly elastic because it’s almost zero marginal cost. Raw cloud costs are largely a function of energy usage and amortized hardware costs — there’s a certain minimum you can’t go under but there remains a huge margin that is open to be negotiated on.
Startups/individuals rarely even think about this because they rarely qualify. But big orgs with large spends do. You can get negotiated cloud pricing.
When a company reached a certain mass, hardware cost is a factor that is considered but not a big factor.
The bigger problems are lost opportunity costs and unnecessary churns.
Businesses lose a lot when the product launch is delayed by a year simply because the hardware arrived late or have too many defects (Ask your hardware fulfillment people how many defective RAM and SSD they got per new shipment).
Churn can cost the business a lot as well. For example, imagine the model that everyone been using is trained in a Mac Pro under XYZ desk. And then when XYZ quit, s/he never properly backup the code and the model.
Bare metal allows for sloppiness that the cloud cannot afford to allow. Accountability and ownership is a lot more apparent in the cloud.
Of course, not all scientific computing workloads require a traditional supercomputer. In fact, I suspect most do not.
Can’t imagine you are paying public prices on any cloud provider if you have a $50M/yr budget.
In addition, if, as the article states, the scientists are ok to wait some considerable time for results, then one can run most, if not all, on spot instances, and that can save 10x right there.
If you don’t have $50M/yr there are companies that will move your workload around different AWS regions to get the best price - and will factor in the cost of transferring the data too.
I was architect at large scientific company using AWS.
That said, I'll repeat something that I commented somewhere else: most of scientific computing (by % of compute) happens in a context that still doesn't make sense in AWS. There's often a physical machine within the organization that's creating data (e.g. a DNA sequencer, particle accelerator, etc), and a well-maintained HPC cluster that analyzes that data.
Spot instances are still pretty expensive for a steady queue (2x of Hetzer monthly costs, for reference), and you still have to pay AWS data transfer egress costs – which are at least 30x more expensive than a colo or on-prem, if you're saturating a 1 Gbps link. Data transfer to optimize for spot instance pricing becomes prohibitive when your job has 100 TB of raw data.
Surely the raw data is input, so ingress costs, which is free?
The problem is if you have large amounts of intermediate data, and you want to transfer that somewhere else to continue analysis. Then it's "expensive". So the logical conclusion is to do all the work on the cloud, so you never have egress costs. That causes anxiety, sure. That said, 100TB costs $1000 to egress, assuming that your $50M/yr covers a Verizon 10GB/s line.
These opinions are my own and not those of my employer or former employer.
I have similar doubts about AWS for certain kinds of intensive business analysis. Not API based transactions, but back-office analysis where complex multi-join queries are run in sequence against tables with 10s of millions of records.
We do some of this with SQL servers running right on the desktop (and one still uses Excel with VLOOKUP). We have a pilot project to try these tasks in a new Azure instance. I look forward to seeing how it performs, and at what cost.
In an environment where there are not too many users and everyone is cooperative, using Google Calendar to reserve time slots works very well and is very low maintenance. Technical restrictions are needed only when the users can't be trusted to stay out of each other's way.
If a policy change is made because of this comment. I'm sorry. For sure let me know though. I'll put it on my resume.
hardware running 100% won't last five years
if hardware is not needed to be running 100% at full steam for five years, you can turn down instances on the cloud and you don't pay anything
in 2 years you'll be stuck with the same hardware, while on the cloud you follow cpu evolution as it arrives to the provider
all in all the comparison is too high level to be useful
Five year is a pretty typical amortisation schedule for HPC hardware. During my sysadmin days, of CPU, memory, cooling, power, storage, and networking, the only things that broke were hard disks and a few cooling fans. Disks were replaced by just grabbing a space and slotting it in, and fans were replaced by, well, swapping them out.
Modern CPUs and memory last a very long time. I think I remember seeing Ivy Bridge CPUs running in Hetzner servers in a video they put out, and they're still fine.
if you have spares, spares need to be in the cost, and value lost to downtime stay minimal. but you have to include spares in the expenses. if you don't have spares, 1-2 day downtime is going to be a decent hit to value.
You have a handful of nodes that the cluster can’t function without (scheduler, fileservers, etc), but you buy spares and 24x7 contracts for those nodes.
Did I misunderstand your comment?
For example, the Blue Waters supercomputer at UIUC was originally expected to last five years, although they kept it in service for nine; it was considered a success: https://www.ncsa.illinois.edu/historic-blue-waters-supercomp...
"Linkedin doesn't make sense for connecting with friends".
The fact that scientific computing has a different pattern than the typical web app is actually a good thing. If you can architect large batch jobs to use spot instances, it's 50-80% cheaper.
Also this bit: "you can keep your servers at 100% utilization by maintaining a queue of requested jobs" isn't true in practice. The pattern of research is the work normally comes in waves. You'll want to train a new model or run a number of large simulations. And then there will be periods of tweaking and work on other parts. And then more need for a lot of training. Yes, you can always find work to put on a cluster to keep it >90% utilization, but if it can be elastic (and has compute has budget attached to it), it will rise and fall.
Spot instances are still pretty expensive for a steady queue (2x of Hetzer monthly costs, for reference), and you still have to pay AWS data transfer egress costs – which are at least 30x more expensive than a colo or on-prem, if you're saturating a 1 Gbps link.
This post was born from frustration at AWS for their pricing and offerings after trying to get people to switch to AWS in scientific computing for years :)
1) A large enough queue of tasks
2) Users/donstream willing to wait
using your own infrastructure always wins (alsuming free labor) since you can load your own infrastructure to ~95% pretty much 24/7 which is unbeatable.
Otherwise completely agree, there might be some cases where the cost of labour means that you're better off running something in AWS, even if that requires someone to do the configuration as well.
- CERN started planning its computing grid before AWS was launched.
- It's pretty complicated (politics, mission, vision) for CERN to use external proprietary software/hardware for its main functions (they have even started to MS Office like products.)
- [cost] CERN is quite different than a small team researchers doing few years research. the scale is enormous and very long lived, like for decades continue
- and more...
HPC and scientific computing aside, I would have loved to be able to use AWS when I worked there, internal infra for running web apps and services wasn't nearly good & reliable, neither had a wide catalog of services offered.
This feels like a similar argument to the one made by people who use Kubernetes to ensure their web app with 100 visitors a day is web scale.
Additionally the typical workloads that run on our HPC system are often some badly maintained bioinformatics software or R/perl/pythong throwaway scripts and often enough a typo in the script causes the entire pipeline to fail after days of running on the HPC system and needs to be restarted (maybe even multiple times). Again on the on-prem system you have wasted electricity (bad enough) but in the cloud you have to pay the computing costs of the failed runs. Again cost transparency might force to fix this but the users are not software engineers.
One thing that the cloud is really good at, is elasticity and access to new hardware. We have seen for example a shift of workloads from pure CPUs to GPUs. A new CryoEM microscope was installed where the downstream analysis is relying heavily on GPUs, more and more resaerch groups run Alpafold predictions and also NGS analysis is now using GPUs. We have around 100 GPUs and average utlizations has increased to 80-90% and the users are complaining about long waiting/queueing times for their GPU jobs. For this bursting to the cloud would be nice, however GPUs are prohibitively expensive in the cloud unfortunately and the above mentioned caveats regarding job resource efficiencies still apply.
One thing that will hurt on-prem HPC systems tough are the increased electricity prices. We are now taking measures to actively save energy (i.e. by powering down idle nodes and powering them up again when jobs are scheduled). As far as I can tell the big cloud providers (AWS, etc) haven't increased the prices yet either because they cover elecriticity cost increase with their profit margins or they are not affected as much because they have better deals with elecricity providers.
First off, the pricing in the article is so disingenuous as to be outright deception.
Here is the Spot price for Azure HB120rs_v2, a popular HPC size with 120 AMD EPYC cores and 456 GB of RAM: https://azureprice.net/vm/Standard_HB120rs_v2?tier=spot&curr...
This is less than $300/month for 2.5x the compute capacity he's referencing! The author's estimate is $200/month for an on-prem server with just 48 cores. Scaled down to that level, the equivalent in cloud spot pricing would be $120.
That's assuming on-prem is 100% utilised and the cloud compute is not auto-scaled. If those assumptions are lifted, the cloud is much cheaper.
The cloud makes sense in several other ways also:
- Once the data is in cloud storage like S3 or Azure Storage Accounts, sharing it with government departments, universities, or other research institutes is trivial. Just send them a SAS URL and they can probably download it at 1GB/s without killing the Internet link at the source.
- Many of these processes have 10 GB inputs that produce about 1 TB of output due to all the intermediate and temporary files. These are often kept for later analysis, but they're of low value and go cold very quickly. Tiered storage in the cloud is very easy to set up and dirt cheap compared to on-prem network attached storage. These blobs can be moved to "Cold" storage within a few days, and then to "Archive" within a month or two at most.
- The algorithms improve over time, at which point it would be oh-so-nice to be able to re-run them over the old multi-petabyte data sets. But on-prem, this is an extravagance, and needs a lot of justification. In the cloud, you can just spin up a large pool of Spot instances with a low price cap, and let it chunk through the old data when it can. Unlike on-prem, this can read the old data in much faster, easily up to 30-100 Gbps in my tests. Good luck building a disk array that can stream 100 Gbps and also have good performance for high-priority workloads!
- The hardware is evolving much more rapidly than typical enterprise purchase cycles. We have a customer that is about to buy one (1) NVIDIA A100 GPU to use for bioinformatics. In a matter of months, it'll be superseded by the NVIDIA "Hopper" H100 series, which is 7x faster for the same genomics codes. In the cloud, both AWS and Azure will soon have instances with four H100 cards in them. That'll be 28 times faster than one A100 card, making the on-prem purchase obsolete years before the warranty runs out. A couple of years later when then successor to H100 is available in the cloud, these guys will still be using the A100!
- The cloud provides lots of peripheral services that are a PitA to set up, secure, and manage locally. For example, EKS or AKS are managed Kubernetes clusters that can be used to efficiently bin-pack HPC compute jobs and restart jobs on Spot instances if they're deallocated. Similarly, Azure CycleCloud provides managed Slurm clusters with auto-scale and spot pricing. For Docker workloads there are managed container registries, and both single-instance and scalable "container apps" that work quite well for one-off batch jobs, Jupyter notebooks, and the like.
- In the cloud, it's easy to temporarily spin up a true HPC cluster with 200 Gbps Infiniband and a matching high-performance storage cache. It's like a tiny supercomputer, rented by the hour. On-prem, just buying a single Infiniband switch will set you back more than $30K, and it'll be just the chassis. No cables, SFPs, or host adapters. A full setup is north of $100K. Good luck buying "cheap" storage that can keep up with that network!
Etc, etc...
Responses to your points below. Azure does have some better HPC infrastructure than AWS, so maybe some of my response to your comment will be wrong. I'm happy to talk about this more if you want, you can reach me using the email address in my profile.
Spot/pre-emptible instances vs. the prices in this post: large CPU/RAM instances have availability issues, especially for spot instances. I spent a lot of time trying to exploit spot instance pricing, and for standard (run command-line tool that reads in file and writes out files) bioinformatics programs, spot instances haven't made sense averaging their performance over a long period after factoring in restarts and their corresponding data transfer vs. other options for decreasing instance cost (reserved instances, negotiating, etc).
Sending S3 links: yeah, AWS makes sending data easier! Although if the destination is not within same same cloud provider (or region), you get hit with a surprising large charge for sufficiently large files.
Input size vs. output size and storing the results: generally, I agree. Cloud storage costs aren't unreasonable for S3, but I want to note how significantly the pipeline can differ. A new Illumina NovaSeq sequencer (about the size of a copy machine) with dual S4 flow cells produces 6Tb every couple days. Some pipelines are definitely inefficient, but others have more raw data. Storing that data in an infrequent access or archive tier decreases the restore speeds and increases the restore cost. If you have 100 TB of data, that increases the cost of re-running data – especially if it's in cold storage or archived.
Improved algorithms and re-running large sets of data: sure, it's a trade-off between cost and the bandwidth of a queue than you can run. For some use cases, the cloud does make sense.
Hardware improvement cycles in bioinformatics: what software are they using in bioinformatics that uses a GPU, AlphaFold? From what I've seen, most computational genomics still happens on a CPU, although fields like computational chemistry use more GPUs.
Infrastructure components and easily installable components: yeah, this is a definite value-add of cloud services, and the off-AWS/GCP/Azure analogues aren't as good yet.
Cost of networking equipment vs. by-the-hour in a cloud: yeah, if you want results quickly and occasionally, this makes sense.
Overall, this post is about most of scientific computing, not all. For this to work, you need a smoothable queue of jobs. Most computational science (by % of compute) run in this context, in universities, larger/growing co's, and government research institutions. If you want instant scalability, the math is different.
That's somewhat surprising to hear, I just assumed AWS has equivalent products. My customer is 80% AWS and 20% Azure, so it's a useful data point to know that some HPC workloads are better off in Azure.
> large CPU/RAM instances have availability issues, especially for spot instances.
I've had two different ~128 vCPU instances running for days and days in my lab environment, but that's probably because my region tends not to have a lot of HPC workloads that would compete for spot instances. I've noticed that "popular" sizes in Azure such as D4, D8, E4, and E8 are pre-empted regularly, but the "special" sizes like HPC not so much.
> surprising large charge for sufficiently large files.
Both Azure and AWS use this as a "roach motel" to encourage vendors and partners to co-locate in the same cloud. It's unfortunate that they charge on the order of $100 per terabyte, but it is what it is. However, bioinformatics files compress well, and compared to getting something out of an on-prem traditional network, the egress fees are a bargain.
My customer has a rural site with a "WAN" link. Their non-cloud option is to buy a NAS, replicate it to another NAS in a data center, and then build a permanent "file sharing solution". This is going to cost tens of thousands of dollars. They might share a few terabytes annually, which makes cloud egress fees look practically free in comparison.
> A new Illumina NovaSeq sequencer (about the size of a copy machine) with dual S4 flow cells produces 6Tb every couple days.
The scientists I talked to raised this, and to be honest I'm also concerned, especially as some of the groups I deal with are in rural areas "far from the cloud." (They analyse samples from cattle ranchers to try and prevent foot and mouth disease.)
Let's say the machine generates 6 terabytes in 2 days, so 3 TB daily. Assuming that's the uncompressed data, it'll be about 1 TB after compression. The location I'm thinking of has a 500 Mbps link, but that can transfer that data volume in just 4 hours: https://www.wolframalpha.com/input?i=%28+1+TB+%29+%2F+500+Mb...
One trick I discovered recently is that both the s3cmd and azcopy tools can take pipeline input. Combine that with a parallel compression tool like 'pigz' that can output to the pipeline, and you can have a workflow where the "raw" input files are compressed and streamed to the cloud storage at the same time. That alone can cut hours off the transfer time!
> If you have 100 TB of data, that increases the cost of re-running data – especially if it's in cold storage or archived.
Not necessarily. Azure Cold storage access has no special charges associated with it. It's a bit slower, but streaming reads were quite fast in my experience. Archive tier of course has some additional costs, but it's not a drama in most cases. For example, "high priority" retrieval costs extra, but normal priority appears to be free.
> what software are they using in bioinformatics that uses a GPU
From: https://developer.nvidia.com/blog/nvidia-hopper-architecture...
"An example is the Smith-Waterman algorithm for genomics processing".
Admittedly, GPU usage for genomics is still rare, but it is becoming more common.
> If you want instant scalability, the math is different.
In my example, think of 10-20 scientists doing semi-regular gene sequencing workloads and running related analyses. Sometimes needing a single machine with 2TB of memory, other times running 10,000 trivial jobs.
In principle, the flexibility of the cloud is nearly optimal for a scenario like this.
Sure, they can advertise that.
But are the vector units on Intel/AMD/ARM processors really not good enough?
"Debunking the 100X GPU vs. CPU Myth" is a somewhat controversial paper, but it lines up with my experience with common genomics workflows.
Data Transfer OUT From Amazon EC2 To Internet
First 10 TB / Month $0.09 per GB
Next 40 TB / Month $0.085 per GB
Next 100 TB / Month $0.07 per GB
Greater than 150 TB / Month $0.05 per GB
Which means if you transfer out 90 TB in one month, it's $0.09 * 10000 + $0.085 * 40000 + $0.07 * 40000 = $7100.Then again a part of ops cost you save is paid again in dev salary that have to deal with AWS stuff instead of just throwing a blob of binaries and letting ops worry about the rest.