Calculating the cost of a Google DeepMind paper
152334h.github.io
152334h.github.io
My wife works on high-throughout drug screens. They routinely use over $100,000 of consumables in a single screen, not counting the cost of the screening “libraries”, the cost of using some of the -$10mil of equipment in the lab for several weeks, the cost of the staff in the lab itself, and the cost of the time of the scientists who request the screens and then take the results and turn them into papers.
My comment was a bit tongue in cheek. Not every research is going to be profitable or eventually profitable, but that also doesn't mean it isn't useful. If we're willing to account for the indirect profits via learning what doesn't work (an important part of science), then this vastly diminishes the number of worthless papers (to essentially those that are fraudulent or plagiarism)
But as specifically for AlphaFold, I'm going to need a citation on that. If I understand the calculus correctly, Google acquired DM in 2014 for somewhere between $525 million to $850 million, and yearly spends a similar amount each year along with forgiving a 1.5bn dollar debt[0]. So I think (VERY) conservatively we can say $2bn (I think even $4bn is likely conservative here)? While I see articles that discuss how the value could be north of $100bn[1] (which certainly surpasses a very liberal estimate of costs), I have yet to see evidence that this is actual value that has been returned to Google. I can only find information about 2022 and 2023 having profits in the ballpark of $60m.
This isn't to say that AlphaFold won't offset all the costs (I actually believe it will), but I your sentence does not suggest speculation but rather actualization ("has basically paid", "been"). I think that difference matters enough that we have a dozen cliches with similar sentiment. In the same way my annoyance is not that we are investing in ML[2], but how quick we are to make promises and celebrate success[3]. Actually, my concern is that while hype is necessary, overdoing it allows charlatans[4] to more easily enter the space. And if they are able to gain a significant foothold (I believe that is happening/has happened) then this is actually destructive to those who actually wish to push forward technology.
[0] https://www.quora.com/How-much-money-did-Google-spend-on-Dee...
[1] https://www.bloomberg.com/news/articles/2024-05-08/deepmind-...
[2] disclosure, I'm an ML researcher. I actually am in favor of more funding. Though different allocation.
[3] I'm willing to concede that success is realistically determined by how one measures success, and that this may be arbitrary and no objective measure actually exists or is possible.
[4] One needs not knowingly be a charlatan. Only that the claims made are false or inaccurate. There are many charlatans who believe in the snake oil they sell. Most of these are unwilling to acknowledge critiques. A clear example is religion. If you believe in a religion, this applies to all religious organizations except the one you are a part of. If you are not religious, well the same sentence holds true but the resultant set is one larger.
I just mean it demonstrated solving something challenging in a convincing way, justifying a great deal of additional resources being dedicated to applying ML to a wide range of biological research.
Not that it actually generates revenue or solves any really important health problems.
The big problem with the latter statement isn't so much exaggeration, but something a bit more subtle. It is that people start to believe you. But then they sit waiting, and in that waiting eventually get disappointed. When that happens the usual response often feeds into conspiracies (perpetuating the overall distrust in science) or generates an overall bad sentiment against the whole domain.
The problem is that companies are bootstrapping with hype. The problem is that this leads to bubbles and makes it a ripe space for conmen, who just accelerate the bubble. There's no problem with Google/Microsoft/OpenAI/Etc talking to researchers/developers in the language of researchers/developers, but there is a problem of them talking to the average person in the language of the future. It's what enables the space for snakeoil like Rabbit or Devin. Those steal money from normal people and takes money from investors that could be better spent on actually pushing the research of the tech forward so that we can eventually have those products.
I understand some bootstrapping may be necessary due to needing money to even develop things, but certainly the big companies are not lacking in funding and we can still achieve the same goals while being more honest. The excitement and hope isn't the problem, it is the lying. "Is/Can" vs "will/we hope to"
Either way, AlphaFold is one of the greatest achievements in science so far, and the funding agencies definitely are paying lots of attention to funding additional work in machine learning/biology, so in some sense, my statement is effectively true, even if not pedantically, literally correct.
If your "causal speech" is lying, then I don't think the problem is someone getting "triggered", I think it is because you lied.
> write long analytic responses
I'll concede that I'm verbose, but this isn't Twitter. I'd rather have real conversations.
That all depends on how you measure discoveries. The most common metric is... publications. Publications are what advance your career and are what you are evaluated on. The content may or may not matter (lol who reads your papers?) but the number certainly does. So the best way to advance your career is to write a minimum viable paper and submit as often as possible. I think we all forget how Goodhart's Law comes to bite everyone in the ass.
Plus, I mean, there are a lot of products that don't work. We all buy garbage and often can't buy not garbage. Though I guess you're technically correct that in either of these situations there can still be a return on investment, but maybe that shouldn't be good enough...
While this is true, it is also true that the sentiment of this line is frequently used outside of the biotech (note that this entire website is primarily dominated by computer science posts). I think it is just worth mentioning that it is perfectly valid for discussions to not be pigeonholded and that you are perfectly allowed to talk about similarities in other fields (exact or inexact).
It is also true that my reply is not invalidated by changing settings. If you pay close attention you'll notice the generality of the comment outside of the specific example. In fact, this is exactly what the OP did, considering the article is about a machine learning paper and the example they used to __illustrate__ their point was about their wife's work in biotech. But the sentiment/point/purpose of their comment would not have changed were their wife to work in physics/engineering/chemistry/underwater basket weaving/whatever. So if this is your issue, I think they are misdirected and I ask that you please take it up with the OP and ensure that they know no comments in this thread may be about anything but ML papers by DeepMind. Illustrative examples out of domain are not allowed.
It is also true that this domain example was a product this company was selling. So I think you're being too quick to dismiss as not only did it run "in a lab" but the actual product runs in the real world. It has real customers who use the software.
It's also true I never accused anyone of doing "just [running] busy work" and that such an interpretation is grossly inaccurate. My final sentence should make this abundantly clear and I would argue is far more important in a setting such as medicine where I've said conveyed that you can sell ineffective or subpar products while still generating a profit and suggested that this probably shouldn't be the metric we care about (it certainly isn't what the spirit of those metrics are about).
But if you want to (implicitly) accuse me of derailing the conversation, I do not think you have the grounds to do so. But I will accuse you of doing so. If you disagree with my comment, you are more than welcome to reply in such a way. If you think my comment does not apply to dug discovery or think it only applies to ML[0], then you are welcome to state as much too and it is encouraged to state why. But the only derailing of the conversation has been the pigeonholing you have applied. If you think this is wrong, I still welcome a response to that as I am happy to learn how to communicate better but you did also catch me on a day where I'm not happy to be unreasonably and willfully misinterpreted.
[0] I didn't tell you the application... a bit presumptuous are we?
It is rather unfortunate that this sort of paper is hard to reproduce.
That is a BIG downside, because it makes the result unreliable. They invested effort and money in getting an unreliable result. But perhaps other research will corroborate. Or it may give them an edge in their business, for a while.
They chose to publish. So they are interested in seeing it reproduced or improved upon.
I suppose it's possible Google's own infrastructure is partitioned from GCP infrastructure, so they have a bunch of idle GPUs even while their cloud division can sell every H100 and A100 they can get their hands on?
(I doubt it's the other way round, that the Deepmind researchers could come in one day and find all their GPUs are being used by some cloud customer).
Call me cynical, but this is not what I experienced to be the #1 reason of publishing AI papers.
The unfortunate part of this is that it can have odd effects like people renaming well known things to make the work appear more impressive, obscure concepts, and drive up their citations.[0] The incentives do not align to make your paper as clear and concise as possible to communicate your work.
Of course it would be a tiny fraction of the $10m figure here, but even 1% would be $100,000. Negligible to Google, but for Google even $10 million is couch cushion money.
$10M is about what Google would spend to get a publication in a top-tier journal. But google's internal pricing and costs don't look anything like what people cite for external costs; it's more like a state-supported economy with some extremely rich oligarch-run profit centers that feed all the various cottage industries.
Not necessarily, publishing also ensure that the stuff is no longer patentable.
That was using idle cycles on Intel CPUs, not GPUs or TPUs though.
I don’t think this is valid, as this point seems to ignore the fact that the data center that this compute took place in required a massive investment.
A paper like this is more akin to HEPP research. Nobody has the capability to reproduce the higgs results outside of at the facility the research was conducted within (CERN).
I don’t think reproduction was a concern of the researchers.
Obviously the 'best' result would be to have a separate collider as well, but no one is going to fund a new collider just to reaffirm the result for a third time.
The point I was trying to make was the fact that nobody (meaning govt bodies) was willing to make another collider capable of repeating the results. At least not yet ;).
I mention this because a lot of universities and small labs are being edged out of the research space but we still want their contributions. It is easy to always ask for more experiments but the problem is, as this blog shows, those experiments can sometimes cost millions of dollars. This also isn't to say that small labs and academics aren't able to publish, but rather that 1) we want them to be able to publish __without__ the support of large corporations to preserve the independence of research[0], 2) we don't want these smaller entities to have to go through a roulette wheel in an effort to get published.
Instead, when reviewing be cautious in what you ask for. You can __always__ ask for more experiments, datasets, "novelty", and so on. Instead ask if what's presented is sufficient to push forward the field in any way and when requesting the previous things be specific as to why what's in the paper doesn't answer what's needed and what experiment would answer it (a sentence or two would suffice).
If not, then we'll have the death of the GPU poor and that will be the death of a lot of innovation, because the truth is, not even big companies will allocate large compute for research that is lower level (do you think state space models (mamba) started with multimillion dollar compute? Transformers?). We gotta start somewhere and all papers can be torn to shreds/are easy to critique. But you can be highly critical of a paper and that paper can still push knowledge forward.
[0] Lots of papers these days are indistinguishable from ads. A lot of papers these days are products. I've even had works rejected because they are being evaluated as products not being evaluated on the merits of their research. Though this can be difficult to distinguish when evaluation is simply empirical.
[1] I once got desk rejected for "prior submission." 2 months later they overturned it, realizing it was in fact an arxiv paper, for only a month later for it to be desk rejected again for "not citing relevant materials" with no further explanation.
> But you can be highly critical of a paper and that paper can still push knowledge forward.
Can you give a concrete example of this?
[1] if the total cost estimate was relatively low, say less than 10k, then of course the lowest rental price and a random training codebase might make some sense in order to reduce administrative costs; once the cost is in the ballpark of millions of USD, it feels careless to avoid optimizing it further. There exist H100s in firesales or Ebay occasionally, which could reduce the cost even more, but the author already mentions 2USD/gpu/hour for bulk rental compute, which is better than the 3USD/gpu/hour estimate they used in the writeup.
MFU can certainly be improved beyond 40%, as I mention. But on the point of small models specifically: the paper uses FSDP for all models, and I believe a rigorous experiment should not vary sharding strategy due to numerical differences. FSDP2 on small models will be slow even with compilation.
The paper does not tie embeddings, as stated. The readout layer does lead to 6DV because it is a linear layer of D*V, which takes 2x for a forward and 4x for a backward. I would appreciate it if you could limit your comments to factual errors in the post.
And the blog post author is talking about the output layer where the model has to produce an output prediction for every possible token in the vocabulary. Each output token prediction is a dot-product between the transformer hidden state (D) and the token embedding (D) (whether shared with input or not) for all tokens in the vocabulary (V). That's where the VD comes from.
It would be great to clarify this in the blog post to make it more accessible but I understand that there is a tradeoff.
Even if it's a small model, one could use ddp or FSDP/2 without slowdowns on fast interconnect, which certainly adds to the cost. But if you want to reproduce all the work at the cheapest price point you only need to parallelize to the minimal level for fitting in memory (or rather, the one that maxes the MFU), so everything below 2B parameters runs on a single H100 or single node.
Looking at [1], the authors there claim that their improvements were needed to push BERT training beyond 30% MFU, and that the "default" training only reaches 10%. Certainly numbers don't translate exactly, it might well be that with a different stack, model, etc., it is easier to surpass, but 35% doesn't seem like a terribly off estimate to me. Especially so if you are training a whole suite of different models (with different parameters, sizes, etc.) so you can't realistically optimize all of them.
It might be that the real estimate is around 40% instead of the 35% used here (frankly it might be that it is 30% or less, for that matter), but I would doubt it's so high as to make the estimates in this blog post terribly off, and I would doubt even more that you can get that "also for small models with plain pytorch and trivial tuning".
From what I can tell, ads made the money and search/ads bought machines with their allocated budget, TI used their budget to run the systems, and then funny money in the form of quota was allocated to groups. THe money was "funny" in the sense that the full reach-through costs of operating a TPU for a year looks completely different from the production allocation quota that gets handed out. I think Google was long trying to create a market economy, but it was really much more like a state-funded exercise.
(I am not proud of how much CPU I wasted on protein folding/design and drug discovery, but I'm eternally thankful for Urs giving me the opportunity to try it out and also to compute the energy costs associated with the CPU use)
The cost was probably the limiting factor.
(A good rule of thumb is that an employee costs about twice their total compensation.)
From the link: "the total compute cost it would take to replicate the paper"
It's not Google's cost. Google's cost is of course entirely different. It's the cost for the author if he were to rent the resources to replicate the paper.
For Google, all of it is running at a "best effort" resource tier, grabbing available resources when not requested by higher priority jobs. It's effectively free resources (except electricity consumption). If any "more important" jobs with a higher priority comes in and asks for the resources, the paper-writers jobs will just be preempted.
TL; DR: It's not ~free.
You get that organically when you are serving lots of users. And, there's not much GPUs etc. used for that. Training LLMs gives you a different utilization pattern. The "best effort" resources aren't as useful in that setup.
For example, if YOU want to rent a backhoe to do some yard rearrangement it’s going to cost you.
But Bob who owns BackHoesInc has them sitting around all the time when they’re not being rented or used; he can rearrange his yard wholesale or almost free.
"Underutilized" isn't the right word here. There's some value in putting your capital to productive use. But, once immediate needs are satisfied, there's more value in having the capital available to address future needs quickly than there would be in making sure that everything necessary to address those future needs is tied up in low-value work. Option value is real value; being prepared for unforeseen but urgent circumstances is a real use.
I believe the larger factor, and someone correct me if they have a better understanding of this, is that for commercially rented properties the valuation used to determine the mortgage terms you get takes into account what you claim to be able to get from rent. Renting for less than that reduces the valuation and can put you upside down on the mortgage. But the bank will let you defer mortgage payments, effectively taking each month of mortgage duration and moving it from now to after the last month of the mortgage duration, extending the time they earn interest for.
So if no one want to lease the space at that price after a prior lessee leaves for whatever reason, it's better for the property owner financially to leave the space vacant, sometimes for years, until someone willing to pay that price comes along, than to lower the rent and get a tenant.
Part of the reason things like Halloween Superstores can pop in is the terms often exclude "short term leases" which are under six months.
Also when you're leasing to companies, they are VERY quick to jump at lower prices if available, which means that if you drop the lease for one tenant, the others are sure to follow, sometimes even before lease terms are up.
We saw the same thing with JIT manufacturing during Covid.
Cloud providers pay capital costs (CapEx) for servers, GPUs, data centers, employees, etc. Utilization allows them to recoup those costs faster.
Cloud customers pay operational expenses (OpEx) for usage.
So Google generally has excess capacity, and while they would prefer revenue-generating customer usage, they’ve already paid for everything but the electricity, so it’s extremely cheap for them to run their own jobs if the hardware would otherwise be sitting idle.
This would be 100% free, as all electricity and "wear and tear" would be required anyhow.
As you run close to 100% utilization, you also run close to infinity waiting times. You don't want that. It might be acceptable for your internal projects (the actual waiting time won't be infinity, and you'll cancel them if it gets too close to infinity) but it's certainly not acceptable for customers.
https://www.bigfishgames.com/us/en/games/5941/roads-of-rome/...
The structure of a time management game is:
1. There's a bunch of stuff to do on the map.
2. You have a small number of workers.
3. The way a task gets done is, you click on it, and the next time a worker is available, the worker will start on that task, which occupies the worker for some fixed amount of time until the task is complete.
4. Some tasks can't be queued until you meet a requirement such as completing a predecessor task or having enough resources to pay the costs of the task.
You will learn immediately that having a long queue means flailing helplessly while your workers ignore hair-on-fire urgent tasks in favor of completely unimportant ones that you clicked on while everything seemed relaxed. It's far more important that you have the ability to respond to a change in circumstances than to have all of your workers occupied at all times.
Ah, sounds like Dwarf Fortress!
Queues are useful to decouple the output of one process to the input of another process, when the processes are not synchronized velocity-wise. Like a shock absorber, they allow both processes to continue at their own paces, and the queue absorbs instantaneous spikes in producer load above the steady state rate of the consumer (side note: if queues are isolated code- and storage-wise from the consumer process, then you can use the queue to prevent disruption in the producer process when you need to take the consumer down for maintenance or whatever).
Running with very small queue lengths is generally fine and generally healthy.
If you have a process that consistently runs with substantial queue lengths, then you have a mismatch between the workloads of the processes they connect - you either need to reduce the load from the producer or increase the throughput of the consumer of the queue.
Very large queues tend to hide the workload mismatch problem, or worse. Often work put into queues is not stored locally on the producer, or is quickly overwritten. So a consumer end problem can result in potential irrevocable loss of everything in the queue, and the larger the queue, the bigger the loss. Another problem with large queues is that if your consumer process is only slightly faster than the producer process, then a large backlog of work in the queue can take a long time to work down, and it's even possible (admission of guilt) to configure systems using such queues such that they cannot recover from a lengthy outage, even if all the work items were stored in the queue.
If you have queues, you need to monitor your queue lengths and alarm when queue lengths start increasing significantly above baseline.
This was my first job after moving into this state. Between my labor and parts, it was about 15% of the sale price.
My most interesting repair was a 1943 Cadillac, a 'war car'.
If the job could easily run for weeks, even when you could buy your way for doing it in a day.
Then have a bidding on this “best effort” resource, where they factor in electricity at any given time
Those effort needs to be added in the cost calculation too.
Those effort needs to be added in the cost calculation too
The problem with neoclassical economics is that it doesn't concern itself with the physical counterpart of liquidity. It is assumed that the physical world is just as liquid as the monetary world.
The "liquidity mismatch" between money and physical capital must be bridged through overprovisioning on the physical side. If you want the option to choose among n different products, but only choose m products, then the n - m unsold products must be priced into the m bought products. If you can repurpose the unsold products, then you make a profit or you can lower costs for the buyer of the m products.
I would even go as far as to say that the production of liquidity is probably the driving force of the economy, because it means we don't have to do complicated central planning and instead use simple regression models.
Isn't that all what high frequency traders would say? :)
Perhaps there is some limit at which additional liquidity doesn't offer much value?
There isn't much there about stocks markets.
I was told by an employee that GDM internally has a credits system for TPU allocation, with which researchers have to budget out their compute usage. I may have completely misunderstood what they were describing, though.
it is a hustle only for the near future while this bubble lasts, but can help reduce costs.
Google makes ~$300b a year in profit. They could make a $10m mistake every day and barely make a dent in it.
To put that into perspective, Alphabet's revenue has increased 13.38% year-over-year as of June 30, arriving at $328.284 billion dollars - i.e. it has increased by $38.74 billion in that time. A $10 million dollar mistake translates to losing 0.0258% of that number.
A $10 million dollar mistake costs Alphabet 0.0258% of the amount their revenue increased year-over-year as of last month. Alphabet could have afforded to make 40 such $10 million dollar mistakes in that period and it would have only represented a loss of 1% of the year-over-year increase in revenue. Taking the year-over-year increase down by 1% (from 13.38% to 12.38%) would have required making 290 such $10 million dollar mistakes within one year.
Let me repeat that because it bears emphasizing: over the past years, every year Google could have easily afforded an additional 200 such $10 million dollar mistakes without significantly impacting their increase in revenue - and even in 2022 when inflation was almost double what it was in the other year they would have still come out ahead of inflation.
So in terms of numbers this is demonstrably false. Of course the existence of repeated $10 million dollar mistakes may suggest the existence of structural issues that will result in $1, $10 or $100 billion dollar problems eventually and sink the company. But that's conjecture at this point.
[0]: https://www.macrotrends.net/stocks/charts/GOOG/alphabet/reve...
I’m confident each one of them were multiple of $10M investments.
And this is just what we know because they were launched publicly.
The equivalent wastage for a self-employed person would be allowing a few cups of Starbucks coffee per year to go cold.
That means Google payed way less than this amount and if you wanted to reproduce the paper yourself, you would potentially pay a lot more, depending on how many engineers you have in your team to squeeze every bit of performance per hour out of your cluster.
How is anyone else going to reproduce the experiment if it's going to cost them $10 million because they don't work at Google and would have to rent the infrastructure?
It works and helps to get a salary raise or a better job, so they continue.
A bit like when someone goes to a job interview, didn't do anything, and claims "My work is under NDA".
That being said, yes, this is hard to reproduce for your average Joe, but there are also a lot of companies (like OpenAI, Facebook, ...) that are able to throw this amount of hardware at the problem. And in a few years you'll probably be able to do it on commodity hardware.
No, it's not. The author clearly states in the very first paragraph that this is the price it would take them to reproduce the results.
Nowhere in the article (or the title) have they implied that this is how much Google spent.