Small Language Models Are Also Few-Shot Learners
aclanthology.org
aclanthology.org
In my mind it would involve some kind of generative grammar that would generate nodes (select from a pool), then these nodes could be trained. I'm thinking about something like a grammar for Bayesian networks or, more broadly, a generative program induction scheme where you'd specify a programming language grammar, generate programs that fit the data and tune the parameters.
I attempted and failed to implement some of these ideas (described here: https://www.quora.com/What-deep-learning-ideas-have-you-trie...). The search space for generative programs is just so huge and irregular and I don't know how to keep things differentiable. In the paper, they mention gradient-based optimization so maybe to figured out part of it.
Why can't they just talk about the reduction of energy or compute resources instead?
In practice algorithms don't work in isolation. You may have to do additional pre-processing on your data to get it into this kind of model, offsetting any energy usage that you reduce through the execution of your model. There are any number of small things that aren't taken into account when looking at this portion of an overall solution and you can't equally make statements about carbon usage based on energy usage alone, nor can you make statements about energy usage based on computation requirements alone.
The actual energy usage is not an independently derived factor and unrelated to the work itself. Adding it as a benefit in your paper without the additional work of proving that is marketing fluff. It doesn't belong and actively hurts studies that are focused on energy usage reduction with the explicit intent of reducing carbon footprints.
https://arxiv.org/abs/1906.02243 (cited in the above paper)
> The U.S. Environmental Protection Agency (EPA) provides average CO2 produced (in pounds per kilowatt-hour) for power consumed in the U.S. (EPA, 2018), which we use to convert power to estimated CO2 emissions:
> CO2e = 0.954 pt
> This conversion takes into account the relative proportions of different energy sources (primarily natural gas, coal, nuclear and renewable) consumed to produce energy in the United States.
(Other countries are also included in the paper.)
https://arxiv.org/abs/2104.10350
The authors note that in many cases accurately estimating the carbon footprint is difficult because the information required is not publicly or readily available. However, they do provide some additional data, improved calculations, as well as motivations beyond CO2 reduction.
Because of the opaqueness of the electrical infrastructure, it's not precise of these carbon fiber footprint measurements, because the data is simply not there.
Therefore, a better measurement is just the electric usage...
Yes, they are approximations in lieu of more accurate data but that doesn't invalidate them as a tool or a motivation for future work such as this one.
Consider further that it's not just OpenAI's models in question: it's every practitioner who attempts to train similarly large models. These practitioners may not be using "green" data centers, even if we generously assume that OpenAI does. (Microsoft's 100% renewable target for data centers isn't until 2025. Read another way: they may be trying but they're not there yet.)
The available data and approximations illustrate that it is not accurate to assume that the average data center are powered by 100% renewables and carbon neutral. Thus the only reasonable conclusion is that more efficient models will have a positive impact on CO2 which is the motivation of the paper.
Even if you don't agree, it's not completely unfounded and is based on at least some research and data. At the end of the day, is this really worth fighting against? Who wants less energy efficient models?
https://www.comet.ml/site/introducing-codecarbon-an-open-sou...
If the paper from the post indeed tried to evaluate and compare various kinds of optimization, then citing the CO2 emissions would be valid. But since it only discusses the improvements in the model itself, then it would be much more productive to just point to the reductions in the required memory, GPU/TPU-time etc.
The readers can do the carbon math themselves depending on how carbon-neutral is their datacenter.
For example, Google CoLab is running in Google data centers, which are carbon neutral. Which is more or less true, but also depends on the exact nature of the source of their power and accuracy of the carbon offsets they purchase, and we all know Google is perfectly accurate and truthful all the time. /s lol
I trust Google insofar as their publicized usage of renewables and extensive use of solar. I also suspect that they're on-grid and a net producer of energy, taking advantage of deals with power companies and governments to ekeout every last penny of value in their infrastructure.
The problem is that without exact reporting of numbers, the margin of error for that source alone creates huge uncertainty in trying to assess the net carbon footprint of their service. How much research is being done using Google infrastructure? How much is being done on college campus data centers that run their hpsc on solar and wind? How much money is spent on offsets by those other sources? Again, the reality requires exact knowledge, since the usage of offsets introduces huge uncertainty, so aggregate reported usage could vary in accuracy as a proxy by more than 100% of the naively assumed footprint.
The studies are only as good as their data, and the data isn't very good unless it's obtained through legal mandate, via subpoena or regulated reporting. To my knowledge, very little of the data available for these estimates is anything except self reported numbers. The math and analyses they do are great, but the margin of error likely exceeds 75%.
Sure, a reduction in computational expenses is in many ways desirable, but I don't think the carbon footprint of a model is a very good metric. There are much better arguments for more efficient models.
I guess you have to do, what you have to do for those grants.
What many researchers can't do is replicating the training process.
Though "learners" should have given it away as well