Google Vizier: A Service for Black-Box Optimization
ai.google
ai.google
"5.3 Delicious Chocolate Chip Cookies
Vizier is also used to solve complex black–box optimization problems arising from physical design or logistical problems. Here we present an example that highlights some additional capabilities of the system: finding the most delicious chocolate chip cookie recipe from a parameterized space of recipes... We provided recipes to contractors responsible for providing desserts for Google employees. The head chefs among the contractors were given discretion to alter parameters if (and only if) they strongly believed it to be necessary, but would carefully note what alterations were made. The cookies were baked, and distributed to the cafes for taste–testing. Cafe goers tasted the cookies and provided feedback via a survey. Survey results were aggregated and the results were sent back to Vizier. The “machine learning cookies” were provided about twice a week over several weeks...
The cookies improved significantly over time; later rounds were extremely well-rated and, in the authors’ opinions, delicious."
This is cute, but "in the authors' opinions, delicious" does not contribute anything to scientific research. Yes, you can use black-box optimization for cookie recipes. No, you should not make any sort of performance claims nor talk about "significant" improvement unless you back it up.
But it's worse than the last 100 SOTA models...
Can there be/Is there an objective way to measure deliciousness? If not what can they really say?
This, from the paper, is a great start on measuring deliciousness. But if you go to that effort, why not provide a benchmark? Such as the existing chocolate chip cookie recipe or the average scores from 10 random chocolate chip cookie recipes?
As a reader, I might suspect that performance was not that great compared to baselines but that the authors were able to mask it by making this qualitative claim ("the cookies were delicious") -- which I'm sure was still true!
Like said elsewhere in this thread, if after a few batches the algorithm optimizes into a local maximum gully that some people don't like, but others really love, then the ones that dislike it will stop taking the cookies and thus stop rating them.
But does that mean they're less delicious? Nobody knows! Because certainly the first and foremost great start on measuring deliciousness, is doing research on what is deliciousness, and are there any good methods to measure deliciousness that can be compared with other literature and research, and if that fails, either measure something else or defend why you chose a target metric that can't be measured properly. The latter being valuable if your way of measuring is novel and you think it's better than what has been used so far in the field. But you gotta provide a good argument for that, and the method used in this experiment wasn't particularly novel, simply flawed.
I mean you can be all giggly-googly about it, but food science is a thing.
This seems more like a spanning/optimization problem (how to cover the majority of preferences while not needing too many varieties) than an actual optimization problem, but then again it's just a sample use case to use to sell their API, so they probably didn't think too hard about it.
I am very disappointed. I, for one, welcome our new robot overlords and also want to try their chocolate chip cookies.
And the recipe I found internally is "The Real California Cookie" in that paper:
Bake at 163C (325◦ F) until brown:
• 167 grams of all-purpose flour.
• 245 grams of milk chocolate chips.
• 0.60 tsp. baking soda.
• 0.50 tsp. salt.
• 0.125 tsp. cayenne pepper.
• 127 grams of sugar (31% medium brown, 69% white).
• 25.7 grams of egg.
• 81.3 grams of butter.
• 0.12 tsp. orange extract.
• 0.75 tsp. vanilla extract.
• 5¼ oz. of all-purpose flour.
• 7 oz. of milk chocolate chips.
• scant ½ + ⅛ tsp. baking soda.
• ½ tsp. salt.
• ⅛ tsp. cayenne pepper.
• 4½ oz. of sugar (31% medium brown, 69% white).
• 1 oz. of egg.
• 2¾ oz. of butter.
• ⅛ tsp. orange extract.
• ¾ tsp. vanilla extract.
Or, equivalently:
• 167 grams of all-purpose flour.
• 245 grams of milk chocolate chips.
• 2.96 mL baking soda.
• 2.46 mL salt.
• 0.61 mL cayenne pepper.
• 127 grams of sugar (31% medium brown, 69% white).
• 25.7 grams of egg.
• 81.3 grams of butter.
• 0.59 mL orange extract.
• 3.7 mL vanilla extract.
Unfortunately my microliter precision pipette is in the shop otherwise I would be enjoying some of these right now!
(say, as spectra)
After you got handed a cookie you were asked to fill out a feedback form on a tablet, which almost always is annoying.
But more importantly, the process of optimizing the ingredients left out many parts of what makes food enjoyable, temperature and texture for example.
I know this was meant to be a fun showcase of ML but to me this is still my favourite example of explaining misuse of ML technology where a simpler statistical model supported by expert opinion would have outperformed the taste of the cookies. Eg, any cookie expert might know that the number of people who like spicy cookies is only a small subset of all cookie lovers.
A good example of training/serving skew.
Black-box optimization is a hugely important problem to solve, especially when experiments require real wet-work (i.e. medicine, chemistry, etc.). Kudos to Google for commercializing this - I expect it will see a lot of use in those fields. But it's bittersweet to know that it's taken this long for this type of application to be promoted like this.
And they're competitive with Spearmint (though not necessarily the closed-source versions of it used at Whetlab), though Vizier remains to be seen: https://arxiv.org/pdf/1603.09441.pdf
You hit the nail on the head. We've been trying to promote more sophisticated optimization approaches since the company formed 4 years ago and are happy to see firms like Google, Amazon, IBM, and SAS enter the space. We definitely feel like the tide of education lifts all boats. Literally everyone doing advanced modeling (ML, AI, simulation, etc) has this problem and we're happy to be the enterprise solution to firms around the world like you mentioned. We provide differentiated support and products from some of these methods via our hosted ensemble of techniques behind a standardized, robust API.
We're active contributors to the field as well via our peer reviewed research [1], sponsorships of academic conferences like NIPS, ICML, AISTATS, and free academic programs [2]. We're super happy to see more people interested in the field and are excited to see where it goes!
> Cloud Machine Learning Engine is a managed service that enables you to easily build machine learning models that work on any type of data, of any size. And one of its most powerful capabilities is HyperTune, which is hyperparameter tuning as a service using Google Vizier.
https://cloud.google.com/blog/products/gcp/hyperparameter-tu...
We've been solving this problem for customers around the world for the last 4 years and have extended a lot of the original research I started at Yelp with MOE [1]. We employ an ensemble of optimization techniques and provide seamless integration with any pipeline. I'm happy to answer any questions about the product or technology.
We're completely free for academics [2] and publish research like the above at ICML, NIPS, AISTATS regularly [3].
We haven't been actively maintaining the open source version of MOE because we built it into a SaaS service 4 years ago as SigOpt (YC W15). Since then we've also had some of the authors of BayesOpt and HyperOpt work with us.
Let me know if you'd like to give it a shot. It is also completely free for academics [1]. If you'd like to see some of the extensions we've made to our original approaches as we've built out our ensemble of optimization algorithms check out our research page [2].
Please note that their work is primarily for expensive optimization where you cannot afford millions of function evaluations like you would need on rotated Rastrigin to solve it exactly for n>64.
It really depends on parameterization though. Did they use off-the-shelf parameters? Did they make any attempt to tune CMAES for their test functions? Did they make any attempt to tune their own optimizer for the test functions? It's really easy to put your thumb on the scale when you're writing a paper about a black-box optimizer, and my default reaction is to assume that any paper claiming to equal or outdo CMAES is pure B.S. This paper did nothing to convince me otherwise.
Stopping rules are also a requirement of tuning. Can't just have a while True:. If an objective function doesn't improve after 100 (or any #), there needs to be logic to stop the trials since this system is apparently serving all of Alphabet.
>Only up to 64 variables? Why such small problems?
This is a tool tuned for optimizing the hyperparameters of machine learning models. In practice, most models have, maybe, 10 hyperparameters in the normal sense (learning rate, nonlinearity, etc.). If you start talking about model structure as a hyperparameter, you can get vastly more than 64, but then these techniques aren't great anyway.
So why such small problems? Because the blackbox function you're optimizing takes hours to run. So as another user mentions, if it takes you 700 function evaluations, you're gonna be running for the better part of a year.
So the domain over which vizier works is one where you probably are never really using more than ~100 evaluations and often even then you would prefer to stop early if you meet certain conditions (because wasting significant compute doing unnecessary optimization is costly).
As a concrete example, a single iteration of a relatively small and non-SOTA algorithm takes 20 minutes and $40 to run. [0]
So stopping a day early saves you a day and $3000. Now scale that up 5 or 10x for larger models. (another way of putting this is that a day and $10K might be worth a .5% increase in accuracy, but isn't worth a .05% increase in accuracy. Stopping rules can encode that).
"The page you are trying to view cannot be shown because the authenticity of the received data could not be verified."