Why train when you can optimize?
justinmeiners.github.io
justinmeiners.github.io
This is possible! One of the cool things about neural networks is that you can try to encode prior understanding into either the structure of the network or choice of activation function. See the paper Neural Networks Fail to Learn Periodic Functions and How to Fix It by Ziwin, Hartweg, and Uweda. Where they propose the activation function f(x) = x + sin(x)^2 that can encode an understanding that the underlying function should be periodic.
[1] https://proceedings.neurips.cc/paper/2020/file/1160453108d3e...
Or you need to optimize without using gradients.
[1] https://geometricdeeplearning.com
[2] https://arxiv.org/pdf/2104.13478.pdf page 27 (23 if you count book pages).
Significant parts of x+sin are close to linear. You don't need to use this activation as the result layer either. Why would we lose anything?
> If we had to build separate programs for each product requirements variations
We pretty much do? We use both extremely generic frameworks both in technical sense (.net) and organisational (sap). But we also have software written to specific requirements where needed (there's millions of very specific ways to invoice someone, companies get invoicing platforms written just for them from scratch). There's space for both approaches.
This is a quote from the parent comment where you state that NNs with periodic activations can't represent non periodic functions.
Siren (what I linked) uses periodic activations and is able to represent non periodic functions.
The problem is improperly evaluating the neural network as a function of time, instead of evaluating the network as a function of previous state.
When we humans approximate functions (let’s say you’re drawing it on a piece of paper, or waving your arm around) we do not simply look at a clock and feed forward that information directly into our motor neurons. Rather, we have sensory neurons that feed in the current state of the function we are approximating as an input, then approximating the next output of a periodic function becomes trivial.
It’s very easy to train a neural network to look at a piece of a sine wave and predict what the next value should be - the fact that the function is periodic actually helps you.
> So it should not be approximated in that way.
We can’t conclude a dynamic approximation is a bad approach based purely on the fact the underlying function isn’t dynamic.
The function might nevertheless be easily approximated via dynamics — as in the case of predicting sine from seeing the recent history.
I’m not sure what you mean. Are you talking about training a network to predict cos(x)? In this case nothing at all changes by training on the derivative.
Or do you mean train the net to take as input sin(t) and produce sin(t+0.01)? The problem with this is that any given value of sin(t) has 2 answers for sin(t+0.01), therefore your optimizer is going to spit out 0 as the answer. Plus, this is a different problem entirely, and you lose the ability to infer sin(t) based on t. It doesn’t answer the same question.
Your suggestion is further going to be seriously confounded if the periodic function is more complex, say, the sum of several sin waves. There can be an arbitrary number of values y that correctly match f(x).
You should try it before making assumptions that it’s easy, see what it takes to train a net to predict sin(x), without embedding knowledge of the fact that sin is periodic.
> The problem is improperly evaluating the neural network as a function of time, instead of evaluating the network as a function of state.
This also sounds like an assumption to me that somehow the entire world of research has failed to consider the most obvious of ideas. The point of both optimizers and neural networks is that they can be black boxes, right? It doesn’t matter at all whether the input is time based or position based or a function of money. The network, in theory, can learn any function, time or otherwise, and there’s nothing special about time.
But, neural networks function better with domain knowledge. When you know the function domain is time and that the output is periodic, you can do things to make a network easier to train, like using a periodic activation function.
As a side note, RNNs explicitly model an NN based on previous state. Also all layered NNs can be viewed as a series of smaller nets that feed state to the next net. In some sense, NNs always evaluate as a function of state.
This is also the classic demo case for LSTMs, I have a notebook open right now that as a LSTM learning the sine function quite well with 32 dimensional state vector.
However the authors point still stands. Neural Networks are not great at computing arbitrary functions where the output is an unbounded real number. The sine function is still limited to range that is bounded to [0,1].
A better example of where NNs really fail is when trying to learn tricky to implement the normal quantile function (inverse CDF).
This would be an excellent place for function approximation because
a.) generating training data is easy (just run random numbers into the CDF and reverse these arguments into a NN)
b.) manually writing the quantile function from scratch is a pain since it involves the inverse error function which is very annoying to implement from scratch.
You can learn the standard quantile function for mean=0, sd=1, however if you try to generalize this to taking not only the desired quantile but an arbitrary mean and standard deviation you will not learn anything useful.
It's a bit of a shame that neural networks are weak in this area because it would be incredible to have a good tool to approximate inverse functions in general. The fact that we almost never see neural networks being used as a tool for this type of work is evidence of this limitation.
In general if you're problem can't be modeled where the output is some vector of probabilities it's not a great fit for NNs.
It's funny you say. I haven't actually used NNs much in my research but my background is in math and in my spare time I'm a maintainer for SciPy working on special and statistical functions, often the exact kind of stuff you mentioned.
I might take up your challenge and write a blog post about it or something if I get any success.
Anyway, I wasn’t trying to invalidate the author’s general point, just point out a fun fact about his example.
Success or failure, I would really enjoy seeing that write up! I would be even more excited to be proven wrong.
The promise of "universal function approximator" is very temping. My personal dream would be to have it so one could essentially run scipy in reverse and learn the entire library with a NN. Even in the case I gave, the idea that you could learn an arbitrary quantile function means you could also arbitrarily learn a sampler for any distribution, since all you have to do compose a uniform sampler with whatever quantile function you learned.
Of course for this example I'm using "solved" using a similar approach with variational inference (pyro has a write up on it: https://pyro.ai/examples/svi_part_i.html, you might find David Blei's "Variational Inference: A Review for Statisticians" useful as well https://arxiv.org/abs/1601.00670)
I'd say that, from "LSTM" and "32 dimensional state vector" alone, the author's point still stands. That may not be a huge network by contemporary deep learning standards, but it's still a pretty darned expensive way to compute sine. That's 100% in line with the original assertion that ANNs can learn any function, but they can't necessarily do so efficiently.
Any suggestions on the best resources for branching out from RL to more classical control and optimization?
Just start learning the basics of supervised learning for classification and regression on common benchmark tasks like CIFAR and UCI. Apply a mix of linear models, neural networks, and trees like random forests and GBDTs. Next try convolutional networks for vision and transformers for NLP. You'll be all set to solve most real world problems.
That's an excellent quote by the way, do you perhaps have any source? Sounds like a good opener slide :)
I think that you can use optimization without having to learn anything about its algorithms (learning the basics is always advised of course, but that takes nowhere near as much effort). Nowadays, an off-the-shelf black-box solver will perform better than a custom implementation in most cases: what you lose in fine-tuning is more than compensated by the implementation quality and the sheer number of algorithms available.
> best resources for branching out from RL to more classical control and optimization
It depends on your problem. For general unstructured (nonconvex continuous) optimization, like the OP, you can use NLopt's documentation as a pragmatic starting point [1].
Optimal control is a whole different thing though, and you may have to start with some academic papers/books. I have been recommended this one [2], but I have not read it.
[1] https://nlopt.readthedocs.io/en/latest/
[2] https://link.springer.com/book/10.1007/978-1-4471-0967-9
In problems like this there are two aspects, (i) designing or specifying the search space of functions (ii) choosing the best function within the search space.
The opposite extremes are a) the search space contains only one function, the right function. In this case the training/optimization is moot. The other extreme is to have a very wide search space, say all smooth functions. In that case searching/training/optimizing is more challenging. The more reasonable example is one uses domain knowledge to design a much more restricted search space (for example, one may encode that the function is periodic with a known period) making the next step easier.
If wonder what classes of problems fit this? For sure, I can't imagine how can you tackle sentiment analysis or text classifiers using optimization.
VADER (Valence Aware Dictionary and sEntiment Reasoner) is a large lookup table mapping words to a sentiment score, calculated by surveying people. It is simple to use and involves no ML.
Decent sentiment analysis is a yet unsolved problem and current state of art solutions (mostly based on large pretrained langugage models like BERT, leveraging transfer learning from very large unlabeled corpora) are not sufficiently accurate, and of course neither is VADER. As soon as you go beyond simple binary polarity (blatantly positive/negative sentiment) and something more informative like various five-way sentiment tasks, or aspect-based sentiment or determining sentiment targets, sentiment analysis still has a long, long way to go.
Current ML solutions are a possible way to tackle sentiment analysis because while the current models are not good enough, we have evidence showing that with increased model and data size they improve and perhaps might result in sufficiently good sentiment analysis someday. Perhaps not, and we'll need something else; but at least they have some potential.
VADER and similar methods are not a way to tackle sentiment analysis because not only the current VADER model is not nearly good enough, it's clear that they can't scale to anything much better than that. Adding extra surveys to increase the lookup table quickly hits diminishing returns (unlike ML systems where we consistently see that ever larger models continue to get improvements), and if they can't even reach the current bar of ML models (which still are not good enough) then they can't bring us to the level of sentiment analysis where we want to be, they are a dead end.
Tackling sentiment analysis also requires handling a wide variety of languages, not only English, which is yet another aspect where data-driven models have an advantage over systems that require extensive human labor for each new language.
I mean, we all (me included!) want to believe that human knowledge encoded in rules can work, intuitively it's a very appealing concept, however, it does not work out in practice and IMHO in the end we all have to learn to accept the Bitter Lesson (http://incompleteideas.net/IncIdeas/BitterLesson.html).
There are other functions, for example, searching, sorting etc where its easier to specify accurately what the desired function does compared to giving a list of pairs of examples. In such cases ML may not be the best choice. Note, reasonably accurate sorting functions can be learned from examples, but that's not the most efficient way to design a sorting function.