Is deep learning a new kind of programming?
tomasp.net
tomasp.net
So he replaced the model with a single line of code with one boolean expression made with those 4 parameters connected with logical operators.
Neural networks make sense with huge number of input parameters where feature selection is really tricky to reason about and decision boundaries are very non-linear such as image classification.
Edited: slight clarification
It is funny how my uni top research (on 1m$ computers) neural nets are considered to make no sense anymore. That went a lot faster than programming.
I would be genuinely interested in examples of problems with a very low number of predictors (say two to five) when a neutral net would be appropriate (where as you say less complex methods have been tried and failed).
I just can't think of one.
I can't think of a method that would use fewer parameters. If nothing else, it's a decent way to compress the data set for interpolation (on nearby averages) as a use case, no?
To your point, I believe you meant "rough interpolation", and it's true in many cases NN's might produce a less overfitted approximating function if one has no prior knowledge of the generating function.
But if one can exploit prior knowledge, one can select an optimal set of basis functions and fit a more parsimonious model than a NN. For instance, if you knew that a nonlinear function was a function of sin, cos and logs, selecting these as basis functions and finding the correct functional form [2] would likely help an optimizer find more parsimonious model than a NN using standard activation functions (ReLU, sigmoid, etc). As a thought experiment, suppose the generating function was this: (5 parameters)
y = a1*log(a2*x)/cos(a3*x) + a4*sin(a5*x)
If one attempted to fit this with log, cos and sin basis functions, one is likely recover this form with ~5 parameters. But suppose we tried to fit this with an NN with the stipulation that the approximation error is under some ε -- I suspect we'll need quite a bit more than 5 parameters.NN's tend to generalize better (assuming proper regularization) than polynomial approximations and have fewer numerical problems like Runge's phenomenon, but I don't think NNs aim for (or have results that demonstrate) parsimony in parameters.
[1] https://en.wikipedia.org/wiki/Polynomial_interpolation
[2] If the functional form is unknown, there are techniques like "symbolic regression" that attempt to do a structure search to find a well-fitting structure. https://en.wikipedia.org/wiki/Symbolic_regression
You cant pick completely the wrong tool, and then complain about how unsuitable it was.
there's really no other magic in it than that.
Quantitatively perhaps, in that deep learning is equivalent to many neurons, whereas a least squares model is only equivalent to a single neuron. Other than that there is not much evidence that a human brain is qualitatively different.
So, could you explain again how the concepts and foundations of deep learning differs from plain old regression techniques?
Transitioning from a function that returns a scalar, to returning a polynomial, to returning an arbitrary function, and now to returning something stateful is a big difference. I suppose the next step is to write an algorithm that trains a Turing machine.
[0]: https://pytorch.org/docs/stable/generated/torch.nn.RNN.html
[1]: https://docs.nvidia.com/deeplearning/cudnn/api/index.html#cu...
The particular encoding -- matrix or otherwise -- is just an implementation detail.
* hidden Markov models
* autoregressive models
* learned LQR
Sure, some of those have finite memory, but RNNs are practically limited–not as easily quantified.
I see what you're saying about the single-step operation appearing unique, but I also think that RNNs can be viewed through a lens such that they look like "normal" programming concepts like generators and folding iterators.
I normally do not feel the need to comment about deep learning, because it is only tangentially related to some of my past projects, but I can also understand those who might want to comment negatively, because I have seen many cases of managers who did not understand at all how exactly certain problems can be solved, but nevertheless they pushed vigorously for the use of deep learning to replace other better suited solutions, because they believed it to be a modern and universally better method.
So there were times when I was tired of seeing one more attempt to misuse deep learning and to have to explain and demonstrate once more which solution is better.
Obviously, any such opinions, about which method is better in a certain case, should be proved with numerical results from tests or simulations, not with guesses, but some times that requires a lot of work to implement both methods, even if you are pretty sure about which will be the result.
And the mistake people make is that they don't start with a simple model. They move straight into the heavy ones.
For instance, FP vs Imperative, where a simple example (map for instance) is being shown and someone always compares it to a for loop, and then disses garbage collection. They're obviously thinking of their existing language (we know from the GC comment) and not of whatever point the OP is trying to make about composability or whatever.
To some degree, examples are to blame. They often aren't good and don't show, to the actual audience, what the author intended. In my FP example, you need something that can't get stuck in a syntax-level debate and is complex enough for the benefits to show through. Like doing something that would take a 50-level deep stack of for loops in an unnested fashion.
It gives you more respect for great teachers who can pull examples out of the air that both illustrate and refuse to misdirect.
Can you (because I certainly can't) write a good post or blog about the minimum useful deep learning project that can't be done equivalently via any other methods?
A hammer, nail, saw and timber can also be used to solve problems but I wouldn't call those programming in and of themselves, but they could be used to built analog computers (where cogs, cams etc. Are like lines of code or procedures).
Building a neutral network to get a result is not at all like programming. There is usually not a "perfect" structure, rather there are hardware, energy and time constraints to training and inference, balanced by over and under training the network.
The same can be said about complex simulations. The difference lies in knowing and being able to determine the limits of the system and the verification process.
In a simulation we can derive the accuracy of our model from parameters like numerical precision and -stability, coarseness and model used. In other words we know the function we want to model because we state it explicitly.
Neural nets can model any function and the challenge is to extract the learned function from the trained network, to examine its limits and correctness. This is what I understood the author meant by the "operational" viewpoint.
We can verify that a given network architecture combined with a given optimisation function will find a local minimum w.r.t a given set of training data. This can be verified and tested.
What's not so easy to verify and test, however, are the properties of the modelled function as well as the function itself. That's why we still have to rely on proxies like error metrics on fixed datasets or failure cases.
With a simulation on the other hand, we can easily control and predict the (quality of the-) outcome by manipulating well understood parameters (number of iterations, coarseness of the simulation, numerical precision, etc.).
I picked simulations as an example, because many other classes of program can be verified using formal methods since the desired results are usually known beforehand. Again, just another reason why the author talks about a distinction in terms of operations, not the fundamental type of programming.
I find this to be a very interesting and thought provoking idea.
I’d argue not. Mostly. If you fit a model using lm() in R, and then apply that model it’s not the same as hand selecting the weights of a linear equation and coding that equation.
You could in theory select the same weights, and code it by hand. But no one ever does.
That seems like a weird distinction to make.
Modern language models (eg GPT-3 et al) offer the capability to take a natural language input, match it against the context of the sentence, then propose a query that is understandable to the layperson. This abstraction allows us to understand the problem better, rather than just analyzing the way the problem manifests itself in code. Having a programming language that mirrors our everyday communication is an important step forward in making the innovations from software broadly available.
The next wave of programmers will need to understand how human language can be used to efficiently guide models to solve problems that can’t be solved by human-written code. This is a big challenge, and we’re just at the very beginning of it, but I think it will open up new and undiscovered ways to create value in the world.
I have quite a few additional thoughts on the topic which I’ve captured here: https://sundayscaries.substack.com/p/whos-the-real-expert
I went to a talk years ago on someone's PhD project involving a certain interactive debugger for Haskell, where the user could traverse the graph, making claims about nodes and eliminating possibilities. I wish I could remember its name.
uu-parsinglib [1] is a parser combinator library that provides error correction.
[1] https://hackage.haskell.org/package/uu-parsinglib-2.3.0
Maybe these algorithms could combine with AI to create something better.
I don't think that's the case at all. Mathematics developed a formalised non-natural language precisely because human language is completely unsuitable for expressing abstract concepts in a concise and unambiguous fashion.
You will find that even in non-technical fields language will quickly converge to a well-defined, coarse and highly coded subset of regular human language when efficiency and correctness are key. You can observe this in the different branches of military, medicine, and trades.
We use programming to formalise algorithms, processes, and models. Those are abstract concepts and the difficulty doesn't lie in expressing them verbally. This has been shown time and again by fruitless efforts to create localised dialects of more accessible programming languages like BASIC or Pascal.
Turns out it doesn't matter whether keywords are written in your native language or if you could write natural language-like sentences: the difficult part remained formalising the abstract concept and ideas in a meaningful, logical and sound way.
What I do think will help tremendously, however, is using system such as GPT-3 to create another level of abstraction. There are many descriptive tasks that don't need to be put into code manually. The structure and behaviour of UIs comes to mind.
Deep learning isn't programming per-se, but it definitely creates results that are of the same kind that programming would be able to create as well (in principle, at least, in many cases).
Machine learning is giving the input and the output to get a program.
Problem is, it's too difficult to summarize or understand the resulting program, while the program you get is tied to the output data which is never really accurate.
I'm still curious how ML specialists are approaching the task of analyzing a resulting deep neural network, and squeeze some science from it (meaning putting words on things they understand and are able to explain).
I've also read that google was using ML to test different learning models, to easily find the best model to use for a given problem. I'm not sure but it sounded like they were feeding the training model and the data into another learning model. I can't remember the details or the article or the reddit comment but it sounded quite interesting.
For example, writing an algorithm that has precise steps and procedures is programming. Putting my input into a box, shaking the box, and taking the result out is not programming, even if the box somehow solved the problem. Merely describing a problem and then having it solved is not enough to delineate programming, because that actually does apply to almost anything.
For instance, imagine you have a black box that observes the horse races, Twitbook, the betting market and so on, and based on those observations executes bets for you with a bookmaker. The execution of the orders has a measurable effect on your net worth.
You might write a traditional programme which takes all of this data, and based on some ETL, statistical models and probability calculations, executes orders.
You might do some ETL, plug it all into a neural network, tune it and execute orders based on the results.
Your traditional programme is very complex, and combinations of small bugs may have large effects on the results. Your unit and integration tests may themselves be wrong. Formal testing possibly reduces the expected value of the system and is an arse to carry out for any large system. The expected value of the system itself becomes harder to reason about as the system grows, based on the operation of reading and understanding the code.
The internals of your neural network are also difficult to reason about in some ways. It is difficult to understand the workings of your neural network and specific parts' effects on the measured effects of its output. It will take time to tune it and build the most profitable model.
Both implementations of the black box may be backtested, and some sort of trust can be established over the expected value of each implementation. Both implementations allow the operations of running, and measuring the results of running. Both implementations are difficult to reason about in various ways.
We are perfectly happy to give money to people for them to do things without fully understanding their inner thoughts and the processes behind those thoughts.
Which is the golden duck?
yes, but we wouldn't say that we're programming them, which is the problem with the operational definition, it applies to everything. If my drunk uncle is great at horse-betting and I just need to give him a nice sixpack of microbrew and get measurable net worth increase out I've not turned into a computer scientist.
Hence my argument that legibility is what matters. Programmers must be able to reason, and rearrange, and understand relationship between syntax and semantics of a program.
I think it's more accurate to compare deep learning to running a sort of physical experiment, rather than programming.
Also, if your uncle is better than your computer, then stop programming it at all. However, if he was actually any good, then he shouldn't be talking to you, and you shouldn't be giving him any beer. Unless he was banned by the bookmaker.
Any insights worth learning about?