Gradients are not all you need
arxiv.org
arxiv.org
Lots of caveats there, of course. First off, I don't know much about the neurology, I just have an amateur interest in second language acquisition research that sometimes brings me into contact with this sort of thing. On the ANN side, which is closer to my actual wheelhouse, we definitely don't actually have any way of knowing if the actual mechanism is all that close, and I'm guessing it probably isn't even close since ANN's don't actually work that similarly to brains. Nor does it need to be, but, intuitively, there's still something promising about an ANN architecture that's vaguely capable of mimicking the behavior of modules in an existing system (human brains) that's well known to be capable of doing the job. I'm not super wild about the bidirectional recurrent layers, either, because they impose some restrictions that clearly aren't great, such as the hard limit on input size. et cetera. But it still strikes me as another big step in a good direction.
Seems relevant to what you're working on. It starts with a randomly initialized, overparameterized neural net, but instead of gradient descent backpropagation, it learns by deleting connection edges.
Where could a person learn more about these?
But it's equally important to create architectures that allow efficient backpropagation of errors.
It does seem like transformers are pretty good at both, already.
I kind of hope we're not getting much something radically better anytime soon, because it seems like AGI is already approaching faster than we can prepare for.
Then again, I would expect that someone somewhere is already using transformer based networks to develop some brand new architecture that does in fact provide such a leap.
What are "well-behaved" gradients?
What type of GPT-4 level stuff?
Sometimes you need to peek at the Hessian.
Seriously though, what is intelligence if not creative unrolling of the first few terms of the Taylor expansion?
(And it goes down all the way)
The cartpoll demo famously tripped up derivative based reinforcement learning for awhile.
Did you mean "Global optimization techniques which do rely on gradients..."? Because exact gradient-based global optimization (GBD or branch-and-bound based) methods for general nonconvex problems are theoretically superior (bounding with McCormick relaxations etc.) but also more challenging to practically deploy than say stochastic methods or metaheuristics like local search.
What does this refer to: cartpoll demo famously tripped up derivative based reinforcement learning
The phrase "cartpole demo famously tripped up derivative-based reinforcement learning" is likely referring to a classic problem in the field of reinforcement learning, which involves balancing a pole on a cart. The pole is attached to the cart via a hinge, and the goal is to keep the pole upright by moving the cart left or right in response to its angle. This problem is often used as a benchmark for testing reinforcement learning algorithms.
The phrase suggests that derivative-based reinforcement learning algorithms, which rely on computing gradients of a function with respect to its parameters, were not successful at solving this problem. This could be due to the fact that the problem is highly non-linear and requires precise control, which may be difficult to achieve with gradient-based methods.
Edit: bard got it too, with more detail, which is surprising
- ChatGPT was trained on a corpus of data containing the answer and is able to give you a decent answer
- ChatGPT was never exposed to the answer and will hallucinate a plausible-sounding response, and because it will answer in a really convincing way, you'll get tricked into believing complete bullshit
Instead, using stochastic gradient approximations (policy gradient method such as proximal policy optimization) might be better suited to solving these kinds of problems. Effectively, they do not compute the exact gradient locally, but rather kind of a global approximation by trying out random sequences of actions and determining which of them are closest to the desired outcome.
Hence, stochastic gradient approximations might be considered some kind of hybrid between greedy local optimization (such as following the exact gradient) and global optimization.
The reason for this is that the algorithm doesn't like to have to "spend" energy, reducing its score. Without huge amounts of trickery to get the gradient descent algorithm to stop getting stuck in the center, this is never solved - due to using a local optimizer for a global optimization problem (finding good weights in a NN)
I was always wondering about a question that is touched in the paper: despite the analytic gradient computation being intuitively more efficient and mathematically correct, it is much harder to learn a policy with it than with the “brute force trial-and-error” black box methods.
This paper brings many new perspectives on why.
I studied this in undergrad, but it’s not the same thing the paper is talking about
Whether quirky titles help or hurt with that i think depends on name recognition and the publication venue
The problem isn’t quirky titles, it’s with websites like Hacker News that only display a headline and not the abstract.
This means that people who have an up-to-date knowledge of a given subfield can quickly get a lot out of a new papers. Unfortunately, it also means that it usually takes a pretty decent stack of papers to get up to speed on a new subfield since you have to read the important segments of the commonly cited papers in order to gain the common knowledge that papers are being diffed against.
Traditionally, this issue is solved by textbooks, since the base set of ideas in a given field or subfield is pretty stable. ML has been moving pretty fast in recent years, so there is still a sizable gap between the base knowledge required for productive paper reading and what you can get out of a textbook. For example, Goodfellow et al [1] is a great intro to the core ideas of deep learning, but it was published before transformers were invented, so it doesn’t mention them at all.
Fixed with ML.
Bing chat does have the ability to read opened PDFs when used from the Edge side bar.
But… yeah, its hardly a title chosen for marketing rather than description.
https://www.harmdevries.com/post/model-size-vs-compute-overh...
This website has an interesting graph which visualizes this:
> For example, the compute overhead for 75% of the optimal model size is only 2.8%, whereas for half of the optimal model size, the overhead rises to 20%. As we move towards smaller models, we observe an asymptotic trend, and at 25% of the compute-optimal model size, the compute overhead increases rapidly to 188%.
So if you train your model you shouldn't just look at compute-optimal training but try to anticipate how much the model will probably be used for inference, to minimize total compute cost. Basically, the less inference you expect to do with your model, the closer the training should be to Chinchilla-optimality.
Branding is eating the world.
Not realizing that it always has is why it's still effective.
History is written by the victors and victors tend to have the best marketting.
40 years ago, when articles went into print publications, you'd just get your paper into a key print journal and then trust that everyone who gets it would at least look through the article headlines and read the abstracts of articles that seemed relevant to them. And it was manageable because you'd only have a few new issues rolling in per month.
But arXiv had an average of 167 CS papers being submitted per day in 2021. An academic who wants to keep their career alive needs to resort to every trick in the book to be heard above that din.
Curation kind of requires gatekeepers, and good gatekeeping is a full-time job that needs to be paid for. In the days of print journals, there was both a motivation, because you can only print so many articles per month due to materials costs, and a steady revenue stream to pay for it, because everyone has to pay for the journals to get any access to articles at all. The Internet's been eroding both of those factors for a good quarter century now.
I think we’re seeing something closer to title optimization in the service of marketing. But even marketing isn’t a great fit.
Upon reflecting on the Wikipedia definition of marketing, it strikes me that title optimization is only a tiny part of what “academic research marketing” could be. In the most generous sense, it could mean reaching out to media sources fairly early in the research/writing process to help craft a paper that will be interesting to the desired audience.
Wikipedia quotes:
> Marketing is the process of exploring, creating, and delivering value to meet the needs of a target market in terms of goods and services; potentially including selection of a target audience; selection of certain attributes or themes to emphasize in advertising; operation of advertising campaigns; attendance at trade shows and public events; design of products and packaging attractive to buyers; defining the terms of sale, such as price, discounts, warranty, and return policy; product placement in media or with people believed to influence the buying habits of others; agreements with retailers, wholesale distributors, or resellers; and attempts to create awareness of, loyalty to, and positive feelings about a brand.
> A brand is a name, term, design, symbol or any other feature that distinguishes one seller's good or service from those of other sellers. Brands are used in business, marketing, and advertising for recognition and, importantly, to create and store value as brand equity for the object identified, to the benefit of the brand's customers, its owners and shareholders.
> Personal branding is the conscious and intentional effort to create and influence public perception of an individual by positioning them as an authority in their industry, elevating their credibility, and differentiating themselves from the competition, to ultimately advance their career, widen their circle of influence, and have a larger impact.
Sorry, are you referring to brandpocalypse?
But then the title can become something catchy that will give you more visibility.
So how do we get more desirable system behavior? It seems we have a collective action problem.
I dunno if it even read the article lols
====
Title: “Deep Reinforcement Learning for Multi-Agent Navigation in Complex Environments” This paper proposes a deep reinforcement learning approach for multi-agent navigation in complex environments. The proposed method uses a centralized critic and decentralized actor architecture, where each agent has its own policy network and learns from its local observations, while the critic network takes in the global state information to provide a centralized value estimate. The method is evaluated on a variety of benchmark environments and achieves state-of-the-art performance compared to existing methods.
"Sometime’ is Sometimes ‘Not Never’" - Lamport
"Reflections on trusting trust" - Thompson
"On the cruelty of really teaching computer science" - Dijkstra
I'm sure there are more
PS: [1] is a good talk in general where he discusses some of the limitations of the paper and things that could have been done better.
[1] https://www.youtube.com/watch?v=YZ-_E7A3V2w
[2] https://papers.nips.cc/paper_files/paper/2018/file/69386f6bb...