18 karma · joined January 17, 2020
Great, they've showed their intentions and we should respond accordingly. Don't call it GPT, call it parroting (or literally whatever you want). I like this for its connection with the Stochastic Parrots work and because it is already a common usage of the word outside of ML/AI. AutoParrot, BabyParrot, ParrotAPI, etc. We train by first parroting a large natural language corpus and then fine tune the parrot model with RLHF.
Companies will do what companies do, but communities work better when they are free (as in thought).
Wouldn't this be more like saying "if you do this, we will make things painful for you"? Am I completely missing something here?
There are a lot of open questions here, so anything I could say about the brain itself would be more of a guess. That said, for our proposed model no negative probabilities are needed, as the distribution is represented by a population of estimators for different predictors of value (in the general sense).
Hope that makes sense and helps clarify.
We can think about asymmetric regression more generally. If you have an error and apply some 'response' function f to that error you change the estimator you learn. In the case of quantile regression f is a sign function, expectile regression it is identity.
In my opinion, and this is entirely speculation, I think with further experiments more completely studying the effect we found in our paper, that we will find the response function (f) in the brain is not linear, but a type of saturating function like if we smoothed the sign function out. We repeated our experiments in the paper using such a function, which has been proposed for dopamine neuron responses before, and the analysis continues to hold because the rewards are all quite small and likely simply in the linear region of a non-linear response function (we know firing rate saturates eventually so this isn't much of a surprise).
Regarding quantiles being more commonly used, it's actually the other way around. The Huber-quantiles we saw perform best in the QR-DQN paper, and which most often get used in the follow-on RL work, are actually more like the type of saturating non-linearity you might expect in the brain (although the Huber loss is not as smooth as you probably would expect the neuron response to be).