Demystifying Differentiable Programming (2018)
arxiv.org
arxiv.org
I haven't seen this discussed before, except in the context of category theory (https://ncatlab.org/nlab/show/differentiation) I'm not sure how this ties into the programming language point of view.
If this "explanation" (scare quotes because it's not really an explanation, just a code dump) doesn't work for you let me know and I can try a different approach.
final case class AD[A](v: A, k: Double => Double) {
// Chain rule
def flatMap[B](f: A => AD[B]): AD[B] = {
val self = this
val next = f(v)
AD(next.v, (d: Double) => next.k(d) * self.k(d))
}
def +(that: AD[Double])(implicit ev: A =:= Double): AD[Double] =
AD(this.v + that.v, (d: Double) => this.k(d) + that.k(d))
def *(that: AD[Double])(implicit ev: A =:= Double): AD[Double] =
AD(this.v * that.v, (d: Double) => (this.k(d) * that.v) + (that.k(d) * this.v))
def sin(implicit ev: A =:= Double): AD[Double] =
this.flatMap(x => AD(Math.sin(x), (d) => Math.cos(x) * d))
def gradient: Double =
this.k(1.0)
}
object AD {
def pure(x: Double): AD[Double] =
AD(x, d => 1.0)
}Shouldn't M actually be a monoid under composition? ie. you should compose like next.k(self.k(d)), not multiply.
So, — × M means something like "type T a = T a M", in Haskell notation.
final case class Mon[A](v: A, m: M) {
def flatMap[B](f: A => Mon[B]): Mon[B] =
Mon(f(this.v).v, f(this.v).m * this.m) // * is M's binary operation
}
object Mon {
def pure(v: A): Mon[A] =
Mon(v, M.unit) // M.unit is M's identity element
}Is there some parenty between the ideas ?
> The implementation proposed by Pearlmutter and Siskind returns a pair of a value and a backpropagator ... Doing this correctly requires a non-local program transformation ... Further tweaks arerequired if a lambda uses variables from an outer scope ... In contrast to Pearlmutter and Siskind [2008], using delimited continuations enables reverse-mode AD with only local transformations. Any underlying non-local transformations are implicitly resolved by shift and reset.
I'll have to look more carefully to understand how the CPS version avoids non-local program transformations.
Thanks for this paper. Along with the string of "compiling to categories" papers, our understanding of AD has improved greatly.
http://h2.jaguarpaw.co.uk/posts/automatic-differentiation-wo...
val x = ...
ys.map(y => y + x)
Each loop iteration needs to contribute a gradient update to x, which is defined in an outer scope. And what if y => y + x is not given as an inline lambda, but defined elsewhere. It doesn't seem like your blog posts discusses any such cases.Could you elaborate on what leads you to think this?
Consider your thermostat. What do you optimize for daily with it? Is it the same in winter as summer? Why not?
To that end, why do we think there is an "optimum" setting?
But I agree that in many cases it is difficult to define a goal.
Week of vacation at home? Probably optimizing for comfort. Standard work week where you aren't home all day? Cost will be big.
Of course, getting costs in there could be a little tricky. Not impossible, just likely a ton of heuristics. And there is some level of stress where maintaining a temp might be cheaper than quickly getting to it, depending on how far you will vary based on being off. Which is just a long way to say it is easier to keep the house warm on warm days. :)
If you drive a lot, this would be akin to trying to find the optimum spot to hold the gas pedal. Sounds like something you could look for, but odds are high that it has to vary based on circumstances.
So, in the end, we aren't looking for a static parameter, but a system of dynamic parameters to keep in tune.
To account for the dynamics / circumstances you can augment the input the thermostat receives from just the current temperature to also include information about the rate of change of temperature (velocity and acceleration). This is the idea behind PID controllers and it allows things like easing off the heating if the house is warming quickly so the thermostat doesn't overshoot the goal.
In particular, my challenge is that in most programs the "loss" function that is of interest to the user is actually not visible. Why I picked the thermostat example. It doesn't know if you are there or not. Or what the cost of electricity is right now. Or what the average temperature outside will be in 2 hours.
It can be made aware of all of these things, but that will itself be a large part of the complexity. Such that "just optimize this loss function" is not wrong. Just, I challenge how much it helps.
https://storage.googleapis.com/nest-public-downloads/press/d...
So, if you throw a whole lot of effort at it today, it can work well! I just want this sort of thing to be more automatic.
The nest is odd, to me, as I live without an air conditioner. So, at best, it can keep things off when we are not at home. But, you can't just turn on when we are home, as at that point it is too late to hit and maintain comfort before bed.
So, we are back to adding more to my house just to support the data and control this is looking to bolster. Which is almost certainly a net negative, all told. :(
For battery optimization use cases, e.g. screen brightness, one metric that works is minimising the amount of times a user goes and changes the settings (with reinforcement learning), and I think that could definitely be a starting point for thermostats.
Obviously, you also care about other things in a thermostat, e.g. energy usage (eg when you're not home), and maybe the amount of changes a user needs to make is not sufficiently tight supervision for a thermostat, but it's a starting point that does take personalization into account.
With more and more cores coming, I expect to see this sort of thing happen once we accept that we can’t really use all of them.
Reminds me of Dynimize https://news.ycombinator.com/item?id=17025627
I think a few compilers sprouted the ability to take perf logs from previous runs into account when compiling, maybe ten years back now, but nothing is there for off-line optimization in situ. And again my choice of input data will never be perfect for all of my users. I can tune for my biggest customer or be populist.