This is standard an unavoidable. There are like a dozen of tricks that solve a few special cases, and they were found after heroic brute force search in the void. (The real fact that is somewhat hidden is that most differential equations can't be solved analytically. You solve analytically only the few cases that are solvable analytically, otherwise you just get a numerical solution or an approximation.)
> In this way there was not a lot to discuss, and from time to time the professor would gently steer us back to merely copying and memorizing the methods, even though no one had questioned him out loud.
That's a horrible way to teach.
> The real fact that is somewhat hidden is that most differential equations can't be solved analytically. You solve analytically only the few cases that are solvable analytically
I discovered this very early in my semester of Differential Equations. We were allowed a single 8.5x11 notesheet for the exams. As there were only a handful of the “most general” cases which are solvable on paper with a basic calculator, I simply copied the step by step solution for each of the very most general case completely worked out in whatever techniques we were going to be tested on for that exam.
The professor was an engineer before becoming a math professor so he only liked to include real-world situation ODE’s on exams which further reduced the potential problem space.
While it greatly confused the professor/grader who scored my exam that I kept adding zero-coefficient terms before solving the differential equation perfectly…I got 100% on all the exams.
The catch was that I didn’t learn anything. The next semester it turned out that I needed to know those techniques for Reaction Kinetics and Heat&Mass Transfer and Biochemical Engineering (these courses involved deriving and solving many equations from first principles).
I had to crawl back to my Differential Equations professors office hours for 3 weeks and beg him to actually teach me differential equations. He was very confused after asking me what grade I got (an A) and I had to explain to him how I got an A without learning anything.
To his credit, he did a fantastic job assigning me custom work for 3 weeks and reviewing it with me and I was able to learn what I needed for the more advanced courses.
But without his help and some additional tutelage from my peers, I would have been completely screwed for the rest of my Chemical Engineering major.
Exactly, so why don't they teach the numerical analysis for actually solving PDEs that matter? These are equations that are very highly relevant to a wide array of real-world science and would be extremely beneficial for many people to know, even if (like calculus or even algebra) most people may not need them later.
I ended up wandering into a career where I work with PDEs nearly every day in some form or other, and would have greatly appreciated some basic training as part of my formal education.
I think there is a little bit of an annoying situation where at least Electrical Engineering students are going to want Differential Equations pretty early on as they are pretty important to circuits (IIRC, I don't touch analog stuff anymore). Like maybe as a first semester 200 level class. This doesn't afford space to put a Linear Algebra class in beforehand (needed for numerical analysis).
Maybe the symbolic differential equations stuff could be stuck at the end of integral calculus, but
1) curriculum near the end of the semester is risky (students are feeling done, and it can suffer from schedule shifts).
2) Transfer students or students who satisfied their calc requirements in highschool (pretty common for engineering students) wouldn't be aware of your curriculum changes.
Or, a numerical-focused PDE class could be added to elsewhere. I bet most math departments have one nowadays, but as an elective.
I wish we had been taught how to use a Computer Algebra System (example, Mathematica or Maple).
Most people will do some computational courses that at least have them solving basic PDEs in their first or second year of undergrad now.
(This reflects the state of those in the UK at least)
Also, in Physics, a lot of ODE are mysteriously integrable if the variable is x instead of t. (One reason is that it's easy to measure the force/fields, but the "real" thing are the potential, so you are measuring the derivative of a hopefully nice object.)
Also a lot of the theoretical advanced stuff to prove analytical solutions and to estimate the error in the numerical integrations use the kind of stuff you learn solving the easy examples analytically.
And also historical reasons. We have less than 100 years of easy numerical integrations, and the math curriculum advance slowly. Anyway, I've seen a reduction in the coverage of the most weird stuff like the substitution θ=atan(x/2) (or something like that, I always forget the details). It's very useful for some integrals with too many sin and cos, but it's not very insightful, so it's good to offload it to Wolfram Alpha.
- Separation of variables. If one is fine with differentials (or their modern cousins differential forms), there isn’t much to explain here.
- Linear equations solved with quasipolynomials. The only ODE-specific observation is that d/dx in the ( x^k e^x / k! ) basis is a Jordan block; the rest is the theory of the Jordan normal form, which makes interesting mathematical points (an embryonic form of representation theory) but exists entirely within linear algebra (even if it was motivated by linear ODEs historically).
- Ricatti equations. Were always a mystery to me, but it appears they could also be called “projective ODEs” to go with linear ones and have pretty nice geometry behind them (even if, as you said, they were first discovered by brute force search).
- Variation of parameters. Despite the mysterious appearance, this is simply the ODE case of Green’s method beloved in its PDE version by physicists and engineers. (This isn’t often included in textbooks, in fear of scaring students with Dirac’s delta, but Arnold does explain it, and IIRC Courant–Hilbert mentions it in passing as well.)
- Integrating factors. Okay, I can’t really explain what that one means, even though it feels like I should be able to.
Not that teaching it like this would make for a good course (too general, and ODEs ≠ methods for solving ODEs), but that’s essentially it, right? There are certainly other methods you could mention, and not unimportant ones (perturbation theory!.. -ries?), but this basically covers the standard litany as far as I can see. And it’s no haphazard collection of tricks—none of these is just pulling solutions out of a hat.
(In the interest of changing things up and not spending an hour on a single comment, I will omit the barrage of references I’d usually want to include with this list, but I can dig them up if somebody actually wants them.)
Here the first ODE course is half a semester. If you spend a week or two proving existence and unicity, you get one week to study each method and make a few examples and then you must change to next week trick.
Fourier/Laplace and other advanced stuff are in a more advanced course.
I never used perturbation theory for ODE. I've seen it for solving eigenvalues/eigenvector of operators in QN. But perhaps it's one tool I don't know.
Those that have taken Integral Calculus may be thinking that solving DEs sounds akin to integration where one may have to apply substitutions, integration by parts, trigonometric substitutions, or partial fractions. Yes Calculus requires learning a bag of trick too, buts its a small bag of simple tricks with wide applicability. So many of the functions one needs to integrate succumb to this small bag of tricks that it's almost fun to hone ones technique. A class on elementary differential equations is just depressing.
To be fair, differential equations are important. Physical phenomena are often best described by differential equations. Fortunately, programs like Mathematica can be used to tackle real world differential equations one way or another (perhaps with numerical methods) to obtain solutions.
I was fortunate to have my Probability course (sadly, not my differential equations course) taught by Gian-Carlo Rota.
I've been coding for years and have been able to fake it with my limited math education but would love to have the time to learn more for the sake of understanding.
I found the classes to be rote. The derivations are truly non-trivial. The book Ordinary Differential Equations by Arnold goes into more detail. Basically if we taught the reasons we'd require everyone to take analysis and differential geometry to truly understand how they work. Given the MAJORITY of students in diffeq are engineers and not math majors 99.9% don't want to know and/or don't care about this detail. You see a similar occurrence in calculus where you're basically told "dont think about it too hard" for your own safety. If you start wondering a little too hard about calculus you end up switching majors to math and taking two semesters of real analysis. It's also EXTREMELY common for engineering professors to teach differential equations rather than math professors. This further waters down the rigor because (obviously) an engineer will not know/care about the rigor. Part of the reason I've pursued a math degree is because there was so much handwaving in engineering/computer science it became just an extremely annoying grab bag of math tricks and I wasn't satisfied.
To me we have too many inter-dependent classes to teach each class with full rigor. As a result you end up with a collection of half-understandings for most of your undergraduate career and only if you take a math major itself (or a minor in math) will you actually unlock the other half. A better path through math might be basic algebra I, II-> geometry -> trig -> abstract algebra I+II -> analytic geometry -> calculus I, II, III -> real analysis I+II -> differential equations I+II, but this would basically make every degree a math degree. What you experienced is the compromise.
Back propagation is (almost) just a fancy word for differential equation, with derivative relative to the error in the output against your training data.
Could be that someone else here remember the exact video
A few resources here:
An overview, with a bias towards finance: https://informaconnect.com/a-brief-introduction-to-automatic...
On the history: Andreas Griewank, Who Invented the Reverse Mode of Differentiation? https://ftp.gwdg.de/pub/misc/EMIS/journals/DMJDMV/vol-ismp/5...
On the history of back propagation: https://en.wikipedia.org/wiki/Backpropagation#History
The article that introduced it to finance: Michael Giles and Paul Glasserman, Smoking adjoints: fast Monte Carlo Greeks https://www0.gsb.columbia.edu/faculty/pglasserman/Other/Risk...
Survey of the application in finance: Cristian Homescu, Adjoints and Automatic (Algorithmic) Differentiation in Computational Finance https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1828503
Backpropagation ≠ Chain Rule: https://theorydish.blog/2021/12/16/backpropagation-≠-chain-r...
I also didn't totally grasp its significance until implementing neural networks from matrix/array operations in NumPy. I hope all deep learning courses include this exercise.
Look into forward- vs reverse-mode automatic differentiation, and you'll understand what I'm referring to.
Linear regression isn't just fitting a line, it's a statistical technique to fit a line of best fit. Hyperparameters are a bayesian term for parameters outside the system of test or "algorithm". User input really misses the bayesian aspect.
These terms actually have meaning so I'd be careful ascribe simpler definitions. The underlying meaning is important to the reason they work. If you don't have a really strong background in probability theory and statistics trying to dig into machine learning will take work. Id recommend taking an MITx course or picking up a textbook on probability so the terminology feels more natural.
Yes, hyperparameters are often set by the user of a model, but more specifically they are parameters that exist separately from the data put into a model (input parameters) or the structure inside of neural networks (hidden parameters). Hyper- meaning above, helps conceptualize these parameters as existing outside the model.
We have to memorize a lot of information without the explanation about why is done in that way (due to the lack of time in those subjects), and also we are more encouraged to study how previous year's exam were made than the content itself. This one of the big reasons only 10% to 15% (IIRC) of the enrolled students pass those exams every year.
That scene, knowing that I have to do a task that is time consuming, pretty hard, artificial, and useless for the rest of my academic life, my work life, or my life in general, is what made me leaving this year. I don't have enough mental health to do such a big thing.
PS: Sorry for the rant. I'm having too much time at home due to COVID and maybe wrote too much.
I am middle aged and completed my EE degree when I was 20, but it was 90% theory with very little practical use (mostly useful if you were to continue climbing up the education chain). Completing the degree made me despise working with electronics, a topic I had deeply loved and had spent my teenage years learning for myself. Most courses were rote learning, and I was very good at passing exams, but it was two years before I realised how pointless the majority of the “knowledge” was, and then I forced myself to finish the degree (sunk cost), which I now regard as one of the few true mistakes of my life (wasted years, for valueless academic “knowledge”). The degree got me a software job, so there is that, but I am sure I would have ended up in software anyway (early love of computers).
The required memorization made things especially difficult for me because I tend to work off intuition rather than memorization. I also usually can't name theorems despite knowing them from practice (this also used to be a huge pain for exams where solutions were unreasonably only considered correct if you named everything used).
In high school, I looked into analytically calculating a ball's maximal trajectory length (or something like that) and was told it required solving differential equations and would be taught in college.
paving the way (or building a wall) such that few can understand how people came up with that stuff. this is intended. this literally constructs knowledge as power.
The ways of thinking used to come up with the techniques are hidden, restricted. The academics who know the whole story (who know the ending -- which is what is taught, as well as how mathematicians of old came up with such ideas) hold this kind of power.
This gets even more interesting when the academics who know the histories, cannot really use the techniques. then the only people who knew both are historical figures (who get bathed in myth).
I cannot forgive them for this, given as they are still actively doing this. e.g. finding out how they make shredded wheat cereal is not possible [1]; and this must be technology from the early 20th or late 19th centuries... anything more recent is just hopeless.
the relation is ideological, cultural (in the sense of being close to the intention of); not direct, causal, material (in the sense of relating to the actual implementation).
One of the standard methods is "integrating factors for first-order linear equations". You are told that, faced with an equation
y' + p(x) y = q(x) you should multiply both sides by e^(the integral of p(x)).
For example, you might have y' + (2/x) y = x.
Then you multiply by e^(integral of 2/x), which is x^2.[Sometimes I wish Hacker News had TeX available.] If that's all you tell people, it looks like some random abracadabra and it's no wonder why people feel they just don't get it. So you might try to explain this way:
"The equation has a derivative in it. To undo a derivative, you need to integrate. But if you integrate as-is, you have no idea how to integrate y' + (2/x) y."
"Well, you know that the integral of df (the derivative of f) would be just f. So if you could make the left side look like the derivative of something, then you could just integrate both sides."
At this point, you scratch your head and think: "What could I do to make the left side be the derivative of something?" This kind of thought is impressionistic - you have to think in a vague way of things the left side "is like". Daydreaming for a while, you might realize it's a sum of two terms, so you think: If this is a derivative and it's the sum of two terms, what derivative rule gives a sum of two terms? And you might think of the Product Rule.
But the given thing is not the derivative of a product as is. What to do? So continuing this line of thought, you might think - maybe I can multiply it by something to make it the derivative of a product. Once again, you have to search through your experience with derivatives and maybe mess around on scratch paper. Finally, you realize x^2 works - multiplying by x^2 makes the equation
x^2 y' + 2 x y = x^3.
The left side is d(x^2 y), so you can integrate both sides and get x^2 y = (1/4) x^4 + c.The final step is to think whether you can generalize what you did with "p(x)" instead of "2/x". After some additional messing around, you come up with the integrating factor I gave at the start.
I have no idea who discovered this method, or what their thought process was (if they even explained it at the time). This was about the extent of the motivation I got when I was taught this stuff in high school/college. I'd tell students this sort of thing when I taught differential equations. But I don't know what other people do in teachng, and I'm not sure this helps. For people who feel their differential equations courses were baffling/unmotivated, is this the kind of explanation you want? Or do you want something completely different, like applications?
At some point, explanation ends. Can a painter say why he put a daub of paint of that color in that place in a painting, or can a writer say why he had a character do this or say that? There are several points in the motivation above where all I can say is "you have to sit there and think and mess around", even after paragraphs of writing. I'm not sure how to do better.