And it's because you need long division to do division in algebra
After looking at the English Wikipedia I found that
a) what I learned in school is long division
b) the notation with the dividend and the divisor separated by a parenthesis and the result on top of a horizontal bar is totally alien to me
c) I've never seen short division before
d) the German Wikipedia doesn't even have an article about short division
Edit: although if someone else can explain it better I'd really like to know.
From what I have searched,(quickly mind you), it seems to be a mind set.
I went to school in Sweden where we do teach all this things, but in a strange order.
the "short" notation relies on the ability to squeeze the remainder in between the digits of the dividend. if you have a divisor greater than ten, it's possible to have a remainder with more than one digit. two or more remainder digits is a lot harder to write legibly between the digits of the dividend. you also end up subtracting larger numbers in your head, which is error prone. finally, most people only memorize the multiplication table up to 12x12 or so. once they can't simply do a lookup into their memorized table, they switch to more of a "guess and check" approach, where they will inevitably have to cross out or erase their previous work.
in the end, it introduces a lot of unnecessary opportunities to make mistakes and hides stuff that people can't do reliably in their heads.
I have a PhD in computer science but I cannot understand long division.
I think you're trolling us a little: you could sit down with pencil and paper and think about how it must work and 20 minutes later you'd understand it, surely.
I have a similar mental block for standard deviation - I always seem to have another ‘but why?’ question that eventually people answer with ‘because!’ and then we’ve both given up.
https://web.stanford.edu/class/ee486/doc/chap5.pdf
Edit: standard deviation is just a fancy name for root mean squared over the sample offsets. The square is there to deal with the fact that there are positive and negative offsets. The mean is the whole point of the exercise, and the root is to negate the fact that we squared earlier.
This is how I always find maths and particularly statistics - I always just want to answer ‘but why’ and then people lose interest in explaining after a while.
Standard deviation is mathematically easier to work with though and has nice properties e.g. you can compute the standard deviation of combining a bunch of independent things by combining their individual standard deviations.
Also, consider that the correlation coefficient consists of similar formulas. They all look like the theorem of pythagoras. This ties in nicely with the fact that observables can be seen as having angles in between them where the correlation coefficient is the cosine of these angles.
You can read up on the central limit theorem if you want to understand more.
This is also where I get frustrated.
I'm taught about the CLT but every place I think I might use it, the samples aren't really IDD when you look hard enough, so it doesn't apply.
Computer benchmarking (which is where I'm usually trying to apply statistics) is the classic example. One iteration changes the state of the computer for the next iteration. They aren't remotely IDD, but people still try to claim the CLT. Seems like all their stats beyond that point is broken.
I'd assume that computer benchmarks have very low correlation between runs, unless you leak ram or your components don't keep a stable temperature. So in that case the CLT seems to apply just fine.
And they are not in a computer program. The state of, to just give one example, the processor cache depends on what happened in the previous run.
We have all this rigour in how statistical techniques are applied, and we make black or white decisions based on the numbers that come out of them, but at their very foundation we're saying 'whatever I'll just pretend this rule applies it's probably fine'. How can we build on that flawed foundation? I don't get it!
"All models are wrong, but some models are useful." Just remember that there's some imprecision in the answer a model gives you, and keep that fact in mind when applying that answer to make real-world decisions.
e.g. https://socratic.org/questions/what-is-the-difference-betwee...
Put simply, standard deviation is arbitrary for most people.
Speaking of language, have you ever thought about how all names and grammatical structures are arbitrary as well? There are parallels.
My larger point is that, something can be arbitrary and still be meaningful. If people agree on a common meaning, it formalizes a communication protocol that allows for much better information sharing.
This means that any distribution with standard deviation of X will behave as a normal distribution with standard deviation X when you have enough of them. Or in other words, any distribution with standard deviation of X will behave like any other distribution with standard deviation of X when you have enough of them. There is no other measure like this as there is no other equivalent to the central limit theorem. Therefore it makes sense to have this as a universal measure of random processes.
Now that we have a normal distribution, the mean is an obvious metric to pick because it captures the notion of 'middle' in a useful sense. Then we have proofs that we can characterise the normal with the mean and 1 other parameter (normal normal) or a matrix (multivariate normal). We call that the standard deviation^.
The exact formula wasn't a coincidence. The normal can obviously be characterised by the mean and a statistic from another formula. There were a bunch of experiments tried (eg, using |x| instead of sqrt[x^2]) but it turned out that sqrt[x^2] had some other nice property that minimised some sort of error so they went with it as a standard. I forget what one, might be error of estimating the true parameters from a sample or similar.
We could characterise the the normal as an infinite sum or something quirky, but when people say 'easy to work with' they mean instead of a function or something quirky we can simply pick a number.
Standard Deviation isn't as important when working with non-normal distributions, although I think it still turns out to be useful. But its importance is that it characterises a normal apart from the information captured in the mean. I'm not a mathematician, YMMV, could be wrong, standard disclaimers.
[0] https://en.wikipedia.org/wiki/Central_limit_theorem
^ I'm not going to edit this but it occurs to me that we call it the Variance. Same thing as std. dev in my opinion.
' There is a fairly subtle observation to make - Normal is characterised by mean and std. dev, but the most efficient ( https://en.wikipedia.org/wiki/Efficiency_(statistics) ) unbiased estimator of the std. dev of the population is the adjusted std. dev of the sample. Therefore, accounting for mean, you can't get a more efficient characterisation of the normal than mean & std. dev. Ie, if you picked a formula other than std. dev then the most efficient estimators to characterise it would still be mean and std. dev. Don't recall if there are equally efficient choices but I think that proves there are none better.
That might have been the logic for why std. dev was chosen. Just a guess.
There are many "right" metrics - this is one of them. The big picture is we want a measure/idea of how spread out the data is. One can come up with many ways to do this, and absolute value and root mean square are two of the most common. They correspond to L0 and L2 norms. There is also the L-infinity norm (just use the farthest point).
The L2 norm is the most complicated of the 3, but as people pointed out, it is one of the easiest to use analytically because of its nice mathematical properties.
Have a look at https://en.wikipedia.org/wiki/Statistical_dispersion for other measures.
To illustrate, perhaps an example closer to home:
If someone asked you to tell them "how many lines of code in this project?" you'd start wondering things like "do I count comments?", "do I count dependent libraries?", "perhaps some metric that is equivalent of a statement count but not strictly counting lines?" and on and on.
Statistical measures are like that. Someone with a decent knowledge of the problem space came up with the best stab at how to characterize the data.
You can. It's called the mean absolute deviation. Which one you use depends on your application. If you just want to give a summary of how spread out the data is then either would work. Lots of people say that we should use the mean absolute deviation as the default rather than the standard deviation.
I too have a problem with math tools passed down without any surrounding context - without telling why are we using this formula, instead of any other variant from the family of formulas that would satisfy the same goals. I'd have much easier time dealing with statistics in school if someone told me that a) deviation with ABS instead of root-square is also a thing, and b) we use the root-square one because it amplifies offsets from the mean.
Also one of my peeves. However over the years I've come to realize that in addition to teachers who omit the context, there is also a class of student that actively doesn't want to hear the context. I'm not sure exactly why this is, but I see it in my immediate family quite a bit. A sort of <cover ears, lalalala...too much detail> kind of thing perhaps due to difficulty taking on too much information?
I feel like the standard deviation is a black box to basically everyone without statistics degree. It also bothers me that it is so widespread, because almost nobody know how to interpret this value. It seems to me like most scientists treat the standard deviation as a magical number that allows to compare spread between datasets and otherwise doesn't mean anything on it's own. I feel like a lot more insight could be gained from the mean absolute deviation.
I'm also curious how are scientific insights affected by defaulting to something that is extra-sensitive to outliers. The effect can't be big per single research, but in collective?
Unless you try to solve more problems yourself, you won't understand the advantages or disadvantages of some methods in solving the problems. So the real answer is: construct the examples, compute and compare, don't approach it "philosophically."
All the methods used today survived because they were the solutions to some problems. It doesn't mean that everybody applies them properly: for that you have to get some experience yourself.
Learning about the historical development of the methods is also for me satisfying experience: e.g. logarithms aren't a concept devised for philosophical purposes, Napier developed them to save astronomers time and limit "slippery errors" of (their) calculations:
https://www.thocp.net/reference/sciences/mathematics/logarit...
"I have a PhD in computer science but I cannot understand long division"
For that, personal experience in devising the examples, evaluating them and comparing the results is the only way to get the understanding. One can't complain that one doesn't understand Greek texts if one personally never tried to learn the Greek letters. One could have learned "about" Greek but at the end one hasn't done the inevitably necessary preconditions to actually read Greek texts.
The error in thinking is believing that because one already learned about something else somewhere else he should somehow "understand" something without doing the necessary work on that something. The solution is "make your own homework" to fill the gaps.
The topic of what, how and why children learn in schools is much more complex.
Edit: Also, now reading the Wikipedia article: https://en.wikipedia.org/wiki/Long_division I do understand why people don't understand it: the taught notation for the process, used in many countries, is utterly confusing to me, so my reason for confusion when looking at various notations is not that I don't understand the nature of the method but that I haven't spent time analyzing what was built on top of the basically simple idea -- I can also "understand" all the notations but I'd also need more time for that. Still, having enough experience, I can claim that I understand the idea and the process even if I never learn all the notations and conventions in different countries (which part of the process you write where, what you write and what you don't when etc). The idea is simple, the conventions enforced don't have to be. Those who only learned the conventions maybe never invested any time to figure out what's behind the conventions: what is actually being done.
This means the quantity of mean/variance depends on your choice of units, whereas the quantity of mean/standard deviation. Is constant no matter the units.
Why do we care about variance? Because it is a nice and linear property. This means it is easy to calculate and manipulate.
Moreover, the variance / standard deviation very nicely describe a Gaussian distribution (bell curve). This is a very important distribution because of the law of large numbers. Because we see bell curves so often, it is nice to have tools (std.dev) that work well with these curves.
It should be noted that, in optimization problems, there can be reasons to try and minimize something other than variance. We tend to pick the variance / std.dev because it is familiar, easy to work with, and very efficient. Notably, the derivative of variance tends to be linear, which makes it pretty efficient to use in gradient descent.
I think the OP's frustration in how computation of standard deviation as a quantity is just taught by rote in a cookbook manner is quite understandable. If one's not thought more in depth about distributions and estimation of their parameters, it really makes little sense.
Whether the statistics should be taught more in depth is another question. The current state of the mathematics of statistics and probability is such that it needs quite heavy tools to do rigorously. And countless of hours of getting to know the quirky inconsistent notation.
But until it is taught more rigorously, more intuitive measures of dispersion like MAD make a lot more sense for most applications. And I don't think it would be an unreasonable ask for the statisticians to lay out the "porcelain" for working with the more intuitive, if somewhat more analytically inconvenient measures. A bit like we don't require programmers to understand how transistors work.
Take a course in ring theory if you want to see how all of this machinery gets built up.
also maybe they are specific about they way they do long division in the US? it was a surprise to me in college too and i have a french background. it's not how the french system does long division...
6240 / 5
= divide(6240, 5)
= divide(1240, 5) + 1000
= divide(240, 5) + 1000 + 200
= divide(40, 5) + 1000 + 200 + 40
= 1000 + 200 + 40 + 8
= 1248
----------
2. Regarding standard deviation, one of these bullets might help:
- The normal distribution has exactly one shape, centered at x=0. But it's useful to apply two transformations to it: translation and horizontal stretch. To translate a distribution left/right, change its mean. To horizontally stretch a normal distribution, change its standard deviation.
- Mean has units of length. Standard deviation has units of length. They tell you where the normal distribution is offset, and how wide it is. Mean and standard deviation are just measuring sticks/rulers for normal distributions.
- When people talk about the standard deviation with any arbitrary data, they're usually assuming the data is normally distributed. If the data is not normally distributed, standard deviation no longer refers to the width of the normal distribution, so we lose that visualization.
- With non-normally distributed data, the standard deviation is still useful as an analytical tool because taking the (sqrt of the) summed squares still gives us a number that grows as the data spreads further apart or if the distribution grows wider. There's a center point for the data (the mean), so to ensure you're measuring the overall spread of the data, you subtract the center point off each data point before squaring them. In other words, you're squaring deviations from the mean. And unlike the sum of absolute differences, the sum of squared differences (variance) is differentiable. Differentiability is a great property, so this is the standard way to compute a sum of deviations from the center.
func Longdiv(A, B) {
// Assume A > 0, B > 0.
N = floor(log10(A))
Q = 0
R = A
while (N >= 0) {
// Invariant: A = B * Q + R
// Invariant: B * 10^(N+1) > R
C = 0
while (B * (C + 1) * (10 ^ N) <= R) {
C += 1
}
// C is the largest C such that B * C * 10 ^ N <= R
R -= B * C * (10 ^ N)
Q += C * (10 ^ N)
// R is still positive.
// Invariant A = B * Q + R is maintained.
// We know B * 10^N > R because otherwise we
// would have picked a larger C.
N -= 1
// Invariant B * 10^(N+1) > R is maintained.
}
// By loop invariants, we have
// A = B * Q + R and B * 10^0 > R.
return (Q, R)
}
This is ten lines, not “astronomically complicated.”I mean your very first line has a logarithmic operation that was not mentioned anywhere in your text or your comment! Just pops in there out of nowhere. I guess it's about the number of decimal digits? But why? Why are we doing things in decimal?
A dense ten-line algorithm like this seems far more complicated than other algorithms we try to get school children to memorise.
We’re subtracting off multiples of 10^N because that’s easy when you write your numbers in base 10. Try subtracting off multiples of 9^N, or multiples of N!, or some other choice, and you’ll see why.
His method for polynomials/derivatives is also interesting if you watch until the end, dead simple calculus.
That could be meant shallow or deep. The shallow answer is because that's all you need to do to function in society (and conversely, while arguing with a police officer about the presumed base on speed limit signs may be fun, it is also pointless).
The other answer probably needs to explore the question a bit further. Perhaps a good starting point is the fact that we almost universally share the physical characteristics of having ten fingers.
I think this is maybe the crucial part that you might be missing. All this long division (and long multiplication) stuff works on base-n represenations of numbers. I.e. if you have an number like 3376, it is actually a short hand for
3*10^3 + 3*10^2 + 7*10^1 + 6*10^0.
And if want to divide it 4 and, suppose, you cannot do it in your head, you do it step by step, by clever regrouping with the distribute law: (3*10^3 + 3*10^2 + 7*10^1 + 6*10^0) / 4 ==
3*10^3/4 + 3*10^2/4 + 7*10^1/4 + 6*10^0 / 4 ==
(3/4 does not work, so let merge the first two again) (3*10^3 + 3*10^2)/4 + 7*10^1/4 + 6*10^0 / 4 ==
(33)/4*10^2 + 7*10^1 /4 + 6*10^0 / 4 ==
(now we have progress) (33)/4*10^2 + 7*10^1 /4 + 6*10^0 / 4 ==
(32+1)/4*10^2 + 7*10^1 /4 + 6*10^0 / 4 ==
(32)/4*10^2 + (1)/4*10^2 + 7*10^1 /4 + 6*10^0 / 4 ==
8*10^2 + (1)/4*10^2 + 7*10^1 /4 + 6*10^0 / 4 ==
(now merge the 1 and the 7 group) 8*10^2 + (17)/4*10^1 + 6*10^0 / 4 ==
8*10^2 + (16+1)/4*10^1 + 6*10^0 / 4 ==
8*10^2 + 4*10^1 + (1)/4*10^1 + 6*10^0 / 4 ==
8*10^2 + 4*10^1 + (16)/4*10^0 / 4 ==
844
The same works for other bases. If you were to implement some bignum library, you would also choose some base n representation for you numbers. Base 10 is not so optimal for computers, so maybe you chose base 2^32. If you then were to implement a division function, you would use similar algorithms.The idea of long division is just that to solve A/B, we can find another problem C/B we know the answer to and break it into the smaller problem A/B = (A-C)/B + C/B.
For example, to solve 742/13, the normal way is to observe that 13×5=65 is as close as we can get to 74 without going over, so we can reduce this to
742/13 = (742 - 13×50)/13 + 13×50/13
= (742 - 650)/13 + 50
= 92/13 + 50
Now we only have to solve the simpler problem 92/13. All the long division stuff is just a book-keeping method for this data.But you could also use 13×4=52 and work out
742/13 = (742 - 520)/13 + 40 = 222/13 + 40
222/13 = (222 - 130)/13 + 10 = 92/13 + 10
92/13 = ( 92 - 91)/13 + 7 = 1/13 + 7
This gives the same answer, but it goes 40->50->57. The normal long division goes 50->57, so it never "goes back" on a digit it has decided on.35/350 (to simplify) Once is 325, twice is 300, X, X, X, we end up at 10 times is 350 divisible by 35.
I'm wondering if I'm missing something (and if I am I'll own up to it lol)
https://commons.wikimedia.org/wiki/File:LongDivisionAnimated...
In essence, repeated subtractions. The blank areas to the right of each number are really just zeros.
https://www.solipsys.co.uk/new/SquareRootByLongDivision.html
NN = (10d + e)(10d + e) = 100dd + 20de + ee
Group the digits of the square by two, and add extra values to the "divisor" so that the extraction of the square root in effect undoes the squaring above.
See, for example, https://sciencing.com/calculate-square-root-hand-5081134.htm...
For anyone interested in going down the rabbit hole of those infinite levels.
Afaikr it's just division, how man X are in y.
083.28
2.26...
___
7 |583.00
That's short division (sometimes called the "bus stop method" in UK), top line is answer [quotient], next line is "remainders". You say "7 in to 5 won't go; 7 in to 58 is 56 [just from knowledge of times tables, it's 56 tens your dividing], with remainder 2 [write remainder down, usually as a superscript to the dividend]; 7 in to 23 goes 3, remainder 2 [write remainder down]". Now you have the answer 583/7 is 83 remainder 2; but you can continue and divide the 20 tenths by 7, and so on.Long division:
083.28
2.26...
___
7 |583.00 [<-dividend]
56 [=8x7]
--
23 [2 from subtracting answer to 8x7 {8 is put in answer line}, 3 from the dividend]
21 [=3x7]
--
2.0 [2 is remainder, 0 from dividend]
1.4 [=2x7 is remainder, 0 from dividend]
---
.60
.56 [=8x7]
--
4 [is the remainder in 100ths]
What we're doing is taking five-hundred and saying can we divide that by seven-hundred, we can't. So then we say well how about fifty-eight tens, can we divide that by seven-tens. The tens cancel each other out ( 10/10==1 ), so 58/7 = 8r2 but this is really saying 580/70 is "80 lots of 7" and 20/7 left over. So now we add that 20 units to the 3 units we have already from the dividend we started with, so now we need to do 23/7. And so on ...I'd do an algebra example but it's a pain in the arse just using ASCII.
Try it with 1/7?
The same process still works if you want to calculate the decimals though. Just pretend you're doing (1000000/7) * (1/1000000) or however many digits you want.