369 karma · joined January 4, 2015
But, I'm pretty sure the assumptions of logistic regression are even stronger than just that. The inputs are assumed to be independent given the output class, and the log odds of the output vary as a linear function of each input. The first one is essentially the naive Bayes assumption, and the second one is completely unreasonable for almost any problem ever (roughly equivalent to assuming every dataset has a multivariate normal distribution). If they are both correct, though, you get a perfectly good Bayesian posterior probability of each output class.
I think the lesson is that gradient descent will build a decent function approximation out of pretty much anything powerful enough, which is why neural networks still work even when probability theory has been thrown completely out the window.
Neural network outputs are not probabilities. I think that's the main lesson here.
I think the correct way to state it is: "all true statements are either tautologies, or can be determined empirically (with arbitrarily high but not necessarily measure-1 probability)". This statement is itself a tautology, because tautologies are by definition true, and because "determined" implies some method of determination, which if it actually can be used to determine truth, means it can be used empricially. Determination and empiricism are secretly defined in terms of each other, basically. The reason this tautology is worth stating is that it gives a simple criterion for discarding non-questions: questions that have no method of determining whether they are true or false (with arbitrarily high probability), are always non-questions.
Just like 42, "atoms and the void" here is just a science-flavored attempt to answer a non-question.
Yes, but this doesn't even come close to describing the typical users' password, which is most likely a 6-letter English word with a capital letter and a 1! appended to the end. Your calculation here isn't really relevant, because it's all about the worst or common case. (You also assume that people are using a GPU for a compute-bound problem, when much faster FPGAs are also available, but either way it's moot.)
Security through obscurity, which is what you're proposing with the shuffled salt idea, is also not normally considered the right way to go. If you wanted to use a similar but much simpler and straightforward method, you could just encrypt the salted hashes before storing them in the database.
The one case where they coincide (sort of) is if you believe your random sequence is generated by a randomly chosen Turing machine, which I've only really seen in philosophical settings.
A uniformly chosen 64-bit integer still has exactly 64 bits of entropy, regardless of how much Kolmogorov complexity the actual bits you generate have.
A deterministic PRNG's sequence has exactly the entropy of it's seed, actually, but it has 0 bits of entropy per symbol, because its sequence is infinite.
The thing most people get confused about with entropy is in thinking that entropy is a property of some single object, like a bit string. Really, entropy is always a measurement about a probability distribution, just like mean or variance is. In the usual case with random streams, the distribution is P(x_i | x_i-1 ... x_0) for bits x_i in the stream, i.e. the distribution remaining for the current bit even if we know all previous bits. For a deterministic PRNG, once we can extract the key from the history (given unlimited compute power) that distribution becomes deterministic, so the entropy is 0.
This has always been possible, but it sounds like they've lowered the minimum entropy needed in the source streams to produce a high-quality output.
So, now it has jumped from (disk -> network -> memory -> ...) to (network -> disk -> memory -> ...), which is a big change.
But hopefully, you've got a STEM degree, and can figure that fact out out yourself ;)
http://slatestarcodex.com/2014/11/03/all-in-all-another-bric...
I'm only a bit disappointed that the author seems not to realize that Bayes' theorem is just a simple consequence of probability theory, and should be attractive not because "maybe the brain is Bayesian", but because it is based on sound set-theoretic and analytic principles. If Bayes' theorem is false, so is probability theory, and so is nearly everything we know about probability.
Edit: Here is a good explanation of the theorem that makes it visually clear how only set theory is involved in deriving it: https://oscarbonilla.com/2009/05/visualizing-bayes-theorem/
One of the simplest is spike-timing-dependent plasticity [1], which is caused by the behavior of NMDA receptors shortly before and after depolarization. (This is why ethanol, an NMDA antagonist, can produce a "blackout" in high doses, where no memories are formed.)
In general, neurons have a lot of mutable long-term state. The graph of how neurons are connected can change, the strength of those connections can change, the internal chemistry of the neurons can change through gene expression factors, and, apparently, levels of long-lived prions can change.
This is not to say that the whole mechanism of long-term memory is understood, but that this discovery is just one of a class of mechanisms that may all be working in parallel or even independently.
[1] https://en.wikipedia.org/wiki/Spike-timing-dependent_plastic...
#define ASSUME(x) if(!(x))__builtin_unreachable()
This is compiler-specific, of course, and introduces undefined behavior if the assumption is violated. There's also no guarantee that the compiler will use the information well. But, it's almost guaranteed to not generate any extra code based on it.Nothing, of course. Thinking that believing reductionism will somehow transform human minds into mathematics is like thinking that believing evolution will transform humans into chimps. And whether or not you believe in reductionism, the reality of humanity won't change. The only thing that changes is that the non-reductionist philosophers lose and the neurobiologists, cognitive scientists, and AI researchers win. Which seems to be happening, in any case.
If you're feeding effectively random data into the block cipher (like if you're using CBC), then because of the birthday paradox, you get at most about 2^32 blocks (far fewer in practice at a good security level) per key if you have 64-bit blocks. This is low enough to be annoying for designers or problematic for suites that don't rekey correctly.
However, because CTR (or GCM) mode uses sequential inputs to the cipher, I think that a 64-bit block size would not be a problem there. At that point, the reason not to use 64-bit block ciphers is because they're all older, weaker, and less-supported than AES-128.
The backwards-moving pattern of "backpropagation" is really just a side-effect of the derivative chain rule application order, but a lot of intro materials treat backprop as if it is some fancy thing specially-designed for neural nets. I suppose "compute the gradient of this function using basic vector calculus" just isn't sexy enough. I complain mostly because it took me a while to figure out whether backprop was exactly the same as gradient descent, or if there were subtle differences.
Whether we're talking about radiation caused by wireless communication (e.g. wifi at 2.4/5 GHz) or high-frequency oscillations in a microprocessor, it's pretty safe to say that all of the significant electromagnetic radiation coming off of a wearable computer is under 10 GHz.
Damage to proteins, DNA, etc. due to radiation is either caused by that radiation stripping electrons / breaking covalent bonds, or through heating.
The sort of electromagnetic radiation that strips electrons and breaks covalent bonds is called ionizing radiation; ionizing radiation only occurs above a certain frequency threshold (depending on the material being ionized). This fact is, in fact, the reason Einstein got his Nobel in physics. Anyway, its pretty safe to say that, say, red visible light (400 THz) does not ionize important human molecules. Ultraviolet is usually considered to be the low end of the ionizing radiation range.
Therefore, because 400 THz > 10 GHz, radiation coming from wearable computers could not possibly cause molecular damage to humans through ionization. The light coming from the screen is significantly more dangerous in this respect than anything coming from the other electronics.
How about heating? Consider that a typical wearable computing device only consumes a few watts. If this power were distributed diffusely, it is harmless, and if it were focused, it would cause obvious and painful burns, which we know doesn't happen.