5,369 karma · joined March 23, 2010
https://terrytao.wordpress.com/2010/10/10/the-cosmic-distanc...
What I like about SEM and graphical model formulation is that it makes the model explicit and easy to communicate. It compactly encodes many hypotheses that you can test on your data.
It means that one has deployed it/is using the tool with a non-trivial (could be anywhere from few 100s to few 1000s of machines) amount of CPU and/or data. In the context of serving, "large scale" could also mean the number of queries/second hitting the serving layer.
1. One of the things Prop 13 does is to cap the maximum rate at which a property tax grows per annum to 2%. However, due to housing demand, the property values grow at a much faster rate (7% is not unheard of). The property tax can change when there is a sale.
2. Bay area prices have skyrocketed.
3. The peninsula topography and zoning laws make it hard to expand housing capacity.
So, if you are an aging home owner, sure you can cash out of a sale. But where would you go? Unless you can afford another place in the same area, existing home owners do not have incentives to sell their property, which results in a very low inventory. And, low inventory further puts pressure on the prices and drives it up.
Corollary: No probabilistic regular grammar
exhibits criticality.
In the next section, we will show that this statement is
not true for context-free grammars (CFGs).
That is, there exists CFGs that exhibit criticality. Programming languages are often parsed by CFGs, so it's likely that some programming languages exhibit the same criticality structure as natural languages.> Statistical dependence can be determined easily if you know the distributions,
Even if you know the distribution, a statistical test will make Type-I/II errors that you would have to take care of.
Actually, I find the text linked above hard to understand, without properly defining 'c.' What's the sample space?
In general, my sentiments are with the xkcd comic strip, but nothing more. Pearl's theories lay a firm foundation for communicating a causal hypothesis and manipulating it algebraically, but the true tests of causal hypothesis are:
- Experimental evidence
- The predictions it makes, in cases where experiments are hard to perform (e.g., in physics, when we make certain causal conjectures about how the universe works).
The notion of "Statistical dependence" is not nebulous. X and Y are independent if the joint distribution factorises as
p(X, Y) = p(X) p(Y)
> even they do not admit Pearson (or any other type) of correlation,Precisely. They operate purely in probabilistic dependence/independence terminology, from a theoretical point of view.
- How do you develop such web apps with html, css, javascript, go, etc. all interacting with each other?
- How are static assets packaged in a single binary?
- Any simple tutorial or stack walkthrough you would recommend me reading?
thanks!
Depending on the study, finding a delta whose p-value is significant does not necessarily mean that the size of the effect (i.e., delta) might be significant enough to be useful.
It would be good to gracefully handle large graphs. I typed in "torvalds/linux" and my tab froze!
Have you folks considered Supersonic engine from Google, which was designed with similar (but not as extensive as Arrow) goals in mind?
Pearl is also an enthusiastic speaker. You can search for his talks online at various venues (Stanford, Microsoft Research, etc.) to learn more.
[1] http://www.michaelnielsen.org/ddi/if-correlation-doesnt-impl...
I've read the topology and data paper; it lays down motivations for TDA, but it doesn't quite connect it to existing literature on dimensionality reduction and manifold learning and explain -- "Here's something you can learn by using tools from TDA, but not existing methods." The best I could see was that it produces results similar to existing methods.
On first glance, the methods here seem a lot like the toolbox of dimensionality reduction techniques (PCA, spectral embedding, or more general manifold learning, etc.) from machine learning literature.
What specific insights from the field of topology have helped further our understanding of data that we missed earlier?