359 karma · joined September 27, 2021
In statistics there are latin letters and greek letters. When you see a symbol denoted as a greek letter then that is a population parameter. When you see a latin letter that is a sample estimate. It could be Frequentist, Bayesian, Likelihoodist, Fiducial, Empirical Bayes, etc. Theoretical population greeks or sample calculated latins.
Statistics can be summarizes as one thing n -> N. Does ‘little n’ represent ‘big N’. In other words, does the sample generalize to the population. Statistics means something like “description of the state”. It was born out of census samples where larger population samples had to be estimated. “n” could be a handful of fish in a “N” lake. “n” could also be the parameter estimated in a linear regression with the sample of data collected while “N” is the true parameter of the relationship if we had all the data. Point estimation is about finding the needle in the haystack, but much more often statistics is about finding the haystack given the needle. One tool statistics uses to get to the haystack is probability.
- they often don’t support missing/corrupt data
Clean data would benefit most models, not just non-deep learning models. Missing data introduces bias even in DL models.
- they focus on linear relationships and not complex joint distributions
1. If seasonality is present. Which is usually the case in practical business problems then you will find that actual ~ lag_actuals explains most of the variance with a linear relationship. non-linearities in time series is not something that I see often. You can usually make a feature that explains those away linearly.
- they focus on fixed temporal dependence that must be diagnosed and specified a priori
Not sure sure you are saying here.
- they take as input univariate, not multiple interval, data
That is not the case. Time series regression can take many lags for inputs.
- they focus on one-step forecasts, not long time horizons
False. Time series regression models are used to forecast revenue many years into the future in the business world.
- they’re highly parameterized and rigid to assumptions
So is F=ma
- they fail for cold start problems
Because a cold start is not a time series data set. Why would time series methods work on non time series data.
sqlite3 :memory: -csv -header -cmd '.import taxi.csv taxi' 'SELECT passenger_count, COUNT(*), AVG(total_amount) FROM taxi GROUP BY passenger_count' | tv
Also, when I use SQLite I do not output using column mode. I pipe to `tv` (tidy-viewer) to get a pretty output.
https://github.com/alexhallam/tv
transparency: I am the dev of this utility
I added VisiData in my README and represented it in a positive light in the description. Again, just wanted to apologize for my mistake.
#better-together
I have a lot of respect your work. Let me know if I can make it up to you. I would be happy to point people to VisiData in my README as a recommendation of a tool that is built to explore and wrangle tabular data.
Also, thanks for the compliment! Like you, I like seeing more data tools in the terminal.