HNHacker News
TopNewBestAskShowJobs

flusteredBias

359 karma · joined September 27, 2021

submissionscomments
flusteredBias··on Positron – A next-generation data science IDE
I use it and like it.
flusteredBias··on Apple typewriter memo (2020)
It’s 2025 and I bought 3 typewriters just this year. I’m fired.
flusteredBias··on Kermit: A typeface for kids
This is anecdotal and I hope someone who has some research experience can say whether this is true or not generally, but I recently got a Kindle and found that if I use really large font sizes where there are fewer than 50 words on a page it's easier for me to stay engaged. Maybe this has something to do with cognitive load or chunking information. Some fonts look quite a bit better at these large sizes. So for me I don't think typography alone is sufficient. I think the interaction between a large font size and a typography that looks pleasing at a large font size helps with engagement.
flusteredBias··on Pipe Syntax in SQL
... so dplyr.
flusteredBias··on WriteFreely: An open source platform for building a writing space on the web
How does this compare to bearblog
flusteredBias··on I've stopped using box plots (2021)
ECDF plots are what I use.
flusteredBias··on Csvlens: Command line CSV file viewer. Like less but made for CSV
See also: (tv) Tidy Viewer. A cross-platform CLI csv pretty printer.

https://github.com/alexhallam/tv

flusteredBias··on Joining CSV Data Without SQL: An IP Geolocation Use Case
I do the same with DuckDB and pretty print with tidy-viewer.
flusteredBias··on IPyflow: Reactive Python Notebooks in Jupyter(Lab)
The only way I do that is with git.
flusteredBias··on IPyflow: Reactive Python Notebooks in Jupyter(Lab)
Given that I have no clue what a multi-user story is I am guessing not.
flusteredBias··on IPyflow: Reactive Python Notebooks in Jupyter(Lab)
I kind of think quarto is a much better solution to the problems that notebooks try to solve plus you get the added bonus of having plain text as the file source.
flusteredBias··on Statistical vs. Deep Learning forecasting methods
Don't take my word for it. https://ocw.smithw.org/csunstatreview/statisticalsymbols.pdf
flusteredBias··on Statistical vs. Deep Learning forecasting methods
How the parameters are estimated it not the message.

In statistics there are latin letters and greek letters. When you see a symbol denoted as a greek letter then that is a population parameter. When you see a latin letter that is a sample estimate. It could be Frequentist, Bayesian, Likelihoodist, Fiducial, Empirical Bayes, etc. Theoretical population greeks or sample calculated latins.

flusteredBias··on Statistical vs. Deep Learning forecasting methods
I apologize. ^This comment was to harsh.

Statistics can be summarizes as one thing n -> N. Does ‘little n’ represent ‘big N’. In other words, does the sample generalize to the population. Statistics means something like “description of the state”. It was born out of census samples where larger population samples had to be estimated. “n” could be a handful of fish in a “N” lake. “n” could also be the parameter estimated in a linear regression with the sample of data collected while “N” is the true parameter of the relationship if we had all the data. Point estimation is about finding the needle in the haystack, but much more often statistics is about finding the haystack given the needle. One tool statistics uses to get to the haystack is probability.

flusteredBias··on Statistical vs. Deep Learning forecasting methods
I am going to defined the readme a little

- they often don’t support missing/corrupt data

Clean data would benefit most models, not just non-deep learning models. Missing data introduces bias even in DL models.

- they focus on linear relationships and not complex joint distributions

1. If seasonality is present. Which is usually the case in practical business problems then you will find that actual ~ lag_actuals explains most of the variance with a linear relationship. non-linearities in time series is not something that I see often. You can usually make a feature that explains those away linearly.

- they focus on fixed temporal dependence that must be diagnosed and specified a priori

Not sure sure you are saying here.

- they take as input univariate, not multiple interval, data

That is not the case. Time series regression can take many lags for inputs.

- they focus on one-step forecasts, not long time horizons

False. Time series regression models are used to forecast revenue many years into the future in the business world.

- they’re highly parameterized and rigid to assumptions

So is F=ma

- they fail for cold start problems

Because a cold start is not a time series data set. Why would time series methods work on non time series data.

flusteredBias··on Statistical vs. Deep Learning forecasting methods
I use my package https://github.com/alexhallam/tablespoon to generate naive forecasts then evaluate the crps of the naive vs the crps of the alternative method. This “skill score” approach is very good.
flusteredBias··on Statistical vs. Deep Learning forecasting methods
You mean CRPS?
flusteredBias··on Statistical vs. Deep Learning forecasting methods
I don’t think I have read anything more false on the internet. XD
flusteredBias··on One-liner for running queries against CSV files with SQLite
https://github.com/alexhallam/tv#inspiration
flusteredBias··on One-liner for running queries against CSV files with SQLite
Here is an example of how I would pipe with headers to `tv`.

sqlite3 :memory: -csv -header -cmd '.import taxi.csv taxi' 'SELECT passenger_count, COUNT(*), AVG(total_amount) FROM taxi GROUP BY passenger_count' | tv

flusteredBias··on One-liner for running queries against CSV files with SQLite
I am a data scientists. I have used a lot of tools/libraries to interact with data. SQLite is my favorite. It is hard to beat the syntax/grammar.

Also, when I use SQLite I do not output using column mode. I pipe to `tv` (tidy-viewer) to get a pretty output.

https://github.com/alexhallam/tv

transparency: I am the dev of this utility

flusteredBias··on Show HN: Tidy Viewer – a cross-platform CSV pretty printer for viewer enjoyment
Thanks and great work!
flusteredBias··on Show HN: Tidy Viewer – a cross-platform CSV pretty printer for viewer enjoyment
You don't have to make the alias `tv`. Feel free to make your alias tidy='tidy-viewer'.
flusteredBias··on Show HN: Tidy Viewer – a cross-platform CSV pretty printer for viewer enjoyment
https://github.com/alexhallam/tv/pull/58

I added VisiData in my README and represented it in a positive light in the description. Again, just wanted to apologize for my mistake.

#better-together

flusteredBias··on Show HN: Tidy Viewer – a cross-platform CSV pretty printer for viewer enjoyment
Love the idea! The issue is already open. I will merge a PR before the next release. https://github.com/alexhallam/tv/issues/49
flusteredBias··on Show HN: Tidy Viewer – a cross-platform CSV pretty printer for viewer enjoyment
Oo, I am sorry. I see I misrepresented VisiData. I apologize. Thank you for the corrections.

I have a lot of respect your work. Let me know if I can make it up to you. I would be happy to point people to VisiData in my README as a recommendation of a tool that is built to explore and wrangle tabular data.

Also, thanks for the compliment! Like you, I like seeing more data tools in the terminal.

flusteredBias··on Show HN: Tidy Viewer – a cross-platform CSV pretty printer for viewer enjoyment
Thanks!
flusteredBias··on Show HN: Tidy Viewer – a cross-platform CSV pretty printer for viewer enjoyment
I see. Thanks for the clarification.
flusteredBias··on Show HN: Tidy Viewer – a cross-platform CSV pretty printer for viewer enjoyment
Oo I see. Thanks for clarifying.
flusteredBias··on Show HN: Tidy Viewer – a cross-platform CSV pretty printer for viewer enjoyment
I have been using tv now for a couple months at work. It has been working well on the data I see. If you find edge cases then please open an issue with an example csv.
Page 1 of 2Next →