HNHacker News
TopNewBestAskShowJobs

samch93

160 karma · joined September 14, 2018

submissionscomments
samch93··on Bayesian Statistics: The three cultures
A (deep) NN is just a really complicated data model, the way one treats the estimation of its parameters and prediction of new data determines whether one is a Bayesian or a frequentist. The Bayesian assigns a distribution to the parameters and then conditions on the data to obtain a posterior distribution based on which a posterior predictive distribution is obtained for new data, while the frequentist treats parameters as fixed quantities and estimates them from the likelihood alone, e.g., with maximum likelihood (potentially using some hacks such as regularization, which themselves can be given a Bayesian interpretation).
samch93··on I don't use Bayes factors in my research (2019)
probably this article: Hoekstra, R., Morey, R.D., Rouder, J.N. et al. Robust misinterpretation of confidence intervals. Psychon Bull Rev 21, 1157–1164 (2014). https://doi.org/10.3758/s13423-013-0572-3

another good article on misinterpretation of p-values and confidence intervals is: Greenland, S., Senn, S.J., Rothman, K.J. et al. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. Eur J Epidemiol 31, 337–350 (2016). https://doi.org/10.1007/s10654-016-0149-3

samch93··on Science and Statistics (1976) [pdf]
I agree that there are many things Fisher got wrong (eugenics, tobacco lobbying, etc.), but his contributions to statistics (e.g., maximum likelihood or ANOVA) and genetics are among the most fundamental of the 20th century.
samch93··on Zelda: Link's Awakening game engine documentation (2021)
The Gaming Historian has a good video about the development of the game (https://youtu.be/pfvk6CJ3v34)
samch93··on Zelda: Link's Awakening game engine documentation (2021)
Link's Awakening is such a weird but fascinating entry in the Zelda series. I was not surprised to learn that its story was directly influenced by Twin Peaks.
samch93··on Show HN: Zelda Breath of The Wild Street View
Nintendo ninjas coming in 3, 2, 1,....
samch93··on Julia ranks in the top 5 most loved programming languages for 2022
try the following R code:

fisher.test(matrix(c(786, 298, 22999, 11156), nrow = 2))

#> Fisher's Exact Test for Count Data

#>

#>data: matrix(c(786, 298, 22999, 11156), nrow = 2)

#>p-value = 0.0003266

#>alternative hypothesis: true odds ratio is not equal to 1

#>95 percent confidence interval:

#> 1.116046 1.469612

#>sample estimates:

#>odds ratio

#> 1.279435

samch93··on Why I got a PhD at age 61
Do you know about Claude Shannon's Master thesis?
samch93··on Guetzli vs. MozJPEG
Guetzli means cookie in Swiss German, I wonder what inspired the devs to choose that name
samch93··on Sustainable Energy without the Hot Air (Revised, Community Edition)
It‘s such a tragedy that he died so young. Prof Mackay was a role model who impacted our world in many positive ways (e.g. his work in information theory, Bayesian statistics, environmental science). I recommend his lecture on information theory which is available on youtube [1]. He also documented the progress of his cancer very detailed on his blog [2], which is a very interesting (and sad) read.

[1] https://youtu.be/BCiZc0n6COY

[2] http://itila.blogspot.com/?m=1

samch93··on Interpreting A/B test results: false positives and statistical significance
The ASA recently published a new statement which is more optimistic about the use of p-values [1]. I myself also think that correctly used p-values are in many situations a good tool for making sense out of data. Of course, a decision should never be conducted on a p-value alone, but the same could also be said about confidence/credible intervals, Bayes factors, relative belief ratios, and any other inferential tool available (and I‘m saying this as someone who is doing research in Bayesian hypothesis testing methodology). Data analysts always need to use common sense and put the data at hand into broader context.

[1] https://projecteuclid.org/journals/annals-of-applied-statist...

samch93··on What Is a Nomogram?
A cool nomogram from statistics available here https://bmcmedresmethodol.biomedcentral.com/articles/10.1186... people often confuse the p-value as posterior probability of the null hypothesis. With this nomogram it is possible to compute a bound on the actual probability of the null hypothesis based on a specified prior and p-value.
samch93··on Why do we use R rather than Excel?
This is the typical „CS-people“ reaction to learning R. R is a language written by statisticians for statisticians which can be a good but also a bad thing. Fact is that R is the lingua franca of statistics, most new methods will be first available in R, rarely in python, almost never in julia. Despite the superior design of julia, there are good reasons to still use R, for example, there are many state of the art libraries such as ggplot2 or data.table which beat any alternative from python or julia.
samch93··on Croatian Wikipedia Disinformation Assessment
Native swiss here and would not sign that swiss german can‘t be understood by germans and that it is a separate language. Austrians and germans who live at the border can very well understand us and their dialects share many features with swiss german, more so than with some german dialects from the north. I think that categorizing swiss german as a separate language is very much an act of balkanization. It is part of the german language whichs spans a large geographic area with many different variations.
samch93··on Bonzi Buddy
Can recommend this video from the science elf about bonzi buddy (https://youtu.be/gU1hesbmaKM), great youtube channel!
samch93··on Pesticides are killing the world's soils
Both initiatives are indeed quite drastic, but so is the decline in biodiversity. Double yes from my side!
samch93··on Franz Kafka: Manuscripts, drawings and personal letters go online
One of my favorite writers. Highly recommend to read also his less famous works like "das Schloss", "Amerika", or "in der Strafkolonie". If you dig such books, give also Robert Walser and Bruno Schulz a try!
samch93··on Researchhub: GitHub for Science
Reminds me of the open science framework (https://osf.io/), which is widely used by researchers from social sciences / psychology.
samch93··on Bill Evans – The Creative Process and Self-Teaching (2007) [video]
Undercurrent is my favorite, closely followed by waltz for debby, stellar artist!
samch93··on Groundhog: Addressing the Threat That R Poses to Reproducible Research
Can recommend the paper "A Reproducible Data Analysis Workflow with R Markdown, Git, Make, and Docker" by Peikert and Brandmaier [1], which shows a much more robust approach to reproducibility.

[1] https://psyarxiv.com/8xzqy/

samch93··on Peer-reviewed papers are getting increasingly boring
Doug Altman actually wrote a paper about this issue already in 1994 [1]. A great quote: „We need less research, better research, and research done for the right reasons“

[1] https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2539276/pdf/bmj...

samch93··on The Elements of Statistical Learning [pdf]
Another hidden gem is Webb‘s Statistical Pattern Recognition https://www.wiley.com/en-us/Statistical+Pattern+Recognition%...
samch93··on Common color mistakes and how to avoid them
For all things related to color in data visualization, I can fully recommend the R-package colorspace. For example, the package features colorblind friendly palettes and also the possibility to emulate different forms of color vision deficieny. See the paper on https://arxiv.org/pdf/1903.06490.pdf and the webpage of the project https://colorspace.r-forge.r-project.org/index.html
samch93··on New Features in Gnuplot 5.4
Amazing project, but seriously why do people spend their time and energy to implement 3d bar charts?
samch93··on Ask HN: Was the Y2K crisis real?
I really recommend this video from LGR about the Y2K crisis https://youtu.be/Xm5OiB3CPxg
samch93··on Why scientists need to be better at data visualization
I can totally recommend knitr by the amazing yihui xie (https://yihui.org/knitr/ ) for this use case. It allows you to write R code chunks directly in your latex code and the output of the code (tables, plots) is then directly inserted in your pdf after compiling. Together with git and docker this gives a fully reproducible workflow!
samch93··on Show HN: Chart.xkcd – Xkcd-styled chart library
Also an R package exists to do this on top of ggplot2: http://xkcd.r-forge.r-project.org/
samch93··on Debunking the Stanford Prison Experiment [pdf]
In my opinion, if such a culture- and time-dependent phenomenon is interesting, it is still worth studying. If this dependence is obvious, researchers should simply not claim that the phenomenon is generalizable.

The replication crisis has provoked many discussions on similar issues among psychologists, i.e. some dinosaurs question the value of direct replication studies and prefer conceptual replications (see for example this paper by Simons https://www3.nd.edu/~ghaeffel/Value.pdf). I believe that direct replication is the only way to verify the credibility of scientific discoveries, and certainly most new generation psychological researchers agree on this.

samch93··on What’s the difference between statistics and machine learning?
I agree that in general statistics is more concerned about inference, while machine learning focuses on prediction. On the other hand, there is such a huge overlap between the fields that it is hard to make a distinction. Also there are statistical fields which focus more on prediction the same way as ML does. For example, in geostatistics prediction is often the only goal, e.g. to predict heavy metal concentration across a domain. People accept that it is impossible to explain every bit of spatial variation and just model it by a gaussian process (the same gaussian process used in ML).
samch93··on Visual Information Theory (2015)
Nice article. For those who are more interested in mosaic plots, statisticians have already done a lot of work on this issue. For R there are many nice solutions, e.g. the strucplot framework which allows to visualize complicated relationships between multiple qualitative variables (https://www.jstatsoft.org/article/view/v017i03).
Page 1 of 2Next →