HNHacker News
TopNewBestAskShowJobs

cscheid

1,841 karma · joined October 13, 2010

building https://quarto.org at Posit (fka RStudio).

https://cscheid.net https://github.com/cscheid https://bsky.net/profile/cscheid.net

submissionscomments
cscheid··on Our brains 'time-stamp' sounds to process the words we hear
Antonio Damasio is a neuroscientist who's done work in the area, and has decent popsci books about this idea. I read the older ones, "Descartes' Error" and "The Feeling of What Happens", and they're fun, good reads. Apparently he's written more on the subject as well.
cscheid··on Grabby Aliens: A Resolution to the Fermi Paradox
> As such, I hold the entire theory as suspect and someone trying to justify a particular mode of human existence. This extends deeper into the trilogy as women continually are painted as too soft and emotionally empathetic to make the cold hard decisions that will kill many humans but ensure the survival of the race.

For folks who've only read the first book and might be somewhat confused, this specific set of values becomes crystal clear about halfway through the second book.

cscheid··on Not the Time to Get Greedy: House Flippers Getting Burned by US Housing Downturn
And also, don't look to get into furniture making where the (previously affordable) good looking, high-quality, easy-to-work-with material is 97% imported from... Russia. Baltic birch plywood, I barely knew ye.
cscheid··on Another scientific body has debunked bitemark analysis
From my understanding, fingerprint analyses don't have widely-accepted specificity or sensitivity rates either. Those "73% match" figures are not derived from any sound principles.
cscheid··on Art Garfunkel's Library
> do you ever shake the feeling that having a book read to you is somehow qualitatively not as good as reading it from the page?

I can't answer for OP, but I did spend a year commuting a bunch and got through ~60 books or so that way.

I went from having time for 10-15 books a year to having time for ~60 books a year, so yes, you do miss a little. But not ~50 books worth of missing. It's easily worth it.

cscheid··on Scientific grant applications are getting heavier on hype
Interesting that you bring up WWII. There's a theory that says that US's big research funding agencies are deliberately set up the way that they are because the US government freaked the _fuck_ out upon realizing that a bunch of nerds could go from zero to nuclear bombs in ~10 years. So, these big bureaucracies were set up to, in effect, keep tabs on physicists. (I first heard of this from a talk from Kim Stanley Robinson of Red Mars fame, and more recently Ministry for the Future. It's on YouTube somewhere but it's a bit hard to find, it was a seminar he gave at Duke University some 15 years ago.)

Then, some 35 years later came the Bayh-Dole act, which however well-meaning it might have been, really provided the incentives for universities to turn into fed money capturing enterprises. The rest is history.

cscheid··on Scientific grant applications are getting heavier on hype
(former tenured prof here)

This is exactly right. Just like it's better to think of McDonald's as a real estate company with a food business on the side, these days it's better to think of big state schools as mechanisms for ingesting federal research dollars, with an education business on the side. (Never mind that state schools shouldn't be education businesses _at all_, they should be _public services_)

cscheid··on DuckDuckGo email protection beta now open
(I've been a happy paying customer of Fastmail for _years_ and I didn't know about masked email. Thank you!)
cscheid··on Perfluorocubane is (as you would expect) weird
Lowe is being funny and calling back to his classic "Things I won't work with" series (which is also linked to in a different comment).

One of the things you learn just from that series is that anything with this much fluorine jammed into it is just asking for trouble. Case in point, https://www.science.org/content/blog-post/things-i-won-t-wor...

cscheid··on ‎Cracking the Code: Sneakers at 30
1) Then Donald Logue got famous, and I'd be the idiot yelling "that's Gunter Janek!" at every episode of Grounded for Life.

2) I can't look at that screenshot and not hear "I leave message here on service but you do not call"

3) RSA's Adleman was the science advisor for the movie, so I'd guess he snuck in the Asiacrypt poster from the beginning

cscheid··on Datalog in JavaScript
I'm sorry, but Datalog has existed since 1977; RDF, since 1999.
cscheid··on Error 404 (Not Found)
Working now, yes. Though deno.land's current IP still nslookup's to something inside googleusercontent.com so maybe google has fixed their Side. But it was definitely offline earlier today (cf. our github CI failures and downforeveryoneorjustme)
cscheid··on Error 404 (Not Found)
https://deno.land down too. I guess a lot of CI actions will be failing just like ours...
cscheid··on Mazette
If you enjoyed that, you'll probably also enjoy the "Maze Generation" section of Mike Bostock's Visualizing Algorithms talk: https://bost.ocks.org/mike/algorithms/#maze-generation
cscheid··on Every Model Learned by Gradient Descent Is Approximately a Kernel Machine
No, they require a full pass over the support vectors, which are potentially a much smaller set. (That’s part of why everyone was so excited about SVMs when they were invented) The support vectors are the training values with nonzero hinge loss, or alternatively, training values sufficiently close to the decision boundary.
cscheid··on The Remarkable Number 1/89 (2004)
I believe that gives you the result starting from Fibonacci, but not that any starting pair of numbers eventually lands at the golden ratio.
cscheid··on The Remarkable Number 1/89 (2004)
If you read this and you’re curious, there’s a proof that is fairly easy to follow if you understand eigenvectors. Write the operation that takes the two last elements of the sequence and produces the following two, notice it’s linear, then analyze the eigenvalues of the associated matrix and relate the original operation to the power method.
cscheid··on Pianojacq, an easy way to learn to play the piano
I’m super curious, if you don’t mind my prodding. I looked into doing this for my own version of a JavaScript midi teacher a while back. Any chance you’re doing it via the neural nets sequence to sequence translation algorithms that are super popular in NLP right now? I think that’s a practical way to do it, except for needing a training corpus. I was going to reverse engineer synthesia’s data to try it out, for what’s worth...
cscheid··on Measure code execution time accurately in Python
I'm not a stats person! But I have been burned in real projects by doing the wrong thing. I just have learned --- the hard way --- to honestly appreciate it.
cscheid··on Measure code execution time accurately in Python
You make a good point I hadn't considered. In that case, not even robust least squares will save you, because the samples are not independent of one another. You'd have to start looking into explicit time-dependent methods and yikes, I have no idea of what the literature says about time dependence and outliers. I wouldn't be surprised if it's "here be dragons" territory.
cscheid··on Measure code execution time accurately in Python
As far as i understand it the RANSAC arguments are very much not Bayesian (they're all about "if you repeat this an infinite number of times under an infinite number of new samples, then...", which is almost caricaturely frequentist).

Still, you just gave me reason to mention one of my favorite "no, that won't work either" paper :) on how Bayes will not save you in the presence of model misspecification. Instead of butchering it any further, I'll just point you to this piece explaining the work, written by the author of the paper himself: http://bactra.org/weblog/601.html

cscheid··on Measure code execution time accurately in Python
Thank you for pointing out some bad phrasing on my part.

When I said "small dataset", I should have said "small subset of all collected points that are initially fed to the model". The issue isn't that you have to collect the data little by little. The issue is that, once you've given a linear model an input that is outside the distribution which you're hoping to model, nothing about the model can be trusted.

So you collect all the data points, but you don't give them all to the model at once. (I'm describing the RANSAC method here now) You start with a large number of "candidate models" that are all fit with a small number of input points, and then test which of the candidate models predict well the points you have not yet given the model. Then you feed the best of these candidate models only the points which it predicts well, and create a more refined, still outlier-free model. This can be proven to work in the presence of a small number of out-of-distribution points.

cscheid··on Measure code execution time accurately in Python
I don't know how else to put it, but the technique suggested in this paper is irreparably incorrect in the setting that it recommends. Attempting to fix a linear least squares fit by removing points with high residuals does not work. Please don't do that: they won't necessarily be the outliers in the dataset and your model will converge to the wrong thing.

Statistics is hard. Like, really really hard. Stuff goes wrong all the time. Please leave it to the experts.

If you're going to do this, please look up the methods behind (for example) robust least squares, outlier detection, L_1 regression, etc. The right way to do this is to start with a small dataset that is with very high probability free of outliers, and slowly grow it by never adding points which have large residuals. (If you've done 3D scan registration and image alignment, this is what RANSAC does.)

The principle is, intuitively, that once an out-of-distribution point gets into your linear model, the model is poisoned forever. You can't trust the model to tell you that the bad points are the outliers. The way this paper does is is irreparably broken, sorry :/

cscheid··on People see it as more acceptable to make passionate employees do extra: study
When I was in a less-than-ideal job situation a few years ago, I found myself saying "... but I love my job" to a friend. I'll never forget the reply: "... you may love your job, but the job doesn't love you back". It was tough to hear it, but in hindsight I'm very thankful for that jolt.
cscheid··on Build a Neural Network
https://cscheid.net/courses/spr19/csc665/

The assignments are not directly available, but my email is easy to find.

cscheid··on Build a Neural Network
Hm. I just finished teaching an ML course where all of the assignments were pure Python (on purpose, so students would actually have the chance to see all of the code). One of the assignments included implementing reverse-mode autodiff and a NN classifier on top. It can be done in ~600 lines of clear python, serious!
cscheid··on Times Newer Roman, a sneaky font designed to make essays look longer
Hey Andrew, funny seeing you here (we went to grad school together!)
cscheid··on Times Newer Roman, a sneaky font designed to make essays look longer
The amount of class-file hackery that goes on with LaTeX submissions is remarkable. (shoutout to the rest of HN procrastinating near the CHI deadline)
cscheid··on Gut Microbes Combine to Cause Colon Cancer, Study Suggests
https://www.nobelprize.org/nobel_prizes/medicine/laureates/2...
cscheid··on Show HN: Observable Notebooks
Fantastic.

One final question/request (for which I'd happily pay for!) : it would be awesome to have a path away from Observable's infra if desired. Say I _really_ want to host a particular notebook locally: is something like that planned? I know this is not a trivial feature since notebooks can call other notebooks, but I'd love to develop stuff on Observable knowing that should the worst happen and it doesn't exist anymore, I can run it all locally on my webpage.

← PreviousPage 5 of 13Next →