HNHacker News
TopNewBestAskShowJobs

neuromantik8086

560 karma · joined October 27, 2016

submissionscomments
neuromantik8086··on Columbia and NYU Would Lose $327M in Tax Breaks Under Proposal
Because endowments aren't piggy banks. They're regulated by UPMIFA [1], which states that universiteis can't draw down more than 7% of the total funds in the endowment unless they can prove that it would be prudent to do so, and the burden of proof is extremely high.

Even without UPMIFA, endowments are a mix of unrestricted and restricted funds, and donor restrictions can and do prevent universities from using money when they might otherwise want to. Even if a university desired to draw down the full 7% allowed without triggering red tape, it's unlikely that they would be able to draw it all without running afoul of donor intent.[2]

If anything, the system is to blame here, not the universities themselves necessarily (not to excuse bad apples in academic administration).

[1] https://en.wikipedia.org/wiki/Uniform_Prudent_Management_of_...

[2] https://en.wikipedia.org/wiki/Donor_intent

neuromantik8086··on Challenge to scientists: does your ten-year-old code still run?
I never answered your last question so here goes:

> Are you saying that the reproductive challenge poses a difficulty to Common Workflow Language?

I don't actually understand how the reproducibility challenge undermines the validity of using CWL / flow-based programming as an approach to promoting reproducible analyses. There certainly wasn't anything in the article that made me think that CWL was challenged, but Hinsen explicitly called out CWL in the abstract, which implies that for some reason he thinks, a priori, that it's a non-solution. He never justifies this implied assumption further, and as near as I can tell, none of the attempted replications used a flow-based language.

If Hinsen really aimed to argue against the viability of CWL/flow-based programming as an approach to reproducibility, he would have done a systematic comparison of historical analyses that used a flow-based system (like National Instruments' Labview or Prograph) vs analyses that are more similar to the approach that he seems to favor (i.e., analyses using Mathematica or Maple).

While I find the challenge interesting to follow, and the retrocomputing geek in me finds it fun, I don't actually understand what it really accomplished other than being a fun diversion. Assuming that an analysis was written in a Turing-complete language and you didn't use non-deterministic algorithms, you should theoretically be able to reproduce the results exactly on modern hardware, and using non-deterministic algorithms I would imagine that a result would be "close enough" within some kind of confidence interval. You may need to go to great lengths (in terms of emulating instruction sets, ripping tapes, etc), but I think a visit to any retrocomputing festival or computer history museum would have made that pretty obvious from the outset.

neuromantik8086··on Challenge to scientists: does your ten-year-old code still run?
I can only speak to Spack in particular, but the main issue that I found with it was balancing researcher expectations for package installation speed with compile times. For most packages, compile times aren't a huge problem, but compilers themselves can take days to build, and it isn't unheard of for researchers to want a recent version of gcc for some of their environments.

In theory this isn't an issue with Spack (assuming that you have a largely homogeneous set of hardware or don't use CPU family-specific instruction sets), since you can set up cached, pre-compiled binaries on a mirror server (similar to a yum repo) and have people install from there.

Spack, however, has a lot of power/complexity. A lot of untamed power that means that bugs can sometimes be more likely than in other, more mature (or mature-ish) package managers. Namely, Spack allows you to not only specify the version number of a package, but also the compiler that you use to make that package, specific versions of dependencies that you want to use, which implementation of an API you want to use (i.e., MPICH or OpenMPI for MPI), and compiler flags for that package. When you run an install command / specify what you want to install, Spack then performs dependency resolution and "concretizes" a DAG that fulfills all of the constraints.

The issue that I ran into was that if you don't specify everything, Spack makes decisions for you about which version of a dependency, which compiler, etc to use (i.e., it fills in free variables in a space with a lot of dimensions). This would be great and dandy normally, although the version of Spack that I used occasionally constructed totally different graphs for the same "spack install gcc" command (if I recall correctly; take all of this with a grain of salt b/c I might be misremembering). This meant that it wouldn't use cached versions of gcc that had already been built, and ended up rebuilding minor variants of gcc with options I didn't care about.

At National Labs and larger outfits, the trade-offs between this kind of complexity/power and the accompanying bugginess (Spack has yet to hit 1.0) seem to favor complexity/power while accepting these sorts of bugs, but I don't work at a larger outfit and my group didn't need that level of power/control over dependencies and rather needed something that "just worked" and would allow researchers to be able to install packages independently of us (IT people). conda (mostly) fit the bill for this. I still think that Spack is the future and it has a special place in my heart, but it will have to be more stable for me to want to use it in production.

neuromantik8086··on Challenge to scientists: does your ten-year-old code still run?
I never said that Konrad Hinsen's agenda was hidden; in fact, it's not at all hidden (which is why I linked the abstract). It's just that this context isn't at all clear in the Nature write-up, and it's relevant to take into account.

I haven't taken the time to seriously contemplate the merits of CWL vs Leibniz, although my gut instinct is that we don't really need another domain-specific language for science given the profusion of such languages that already exist (Mathematica, Maple, R, MATLAB, etc). That's the extent of my bias, but again, it's a gut instinct and not a comprehensive well-reasoned argument against Leibniz.

neuromantik8086··on Challenge to scientists: does your ten-year-old code still run?
There are some efforts in this vein within academia, but they are very weak in the United States. The U.S. Research Software Engineer Association (https://us-rse.org/) represents one such attempt at increasing awareness about the need for dedicated software engineers in scientific research and advocates for a formal recognition that software engineers are essential to the scientific process.

In terms of tangible results, Princeton at least has created a dedicated team of software engineers as part of their research computing unit (https://researchcomputing.princeton.edu/software-engineering).

Realistically though even if the necessity of research software engineering were acknowledged at the institutional level at the bulk of universities, there would still be the problem of universities paying way below market rate for software engineering talent...

To some degree, universities alone cannot effect the change needed to establish a professional class of software engineers that collaborate with researchers. Funding agencies such as the NIH and NSF are also responsible, and need to lead in this regard.

neuromantik8086··on Challenge to scientists: does your ten-year-old code still run?
Just as a quick bit of context here, Konrad Hinsen has a specific agenda that he is trying to push with this challenge. It's not clear from this summary article, but if you look at the original abstract soliciting entries for the challenge (https://www.nature.com/articles/d41586-019-03296-8), it's a bit clearer that Hinsen is using this to challenge the technical merits of Common Workflow Language (https://www.commonwl.org/; currently used in bioinformatics by the Broad Institute via the Cromwell workflow manager).

Hinsen has created his own DSL, Leibniz (https://github.com/khinsen/leibniz ; http://dirac.cnrs-orleans.fr/~hinsen/leibniz-20161124.pdf), which he believes is a better alternative to Common Workflow Language. This reproducibility challenge is in support of this agenda in particular, which is worth keeping in mind; it is not an unbiased thought experiment.

neuromantik8086··on Challenge to scientists: does your ten-year-old code still run?
Guix is one of several solutions that has been touted as a solution. Another one that is quite popular in HPC circles is Spack (https://spack.readthedocs.io/en/latest/).

At my institute, we actually tried out Spack for a little bit, but consistently felt like it was implemented more as a research project rather than something that was production-level and maintainable. In large part, this was due to the dependency resolver, which attempts to tackle some very interesting CS problems I gather (although this is a bit above me at the moment; these problems are discussed in detail at https://extremecomputingtraining.anl.gov//files/2018/08/ATPE...), but which produces radically different dependency graphs when invoked with the same command across different versions of Spack.

I've since come to regard Spack as the kind of package manager that science deserves, with conda being the more pragmatic / maintainable package manager that we get instead . Spack/Guix/nix are the best solution in theory, but they come with a host of other problems that made them less desirable.

neuromantik8086··on The Shadow Inc. app that failed in Iowa last night
The COO majored in Music Technology at Oberlin. That's quite a bit more technical than most people realize. TIMARA (the music tech program at the Oberlin Conservatory) involves a decent amount of programming and/or audio engineering. To put that in perspective, the founder of Macromind/Macromedia (Marc Canter) is also an alumnus of TIMARA.
neuromantik8086··on WeWork chases new financing as cash crunch looms
I prefer the following:

"It felt like a yuppie aquarium."

https://www.reddit.com/r/finance/comments/c93agd/wework_isnt...

neuromantik8086··on Why Enterprise Software Sucks
As others have pointed out, what you're describing isn't a fundamentally new idea or even that revolutionary. You're basically describing a database filesystem. Onne Gortner attempted an implementation of this concept in 2004 as part of his/her master's thesis (see http://dbfs.sourceforge.net/). Systems like Spotlight are effectively a partial implementation of this concept- OS X essentially has a hybrid setup where there's both a database and conventional filesystem running in parallel. Going back further, locate (first implemented in 1982) could almost be viewed as a proto-Spotlight. Gmail's labels/tags are another example of a mainstream implementation of this.
neuromantik8086··on AWS EC2/RDS Outage in us-east-1
Modern services such as reddit and Twitter effectively usurp the role that Usenet/NNTP and similar distributed protocols used to fulfill, but without the advantage of decentralization / lack of large single points of failure that such protocols embraced. That's what I was getting at, and maybe I'm full of shit.

In the 80s if a university campus internet connection went down, only that university was affected. Now, when a single AWS availability zone goes down, a much wider swath of users is impacted. Such consolidation / centralization shows a disregard for the spirit of the early internet and design considerations that went into it.

Again, maybe I'm full of shit. Lots of people here seem to think so.

neuromantik8086··on AWS EC2/RDS Outage in us-east-1
We wouldn't have this problem if people just used application-layer protocols and federated services like the early internet.
neuromantik8086··on Ask HN: Is Google Compute down?
Resource Public Key Infrastructure, but ISPs are too cheap to actually implement it.
neuromantik8086··on Show HN: HomelabOS – Ansible scripts to deploy self hosted cloud services
Maybe I'm being obtuse, but doesn't using a configuration management tool to deploy black-box Docker containers eliminate many of the advantages of using config management in the first place?
neuromantik8086··on The Art Institute of Chicago Has Put 50k High-Res Images Online
I mean, part of the appeal of the Louvre at least isn't just that you can see the art in the physical world, but that you're practically bathed in it. This sense of being overwhelmed and the serendipity factor in discovering new works constituted a lot of the appeal when I used to visit there. You could go into there once every weekend for a year and never have the same experience twice.
neuromantik8086··on 'Human brain' supercomputer with 1M processors switched on for first time
Just as a bit of a nitpick the Connection Machine:

a) Wasn't manufactured by Cray. It was made by Thinking Machines Corporation in the greater Boston area.

b) Didn't have anything to do with neural nets, as it was developed during the period of time when GOFAI / symbolic AI was still in vogue (although by the late 80s the Japanese had revived connectionism / neural nets), and thus had far more in common with a LISP machine.

c) Was mostly about developing a decent SIMD architecture.

neuromantik8086··on 'Human brain' supercomputer with 1M processors switched on for first time
I mean, a lot of the brain is devoted to sensation, so if you don't care about simulating how the brain interprets certain aspects of sensation (motion, depth, vision more generally) you could probably simulate other functions. For memory, however, at least, there's a lot of evidence that you'll need to simulate sensory systems to be able to accurately simulate recall [0].

[0] https://books.google.com/books?id=VjZyDwAAQBAJ&pg=PT597&lpg=...

neuromantik8086··on 'Human brain' supercomputer with 1M processors switched on for first time
> Science used to be culturally important in the 50s as a way to [beat the communists].

Fixed that for you

neuromantik8086··on How hustle culture took over advertising
It's a bit different in the U.S. in the sense that, whereas in Britain and France most mid-sized cities have decent public transit and municipal services, in the U.S. only a handful large urban centers have these kinds of things.

For example, in Clermont-Ferrand, which is roughly to Paris what Albany is in NYC, there is a well-established tram system that can get you to many places you'd want to go to without a car. This is especially essential if you're low income. Albany in contrast, just has buses (the CDTA), and pretty crap buses at that. The only other city in New York State with anything even remotely approaching the utility of the subway system is the light rail system in Buffalo, which is mostly useless since it only comprises one marginally useful line.

Believe me, as someone not involved in finance in NYC, if I could get the same services I get here in a mid-sized city like Pittsburgh, Cincinnati or Buffalo, I'd strongly consider relocation.

neuromantik8086··on 1 in 4 Statisticians Say They Were Asked to Commit Scientific Fraud
The HBR article's discussion of incentives is not really quite what I was thinking of when I wrote my comment. Specifically, the article you cite refers to the well-known phenomenon of how introducing extrinsic rewards via positive reinforcement is counterproductive in the long run. I've often noticed this form of "incentive" / reward being offered in the gamification of open science, such as via the Mozilla Open Science Badges [0], which in my opinion are a waste of time, effort, and money that do little to address systemic problems with scientific publishing.

With regard to the issue of grad students being unwilling to come forward and report mistakes, incentives wouldn't be added, but rather positive punishment [1] would be removed, which would then allow rewards for intrinsically motivated [2] actions.

[0] https://cos.io/our-services/open-science-badges/

[1] https://en.wikipedia.org/wiki/Punishment_(psychology)

[2] https://msu.edu/~dwong/StudentWorkArchive/CEP900F01-RIP/Webb...

neuromantik8086··on 1 in 4 Statisticians Say They Were Asked to Commit Scientific Fraud
Grad students / postdocs / human lab rats aren't scum, the incentives just aren't in place to promote good behavior (such as calling other researchers out on their bullshit). If you're trying to acquire a vaunted tenure track job, you can't afford to piss off $senior_tenured_researcher_at_prestigious_institution, since $senior could blacklist you so that you won't get hired at the incredibly small set of universities out there. Sometimes things work out despite pissing off major powers (Carl Sagan technically had to "settle" for Cornell due to being denied tenure at Harvard, in no small part because of a bad recommendation letter from Harold Urey [0]), but not often.

Even if you do manage to get a tenure track job, you pretty much have to keep your head down for 7 years in order to secure your position.

And once you have tenure, you still get attacked vociferously. Look at what happened when Andrew Gelman rightly pointed out that Susan Fiske (and other social psychologists) have been abusing statistics for years. Rather than a hearty "congratulations", he was called a "methodological terrorist" and a great humdrum came about [1].

When framed against these circumstances, it should be evident that there is literally nothing to gain and everything to lose from sending out a short e-mail pointing out that someone's model doesn't work.

[0] https://www.reddit.com/r/todayilearned/comments/15m8om/til_c...

[1] https://www.businessinsider.com/susan-fiske-methodological-t...

neuromantik8086··on Colorizing and restoring old images with deep learning
> And yes, I'm definitely interested in doing video

As someone familiar with the libraries space, I'd actually be very interested in seeing a machine learning model that could deal with "cleaning up" old film (I've actually brought this up w/ several of my ML friends occasionally). One of the biggest challenges in the world of media preservation is migrating analogue content to digital media before physical deterioration kicks in. Oftentimes, libraries aren't able to migrate content quickly enough, and you end up with frames that have been partially eaten away by mold.

As a heads-up, these are some of the problems you might encounter on the film front (which you might not otherwise find with photos due to differences in materials used, etc):

https://www.nyu.edu/tisch/preservation/program/05fall/physic...

https://www.filmpreservation.org/preservation-basics/vinegar...

neuromantik8086··on Study challenges conventional wisdom of how cell membranes work
One of the side effects of the corporatization of modern universities is that pretty much every scientific finding is accompanied by a PR fluff piece and scientific journalists usually just recycle aforementioned fluff piece. :/
neuromantik8086··on Why Jupyter is data scientists’ computational notebook of choice
> Those fields don't even know what their best practices are

"Best practices" are a chimera. The issue at hand isn't about what is "best", but whether or not a software engineer's "good enough" practices are more likely to achieve science's goals than a graduate student's "good enough" practices.

It's also disingenuous to claim that classical music composition doesn't have "best practices" when the field of music theory exists as an explicit manifestation of "best practices" in music. Having gone to a school with a conservatory, I also believe that I know several individuals who would would disagree with your mindset regarding how the creative process can't be managed. Indeed, if creativity, as it relates to musical composition, couldn't be managed most orchestras would be brimming with anger at the number of commissions that weren't finished on time for the concert, and most Hollywood studios and Broadway shows would screech to a halt.

neuromantik8086··on Why Jupyter is data scientists’ computational notebook of choice
I agree that science can't be bound by the rigid structures of most applied disciplines, and that the freedom to combine technologies in novel ways is a pre-requisite to novel findings.

What I find objectionable is the inability of scientists to explicitly delegate tasks to domain specialists in their everyday work when it makes sense. I think that it's unrealistic of you to believe that engineers always work with "a fully dimensioned and toleranced drawing" before starting work on a project and that your would work "grind to a halt". Indeed, there's a reason for the qualifier rapid in the term "rapid prototyping". If you can give an engineer general specifications for what you want and then leave him/her alone, he/she should be able to produce something that mostly fits your needs while avoiding all of the pitfalls that wouldn't have occurred to you. It would also be incorrect to assume that engineering does not involve creativity and is purely bound by rigid processes- if your requirements were strange enough, something fresh would inevitably be built.

This sort of delegation of course, is actually more efficient, since you can work on other tasks in parallel with the engineer (such as writing your next grant proposal or article or gasp teaching). Most scientists also already do this implicitly by choosing to purchase instrumentation from manufacturers like Olympus, Phillips, or Siemens rather than building it themselves.

Part of the reason for why I have such strong opinions about this matter, is that I've actually witnessed scientists waste more time messing around in fields where they were clearly out of their depth. As an example, there was a thread on a listserv in my (former) field that lasted for literally months that was solely devoted to the appearance of a website. Everyone wanted to turn the website design into an academic debate, when the website's creation (which had little to do with the substance of the scholarship itself) could have been turned over to a seasoned web developer and finished in less than a week or two.

neuromantik8086··on Why Jupyter is data scientists’ computational notebook of choice
> When I see stuff around notebooks for "reproducibility", I'm a bit confused in that notebooks often don't specify any guidance on installation and dependencies, let alone things like arguments and options that a regular old script would.

At the core of this, as some others may have already alluded to already, is that many academic scientists have not been socialized to make a distinction between development and production environments. Jupyter notebooks are clearly beneficial for sandboxing and trying out analyses creatively (with many wrong turns) before running "production" analyses, which ideally should be the ones that are reproducible. For many scientific papers, the analysis stops at "I was messing around in SPSS and MATLAB at 3 AM and got this result" without much consideration for reformulating what the researcher did and rewriting code/scripts so that they can be re-run consistently.

neuromantik8086··on How an outsider bucked prevailing Alzheimer's theory, clawed for validation
This isn't a particularly surprising story considering the history of science. To give an example of how this sort of thing played out historically, Marshall Nirenberg was basically persona non grata (or at least dutifully ignored) at many scientific conferences prior to his postdoc's discovery of the codon for phenylalanine. At the time, he was working at the NIH, which was considered very low prestige by many contemporary scientists. Somewhat fortuitously, Francis Crick heard a lecture by Nirenberg at a conference in 1961 and considered it good enough to bring to the attention of the other key players of the day, and Nirenberg was elevated from obscurity to stardom. For whatever reason, Nirenberg's postdoc never really achieved stardom, despite being the individual who actually made the discovery.

For more information:

https://www.telegraph.co.uk/news/science/science-news/854683...

neuromantik8086··on How an outsider bucked prevailing Alzheimer's theory, clawed for validation
Scientific / conference culture has a somewhat concerning relationship with alcohol, which was a contributing factor to my ultimately leaving science. The amount of pressure to go out to a pub with your colleagues was much higher for me than it was in my current job, in large part because serious discussions and informal deals that could directly impact your academic career tend to happen in these sorts of contexts. I suspected that a lot of my co-workers in science were high functioning alcoholics.
neuromantik8086··on How an outsider bucked prevailing Alzheimer's theory, clawed for validation
Yup. And the NIH even has pretty reports to hit you over the head with just how depressing the situation actually is for researchers (especially if you're young).

https://report.nih.gov/NIHDatabook/Charts/Default.aspx?chart...

https://report.nih.gov/DisplayRePORT.aspx?rid=827

neuromantik8086··on How an outsider bucked prevailing Alzheimer's theory, clawed for validation
He (or she)'s not disagreeing with the notion that unfettered funding produces innovation, he (or she)'s disagreeing with the notion that HHMI provides unfettered funding.

Part of the beauty of Bell Labs was that it wasn't necessarily all about fame and grandstanding- a lot of folks who were working there were just normal, unassuming New Jersey folk who put in a good days work messing around with whatever they were tasked with. The only analogous example to that sort of dynamic today might be amongst the workers at large government agencies like NIST and the NIH; it's unlikely that many of them will ever become rockstars in today's research climate, but they dutifully carry out experiments nonetheless.

Page 1 of 11Next →