HNHacker News
TopNewBestAskShowJobs

uniqueuid

3,775 karma · joined June 17, 2021

Computational social scientist.
submissionscomments
uniqueuid··on AI systems with 'unacceptable risk' are now banned in the EU
Unfortunately yes, the article is a simplification, in part because the AI act delegates some regulation to existing other acts. So to know the full picture of AI regulation one needs to look at the combination of multiple texts.

The precise language on high risk is here [1], but some enumerations are placed in the annex, which (!!!) can be amended by the commission, if I am not completely mistaken. So this is very much a dynamic regulation.

[1] https://artificialintelligenceact.eu/article/6/

uniqueuid··on ZFS 2.3 released with ZFS raidz expansion
True, but I have a gut feeling that a lot of these thorny issues would come up again:

https://github.com/openzfs/zfs/issues/3582

uniqueuid··on ZFS 2.3 released with ZFS raidz expansion
Yes but see my sibling comment.

When you expand your array, your existing data will not be stored any more efficiently.

To get the new parity/data ratios, you would have to force copies of the data and delete the old, inefficient versions, e.g. with something like this [1]

My personal take is that it's a much better idea to buy individual complete raid-z configurations and add new ones / replace old ones (disk by disk!) as you go.

[1] https://github.com/markusressel/zfs-inplace-rebalancing

uniqueuid··on ZFS 2.3 released with ZFS raidz expansion
It's good to see that they were pretty conservative about the expansion.

Not only is expansion completely transparent and resumable, it also maintains redundancy throughout the process.

That said, there is one tiny caveat people should be aware of:

> After the expansion completes, old blocks remain with their old data-to-parity ratio (e.g. 5-wide RAIDZ2, has 3 data to 2 parity), but distributed among the larger set of disks. New blocks will be written with the new data-to-parity ratio (e.g. a 5-wide RAIDZ2 which has been expanded once to 6-wide, has 4 data to 2 parity).

uniqueuid··on Monkeys Can Predict Election Outcomes
Ok so the correlation coefficient for vote share is .048, meaning the R squared or explained variance is .0028 i.e. around three percent. No, monkeys cannot predict vote share.

That said, for a paper with such an obvious bullshit design the text itself is not that bad, at least they seriously investigate the facial features that seem to be relevant to these cases.

If you must remember this paper (please don’t), do it as “monkey ganze is a better predictor for election races in the US than a coin throw”.

uniqueuid··on WhisperNER: Unified Open Named Entity and Speech Recognition
It's so great to see that we finally move away from the thirty year old triple categorization of people, organizations and locations.

This of course means that we now have to think about all the irreconcilable problems of taxonomy, but I'll take that any day over the old version :)

uniqueuid··on Discarded delights: The joy of ex-library books (2021)
And sometimes you find little gems, too!

I have a copy of "New Rules for the New Economy" by Kevin Kelly, signed as part of the Global Business Network that he and Steward Brand founded a long time ago.

Having read Fred Turner's immensely great book "From Counterculture to Cyberculture", that is a valuable little piece of history to me.

uniqueuid··on New sequencer from Teenage Engineering: OP-XY
Looks interesting - they even added a pitch bend. But I'm wondering if it's worth the 2.3k euro.
uniqueuid··on Hong Kong Jails Benny Tai for 10 Years in Longest Security Law Sentence
The promise to implement long-announced and in part contractually guaranteed democratic instruments.

https://www.goodreads.com/book/show/57873484-freedom

uniqueuid··on Analysis of economic and productivity losses caused by cookie banners in Europe
Sure I find it reasonable to disagree on these points.

I personally find informed consent to be a very desirable thing, because it aims at the goal of legislation, not at the means. If you think that citizens cannot, should not, or should not be required to profoundly understand what is happening to them in digital contexts, that's a specific point of view. From this you evaluate the trade-offs.

My personal (humanistic) perspective is that a profound understanding and practical control over our digital lives are the prerequisite for dignity, which is the ultimate goal of a state.

uniqueuid··on Analysis of economic and productivity losses caused by cookie banners in Europe
This is not what I meant. Laws are made concrete and understandable through either case law (harder for citizens to anticipate IMO) or through statutory interpretation in civic law traditions. Both (eventually) offer a clear understanding of the meaning and scope of a law.
uniqueuid··on Analysis of economic and productivity losses caused by cookie banners in Europe
I am kind of frustrated by the widespread misunderstandings in this thread.

Laws are best when they are abstract, so that there is no need for frequent updates and they adapt to changing realities. The European "cookie law" does not mandate cookie banners, it mandates informed consent. Companies choose to implement that as a banner.

There is no doubt that the goals set by the law are sensible. It is also not evident that losing time over privacy is so horrible. In fact, when designing a law that enhances consumer rights through informed consent, it is inevitable that this imposes additional time spent on thinking, considering and acting.

It's the whole point, folks! You cannot have an informed case-by-case decision without spending time.

uniqueuid··on TinyTroupe, a new LLM-powered multiagent persona simulation Python library
Yup, looks like their azure api configuration is just a generic wrapper for openapi in which you can plug any endpoint url. Nice.

https://github.com/microsoft/TinyTroupe/blob/7ae16568ad1c4de...

uniqueuid··on TinyTroupe, a new LLM-powered multiagent persona simulation Python library
Needs openai or azure APIs. I wonder if it's possible to just use any openapi-compatible local provider.
uniqueuid··on Programmer in Berlin: Culture
Although you probably agree: If someone wants to describe politics with a one-dimensional scale, left-right is not so bad. And that's why and how it developed.
uniqueuid··on Moving to a World Beyond "p < 0.05" (2019)
Unfortunately it‘s true that most studies are useless or even harmful (limited to certain disciplines).
uniqueuid··on Moving to a World Beyond "p < 0.05" (2019)
It is also not easy if you have many potential covariates! Because statistically, you want a complete (explaining all effects) but parsimonious (using as few predictors as possible) model. Yet you by definition don‘t know the true underlying causal structure. So one needs to guess which covariates are useful. There are also no statistical tools that can, given your data, explain whether the model sufficiently explains the causal phenomenon, because statistics cannot tell you about potentially missing confounders.

A cool, interesting, horrible problem to have :)

uniqueuid··on Moving to a World Beyond "p < 0.05" (2019)
The real underlying problem is that in your case, genetic variants are not accounted for. As soon as you include these crucial moderating covariates, it‘s absolutely possible to find true effects even for (rather) small samples (one out of a hundred is really to few for any reasonable design unless it‘s longitudinal)
uniqueuid··on Moving to a World Beyond "p < 0.05" (2019)
The best way to talk about this is IMO effect heterogeneity. Underlying that you have the causal DAG to consider, but that‘s (a) a lot of effort and (b) epistemologically difficult!
uniqueuid··on SSH Remoting
"Man, just put it on tape and walk over to that machine to load it there".

Sorry, I know such low-effort puns are shunned on HN, but once in a decade I grant myself the permission to not resist.

uniqueuid··on OpenZFS deduplication is good now and you shouldn't use it
That's not true, you commonly have CDX index files which allow for de-duplication across arbitrarily large archives. The internet archive could not reasonably operate without this level of abstraction.

[edit] Should add a link, this is a pretty good overview, but you can also look at implementations such as the new zeno crawler.

https://support.archive-it.org/hc/en-us/articles/208001016-A...

uniqueuid··on OpenZFS deduplication is good now and you shouldn't use it
I get the use case, but in most cases (and particularly this one) I'm sure it would be much better to implement that client-side.

You may have seen in the WARC standard that they already do de-duplication based on hashes and use pointers after the first store. So this is exactly a case where FS-level dedup is not all that good.

uniqueuid··on You-get: Dumb downloader that scrapes the web
That's an interesting question. They only depend on a single library, but I wonder how much code is really their own. I found it curious, for example, that there is a dedicated mp4 joiner (I mean, if you already have ffmpeg, there is probably no way you can do it better yourself).

https://github.com/soimort/you-get/blob/develop/src/you_get/...

uniqueuid··on The Internet Archive is back online
The problem is that it's hard to do this in a way that ensures good archival of ALL resources.

Bittorrent works well for popular things but fails for marginal content (unless some really dedicated individuals step in.)

What the internet archive provides is a way to have access to many many resources which you didn't know you needed in advance.

uniqueuid··on Why has nuclear power been a flop? (2021)
Haha thanks fixed.
uniqueuid··on Why has nuclear power been a flop? (2021)
I agree that it's stupid not to take the game-changer of solar and wind (and batteries) into account.

Even though I believe that nuclear no longer has a role in energy (fuel sources, disposal, concentration of risk into few small units), the statement is still correct and it has another dimension:

There is a conflict between poverty, climate change, gravity of climate change harms and the speed at which we can reduce climate change impacts.

The problem is that there is a shrinking window to limit harm, and the harms will disproportionally affect poor nations which at the same time lack resources to mitigate. I'd definitely call that a gordian knot.

uniqueuid··on The AI Scientist: Towards Automated Open-Ended Scientific Discovery
Maybe, maybe not. It's a tiered system - you get the deluge at the unfiltered bottom and a narrower selection the more prestigious and selective the outlets / conferences / journals are.

Problem is, of course, that selection criteria are in large parts proxies, not measures of quality. With AI, those proxies become tainted and then you get an explosion of effort.

If anyone has a good recommendation for scalable criteria to assess the quality of papers (beyond fame haha) I'm all ears.

uniqueuid··on The AI Scientist: Towards Automated Open-Ended Scientific Discovery
It would also be sad to see the scientific system destroyed by a wave of automatically generated papers that no human has the capacity to verify.

It's not hard to generate ideas, it's hard to generate reliable and relevant ideas. Such AI science generators are destroying the grass they graze on unless they take science more seriously (and not as a toddler idea of "generating and testing ideas", which is only a small part of the story).

uniqueuid··on The EU should be the heat-pump pioneer
That's an overstatement. Also, people routinely rebel against the introduction of change, but soon become accustomed to it. For quite over a year, heat pumps have made up over 50% of new installs [1]. I don't think it's useful to focus too much on short-term pains here.

[1] https://www.bdew.de/service/daten-und-grafiken/entwicklung-b...

uniqueuid··on Enhancing R: The Vision and Impact of Jan Vitek's MaintainR Initiative
Interesting that they mention Unicode problems on Windows. I've ran into these a couple of times where data exported from Windows R had unicode codepoints swapped and/or double-encoded.

The all-in-one ecosystem of R is nice, but text encoding is still a major pain point (e.g. people try to put emoji into RMD to translate into pdf via tinytex, and fail miserably, of course).

← PreviousPage 4 of 27Next →