HNHacker News
TopNewBestAskShowJobs

andreasvc

1,599 karma · joined December 22, 2008

http://andreasvc.github.io
submissionscomments
andreasvc··on I Violated a Code of Conduct
Why would we never know? Quoting from the blog post: "The specific reasons given were that [...]" followed by a bullet list. Your guess is unwarranted speculation.
andreasvc··on Python 3.9
Ah, I see. That really doesn't deserve to be called an idiom, it's a clever hack. But it's nice to know about it. It seems less ugly than the walrus operator to me, and it doesn't leak the variable outside of the comprehension.
andreasvc··on Python 3.9
I'm puzzled by this example:

    sums = [s for s in [0] for x in data for s in [s + x]]
Why would you do "for s in" twice? Is that intentional? It would make more sense to me if the variables would have been different. And why would you want to add 0 to numbers?! Curious about a real world use case for this.
andreasvc··on What the interns have wrought, 2020 edition
What are the advantages of Re over Re2?
andreasvc··on Differences between the word2vec paper and its implementation
bollu, multiple people have shown that your claim that the paper doesn't match the code is flat out wrong. I think at this point you should issue a retraction of your wildly inappropriate suggestion of academic dishonesty.
andreasvc··on Tips and tricks to write LaTeX papers in with figures generated in Python
It is an extremely thin font which makes it unsuitable for screen reading.
andreasvc··on The World's Most Efficient Languages (2016)
Why do you think English would be least compressible? Is that based on conjecture or have you investigated this? Why would artificial language be more compressible? That seems completely orthogonal to me (by definition, an artificial language can be designed with whatever properties you choose). Fortran may be more compressible due to its limited set of keywords, but it's my impression that Ithkuil is by design more information dense and thus harder to compress than English.

The most efficient language is the least compressible language only in a narrow and arbitrary sense of efficient. There are many considerations such as what is efficient for the speaker, the hearer, redundancy to noise, efficiency with respect to particular purposes, etc. We can assume that natural languages will generally make a good trade-off across these factors, and searching for the most efficient language in one particular narrow sense is not very useful. Moreover, compression of text focuses only on surface form, completely ignoring the dimension of meaning.

andreasvc··on Practical Text Classification with Python and Keras
Yes, in the context of text classification a bag of words model will refer to that, or combined with some other linear model like linear SVM or naive bayes.

The queen - woman example is when you try to make a model of word semantics, such as with word2vec. In a document classification task the vectors represent documents.

andreasvc··on Practical Text Classification with Python and Keras
You mention a single counterexample, but is it "definitely not true"? I think there is a strong publication bias: papers will report when they improve on the baseline, but that's likely not representative for the common case (e.g., limited data, no pretrained model available, no time for extensive parameter tuning).
andreasvc··on Practical Text Classification with Python and Keras
You're right that it is a representation, but also an instance of the vector space model of language. Coupled with a linear model for prediction it is a strong baseline for text classification problems. See e.g. http://scikit-learn.org/stable/tutorial/text_analytics/worki...
andreasvc··on Unpublished and Untenured, a Philosopher Inspired a Cult Following
The article states he's perfectionist. He probably wouldn't be OK with that.
andreasvc··on Mmm, Pi-hole
Yup, that's a downside. The advantage is that it's much simpler and will also work when you're not on your home network.
andreasvc··on Mmm, Pi-hole
The simplest approach is to use a hosts file: https://someonewhocares.org/hosts/
andreasvc··on Doug McIlroy's C++ regular expression matching library
Regarding that idea for a universal integer set, roaring bitmaps implement that.
andreasvc··on Women Once Ruled the Computer World. When Did Silicon Valley Become Brotopia?
Would love to read more about this. I've read that early programming was thought of like a glorified telephone operator, but that these women had advanced mathematics degrees.

I can't find anything non-biased though, everyone just wants to push the female victim narrative.

andreasvc··on Lying to TCP makes it a best-effort streaming protocol
To clarify, what I meant is that with TCP, I can set up a two-way communication channel, even if I'm behind a NAT/firewall I don't control. As far as I understand, with UDP this is harder (i.e., does not work with all NAT types), because UDP does not establish a connection, and it does not provide a two-way communication channel. However, I am not up to date on NAT traversal techniques, so I might be wrong.
andreasvc··on Lying to TCP makes it a best-effort streaming protocol
No. UDP is connectionless, which makes it harder to use with NAT.
andreasvc··on ‘Saluton’: the surprise return of Esperanto
> machines like Parsey McParseFace can now accurately parse human universal grammar. unambiguous parsing was a big feature of lojban, so it's lost a major selling point.

Slow down there ... Most linguists (or a substantial minority) do not subscribe to the theoretical idea of universal grammar. Parsey McParseFace is only an incremental improvement in decades of statistical parsing. The problem of natural language understanding remains largely unsolved. Like other deep learning models, it is not hard to come up with adversarial examples that will confuse the parser but not any competent human language user. The parser is only as good as the data it is trained on; this data is expensive to acquire but there is never enough. Additionally there are a myriad other kinds of ambiguity in language beyond the sentence-level syntactic ambiguity which is resolved by a statistical parser such as Parsey McParseFace.

In my opinion Lojban provides a good illustration of just how hard it is to remove ambiguity from human language.

andreasvc··on The Neural Net Tank Urban Legend
For a better, actual example of this problem, see the leopard sofa: http://rocknrollnerd.github.io/ml/2015/05/27/leopard-sofa.ht...
andreasvc··on Linked Lists Are Still Hard
Why would one ever prefer a linked list over a dynamic array? I know about the asymptotic performance, but if you take actual memory performance into account (locality, cache, pointers are slow, etc), it seems dynamic arrays should be simpler and perform better.
andreasvc··on Ask HN: Vim users, what's your favourite colorscheme?
My own, https://github.com/andreasvc/vim-256noir
andreasvc··on Semantics derived automatically from language corpora contain human-like biases
A fact is a statement corresponding with reality, a stereotype is a belief/generalization deemed harmful/undesirable.
andreasvc··on Semantics derived automatically from language corpora contain human-like biases
Again, from humans and from reality is not different. Whether it is desirable to avoid stereotypes depends on your values.
andreasvc··on Semantics derived automatically from language corpora contain human-like biases
I think you are making a distinction without a difference. If the word vectors pick up biases from wikipedia text, than for all practical purposes, they are (indirectly) absorbing stereotypes from humans. This is an expected result, but not necessarily desirable in the end.
andreasvc··on Semantics derived automatically from language corpora contain human-like biases
Again, per the article "bias refers generally to prior information, a necessary prerequisite for intelligent action (4)." This includes a citation to a well-known ML text. This seems broader than the statistical definition you cite.

Think for example of an inductive bias. If I see a couple of white swans, I may conclude that all swans are white, and we all know this is wrong. Similarly, I may conclude the sun rises everyday, and for all practical purposes this is correct. This kind of bias is neither wrong nor right, but, in the words of the article "a necessary prerequisite for intelligent action", because no induction/generalization would be possible without it.

There are undoubtedly examples where the prejudiced kind of biases lead to both truthful and untruthful predictions, but that seems beside the point, which is to design a system with the biases you want, and without the ones you don't.

andreasvc··on Semantics derived automatically from language corpora contain human-like biases
It appears that the linked paper has examples.
andreasvc··on Semantics derived automatically from language corpora contain human-like biases
I think your questions would be answered by reading the article. Particularly:

"In AI and machine learning, bias refers generally to prior information, a necessary prerequisite for intelligent action (4). Yet bias can be problematic where such information is derived from aspects of human culture known to lead to harmful behavior. Here, we will call such biases “stereotyped” and actions taken on their basis “prejudiced.”"

This definition is not unusual. This is about inferences that are wrong in the sense of prejudiced, not necessarily inaccurate.

andreasvc··on Unreasonable Ineffectiveness of Machine Learning in Computer Systems Research
Cute title but the post didn't really address either reasonableness or effectiveness, but mostly claimed that the potential has not yet been realized. It's a pet peeve of mine to see these hackneyed joke titles referencing famous papers, "considered harmful" is another case in point. Let's just stick to descriptive titles.
andreasvc··on Statistician Proves Gaussian Correlation Inequality
Mathematics has this funny habit of throwing around "it is easy to say that ...", or in this case, calling a proof that eluded mathematicians for decades "simple."
andreasvc··on Evidence Rebuts Chomsky's Theory of Language Learning
He's hugely influential in the field; arguably dominant in a way that is unhealthy for a field. Acolytes will say he's influential because of merit, detractors will say it is because of rhetorical skill, sophistry. I hope there will come a time when press releases no longer breathlessly announce Chomsky's theories falsified for the umpteenth time, but rather actual progress.
Page 1 of 34Next →