HNHacker News
TopNewBestAskShowJobs

throwawaywego

92 karma · joined June 19, 2019

submissionscomments
throwawaywego··on YouTube bans 14 year old for violating hate speech guidelines
She can't keep getting away with this.
throwawaywego··on YouTube bans 14 year old for violating hate speech guidelines
And for context:

https://i.imgur.com/NpvySvG.jpg (Pewdiepie)

(It's edgy humor/let kids play on their skateboard)

throwawaywego··on YouTube bans 14 year old for violating hate speech guidelines
Not a credible threat, according to police. Also, only after she was targeted by the media (a common pattern for Youtube to take action).

> “Susan, I’ve known your address since last summer. I’ve got a Luger and a mitochondrial disease. I don’t care if I live. Why should I care if you live or your children? I just called an Uber. You’ve got about seven minutes to draft up a will. … I’m coming for you, and it ain’t gonna be pretty.”

throwawaywego··on Sites using Facebook ‘Like’ button liable for data, EU court rules
Yup. And Google Analytics should (and I believe already has) be treated similarly to a tracking beacon.
throwawaywego··on GitHub sued for aiding hacking in Capital One breach
Worse. You can make the ML model spit out the SSNs it was trained on. That's a problem when you can't manually curate billions of documents. If you didn't look, you wouldn't even know they were there.
throwawaywego··on GitHub sued for aiding hacking in Capital One breach
Can you yourself reliably detect if something is a list of SSN's or some other type of identifier?

I think you can. Now when we have neural nets capable of distinguishing 100s of different dog breeds, we still have to trip on the most basic and structured type of entity extraction? No. A simple regex, combined with heuristics, and a linear model on top can reliably detect SSN's.

Just that there will be a few false positives (no matter if you automate this, or do this manually) does not mean it is a Herculean technical challenge to do this.

throwawaywego··on GitHub sued for aiding hacking in Capital One breach
In most modern browsers you can't capture this when find-in-page box has focus. Only if you manually select text after searching can you capture it (or, like the other reply stated: capture the scroll distance).

But you could easily disable ctrl+f and throw up your own search box with keypress capture. Not all browsers show the search box outside of the browser viewport, and even for those that do (such as Chrome), you could display a hovering modal inside the viewport, as most users won't remember the exact location of the search box.

throwawaywego··on Sites using Facebook ‘Like’ button liable for data, EU court rules
That's a very D&D Rulebook type interpretation :). Not that rulings should be ambiguous, but usually some common sense can be applied (and is expected to be reasonably applied).

The Facebook like button is a web tracker, disguised as a social engagement button. If not its primary -, then its secondary function is to (indiscriminately) track users and non-users outside of its walled garden, like some reversed Trojan Horse.

Hotlinking an image is just that: hotlinking an image. Facebook relies on us and lawmakers to say: "We just can't ban third party content!", while we perfectly could leave innocent third party content alone, and focus our sights on the spy button. It isn't reasonable, nor common sense to conflate the two: even if similar in syntax, the context is vastly different.

throwawaywego··on GitHub sued for aiding hacking in Capital One breach
What I think happened: Someone contacted Capital One by email to responsibly disclose to them that there were SSN's and other data on a Gist. That person found them with a simple crawler or search.

Then Cap 1 thought: If some rando can find this after a lot of damage has been done, why can't Github find these seconds after upload?

And, really, there is no technical excuse. It is perfectly possible to do this, and lots of big companies do this (or hire security companies to do this for them). Mention their name on some deep web hacking forum, a pastebin, or inside Github code, and somewhere an alarm goes off.

Github could (and should) warn if a user uploads loads of PII-like data. For the cost of running a search server and a few moderators. "Are you sure you want to upload your AWS credentials in a public repository?".

Github is somewhere halfway between moderated and a content platform. They already have a history of taking down repositories if they link to PII data (or infringe copyright, or damage U.S. national security): http://web.archive.org/web/20180619172528/https://github.com... so not acting on this specific repo with SSN numbers could be seen as a poor/shoddy job on their part. Github is certainly in the dominant position to mitigate spread of PII data, so they should have their stuff in order.

throwawaywego··on Homomorphic encryption
https://www.microsoft.com/en-us/research/publication/crypton... (2016)

> We demonstrate CryptoNets on the MNIST optical character recognition tasks. CryptoNets achieve 99% accuracy and can make more than 51000 predictions per hour on a single PC. Therefore, they allow high throughput, accurate, and private predictions.

throwawaywego··on Homomorphic encryption
Cloud Computing. Use the compute from a big company, without that company possibly knowing what they are computing on.

Machine Learning. Apply neural networks to encrypted data and avoid privacy issues.

throwawaywego··on Microsoft is investing $1B in OpenAI

    1943 First mathematical neural network model
    1958 Learning neural network classifies objects in spy plane photos
    1965 Deep learning with multi-layer perceptrons

    2010 ImageNet error rate 28%
    2011 ImageNet error rate 25%
    2012 ImageNet error rate 16%
    2013 ImageNet error rate 11%
    2017 ImageNet error rate 3%
    2019 Pre-AGI
throwawaywego··on Microsoft is investing $1B in OpenAI
I think any AI researcher has a tale where an algorithm they wrote genuinely took them by surprise. Not due to wrong calculations, but by introducing randomness, heaps of data, and game bounderaries where the AI is free to fill in the blanks.

A good example of this is "move 37" from AlphaGo. This move surprised everyone, including the creators, who were not skilled enough in Go to hardcode it: https://www.youtube.com/watch?v=HT-UZkiOLv8

throwawaywego··on Microsoft is investing $1B in OpenAI
https://arxiv.org/abs/1502.04156

> Towards Biologically Plausible Deep Learning

> Neuroscientists have long criticised deep learning algorithms as incompatible with current knowledge of neurobiology. We explore more biologically plausible versions of deep representation learning, focusing here mostly on unsupervised learning but developing a learning mechanism that could account for supervised, unsupervised and reinforcement learning. The starting point is that the basic learning rule believed to govern synaptic weight updates (Spike-Timing-Dependent Plasticity) arises out of a simple update rule that makes a lot of sense from a machine learning point of view and can be interpreted as gradient descent on some objective function so long as the neuronal dynamics push firing rates towards better values of the objective function (be it supervised, unsupervised, or reward-driven). The second main idea is that this corresponds to a form of the variational EM algorithm, i.e., with approximate rather than exact posteriors, implemented by neural dynamics. Another contribution of this paper is that the gradients required for updating the hidden states in the above variational interpretation can be estimated using an approximation that only requires propagating activations forward and backward, with pairs of layers learning to form a denoising auto-encoder. Finally, we extend the theory about the probabilistic interpretation of auto-encoders to justify improved sampling schemes based on the generative interpretation of denoising auto-encoders, and we validate all these ideas on generative learning tasks.

throwawaywego··on Microsoft is investing $1B in OpenAI
All sciences that collaborate with the field of AI: Cognitive Science, Neuroscience, Systems Theory, Decision Theory, Information Theory, Mathematics, Physics, Biology, ...

Any AI curriculum worth its salt includes the many scientific and philosophical views on intelligence. It is not all alchemy, though the field is in a renewal phase (with horribly hyped nomenclature such as "pre-AGI", and the most impressive implementations coming from industry and government, not academia).

And eventhough the atom bomb was based on science too, there is this anecdote from Hamming:

> Shortly before the first field test (you realize that no small scale experiment can be done—either you have a critical mass or you do not), a man asked me to check some arithmetic he had done, and I agreed, thinking to fob it off on some subordinate. When I asked what it was, he said, "It is the probability that the test bomb will ignite the whole atmosphere." I decided I would check it myself! The next day when he came for the answers I remarked to him, "The arithmetic was apparently correct but I do not know about the formulas for the capture cross sections for oxygen and nitrogen—after all, there could be no experiments at the needed energy levels." He replied, like a physicist talking to a mathematician, that he wanted me to check the arithmetic not the physics, and left. I said to myself, "What have you done, Hamming, you are involved in risking all of life that is known in the Universe, and you do not know much of an essential part?" I was pacing up and down the corridor when a friend asked me what was bothering me. I told him. His reply was, "Never mind, Hamming, no one will ever blame you."

throwawaywego··on Microsoft is investing $1B in OpenAI
> The atomic bomb was based on science theory.

Our study of (automated) intelligence is based on science too.

> A computer ... will never be able to think by itself.

Turing wrote an entire paper about this (Computing Machinery and Intelligence), where he rephrases your statement (because he finds it to be meaningless) and devises a test to answer it. He also directly attacks your phrasing of "but it will never":

> I believe they are mostly founded on the principle of scientific induction. A man has seen thousands of machines in his lifetime. From what he sees of them he draws a number of general conclusions. They are ugly, each is designed for a very limited purpose, when required for a minutely different purpose they are useless, the variety of behaviour of any one of them is very small, etc., etc. Naturally he concludes that these are necessary properties of machines in general.

> A better variant of the objection says that a machine can never "take us by surprise." This statement is a more direct challenge and can be met directly. Machines take me by surprise with great frequency. This is largely because I do not do sufficient calculation to decide what to expect them to do, or rather because, although I do a calculation, I do it in a hurried, slipshod fashion, taking risks.

throwawaywego··on A brief history and future of credit scores
Please reread the article. The article itself admits that non-financial data is useful and used (where this is allowed). I agree 100%.

You drew the wrong conclusion about something you don't know a lot about and doubled down. Good luck with that and I forgive you.

throwawaywego··on A brief history and future of credit scores
The evidence is in the article. It tells of a US company using 10.000 data points in a jurisdiction where this is allowed.

The article cites Turner, who a few years back wrote this:

> The non-financial data was found predictive in all three outcomes examined when no other ‘traditional’ credit information was used, strongly suggesting that alternative data would be useful to lenders in underwriting the so-called ‘no-file’ or ‘no-score’ consumer who have little or no payment/credit information available.

I think his quote about "laughing out" is taken out of context. No predictive modeler will throw away informative features, because she can not distinguish noise from signal after 26 variables. That's 10 to 1% of a modern credit scoring model. It may be the perspective of a regulator though (they start drowning in noise after reviewing 100+ variables).

Yes, all data that is legal to use and predictive, will get used, if not by you, then by your competitor.

And informative variables that can not be used in the decision to give a loan, are used internally to predict if the loan will be paid back. There is more to credit scoring than the initial yes-no.

throwawaywego··on A brief history and future of credit scores
Other comments have pointed out the risk of redundant encodings /proxy variables.

I like to point out that sources of unfairness are possible, even with 100% decorrelated non-protective variables. For instance, the data collection and labeling of the data may be biased (You label re-admission to jail with current criminals in jail, and just overfitted to the war on drugs).

Or you make a biased decision based on the output of a fair model. Things like not taking into account sample size / uncertainty.

Also, breaking the law is at odds with unethical behavior. Unethical behavior is not necessarily breaking the law, but it is still nasty. For instance, from your Facebook likes (not a protected variable) I could deduce all of race, sex, origin, religion. Not against the law, but still discrimination.

throwawaywego··on A brief history and future of credit scores
> The key here being that non-financial data isn't actually useful in predicting ability to repay loans.

Sure it is. Even religion correlates with loan risk. (Could be a proxy for social status and people from your tribe helping you out when you can't pay back). I could probably get a predictive model better than random guessing by mining your HackerNews comments or Facebook likes.

> I'm sure it'd be used somehow by modern financial institutions if it were predictive.

It is used. All data that is even remotely informative is used. To the fullest extend made possible by jurisdiction/ anti-discrimination laws.

throwawaywego··on A brief history and future of credit scores
There is multiple elements to your question.

Two ways where these systems may give out more loans than is strictly profitable, and they are both investments:

- Fairness. If you have a variable race and a zip code, you could account for discrimination via redundant encodings, while still using the feature for the optimal trade-off between a fairness criteria and model performance.

- Exploration. Concept drift (the correlational and causal meaning of variables shifts over time) can introduce wrong predictions. If all you have is few samples from a zip code, the model will always be uncertain. You can counter this by exploration and active learning: gather samples, not because this makes the model max-profit, but gather samples, to better learn how to predict these samples in the future.

But yes, giving out too many loans to minorities, may very well lead to further crisis and defaults, and tainting the credit scores of people will low access to finances even further. A bit like how well-meaning people donate money and food to Africa during Christmas time, then a few months later when donations subside, there are increases in famines. There is such a thing as "being too good".

throwawaywego··on YouTube bans content “showing users how to bypass secure computer systems”
There is an ongoing lawsuit against Jones. The lawsuits contain unproven claims of harassment.

There is no direct inciting, and the indirect inciting claims are half-baked. There is a class action lawsuit office looking for a pay day, so they benefit by keeping Jones in the news.

All people factually saying that Alex Jones incited harassment against the parents, are parroting a non-concluded court case PR campaign ("OJ did it!"), and much like these behemoth companies, are playing judge, jury, and executioner. They most likely haven't seen a clip of Alex Jones inciting, or when they have, it is taken grossly out of context.

Remember, when CNN was targeting Jones for weeks, going through his hours of video, and noting anything they found controversial. "Alex Jones is transphobic, which is against Youtube TOS, but Youtube does nothing". Alex Jones was talking about "public library drag queen reading hour". He said these drag queens looked demonic and that it is not normal to normalize this for little kids. Now, just recently, someone ("a Trump supporter") showed up with a gun to these reading hours. Two possible conclusions: "Alex Jones is a transphobe and incited his followers to violently get rid of them". "Alex Jones practices unpopular, but legal, free speech, and they try to punish him for what any of his million followers might do that is against the law". I am leaning heavily towards the latter and it is a downright shame that the first conclusion is drawn by many, without doing any deeper research.

Imagine what you could do if you knew that many people would go for the first conclusion, regardless of the facts?

throwawaywego··on How Artificial Intelligence Is Changing Science
The article states positive impacts on science, but there are also negative impacts on science. For instance, the hype of AI has caused a brain-drain on related fields (such as cognitive science or applied mathematics). AI research itself suffers from companies buying up the academic talent. And researchers slap AI (which is usually deep learning) on a decade-old problem, without any care for complexity/benchmarks, implementation/usage, and proper validation methods, just to get published or receive funding.
throwawaywego··on Dutch Telephone Outage Takes Out Nation's Emergency Number
Me too. But it is a mere suspicion. We know that:

- GRU has been stepping up its attacks on the West (anti-doping, chemical weapons, MH17 investigation)

- US, Britain, and Israel attacked Iran with Stuxnet

- Iran retaliated with cyber attacks on banks and electrical grids

- Iran downed a drone

- Number stations go crazy

- US almost goes for military strike, but Trump calls it off and orders US Cyber Command to cyber-attack Iran, with focus on missile systems and spy networks.

- BGP route leak (Allegheny Technologies Inc. caused the leak, and was a target for Titan Rain around 2014).

- Dutch payment - and emergency response network goes down.

Edit: According to a KPN spokesperson: "We have an indication of the cause, but a lot remains unclear. Anyway, this was not the result of a hack."

throwawaywego··on The case of Chinese LinkedIn spy recruitment [pdf]
https://gizmodo.com/job-networking-site-linkedin-filled-with...

> Want to know all the code names for America's massive intelligence gathering programs? Just browse through the "intelligence analysts" who post their resumes on the public career networking site LinkedIn.

throwawaywego··on The case of Chinese LinkedIn spy recruitment [pdf]
Got an invite from a fake account claiming to be a professor at my alma mater. The name only exists on LinkedIn and has currently a handful of connections. I reported the account, but it remains up as of today. I do not have a public account.

Made me paranoid enough to stop accepting invites (~1000+ pending) and drop all usage of LinkedIn. Very discomforting to know you are a target. LinkedIn will have to visibly move on this for me to change my outlook and regain trust.

← PreviousPage 2 of 2