92 karma · joined June 19, 2019
https://i.imgur.com/NpvySvG.jpg (Pewdiepie)
(It's edgy humor/let kids play on their skateboard)
> “Susan, I’ve known your address since last summer. I’ve got a Luger and a mitochondrial disease. I don’t care if I live. Why should I care if you live or your children? I just called an Uber. You’ve got about seven minutes to draft up a will. … I’m coming for you, and it ain’t gonna be pretty.”
I think you can. Now when we have neural nets capable of distinguishing 100s of different dog breeds, we still have to trip on the most basic and structured type of entity extraction? No. A simple regex, combined with heuristics, and a linear model on top can reliably detect SSN's.
Just that there will be a few false positives (no matter if you automate this, or do this manually) does not mean it is a Herculean technical challenge to do this.
But you could easily disable ctrl+f and throw up your own search box with keypress capture. Not all browsers show the search box outside of the browser viewport, and even for those that do (such as Chrome), you could display a hovering modal inside the viewport, as most users won't remember the exact location of the search box.
The Facebook like button is a web tracker, disguised as a social engagement button. If not its primary -, then its secondary function is to (indiscriminately) track users and non-users outside of its walled garden, like some reversed Trojan Horse.
Hotlinking an image is just that: hotlinking an image. Facebook relies on us and lawmakers to say: "We just can't ban third party content!", while we perfectly could leave innocent third party content alone, and focus our sights on the spy button. It isn't reasonable, nor common sense to conflate the two: even if similar in syntax, the context is vastly different.
Then Cap 1 thought: If some rando can find this after a lot of damage has been done, why can't Github find these seconds after upload?
And, really, there is no technical excuse. It is perfectly possible to do this, and lots of big companies do this (or hire security companies to do this for them). Mention their name on some deep web hacking forum, a pastebin, or inside Github code, and somewhere an alarm goes off.
Github could (and should) warn if a user uploads loads of PII-like data. For the cost of running a search server and a few moderators. "Are you sure you want to upload your AWS credentials in a public repository?".
Github is somewhere halfway between moderated and a content platform. They already have a history of taking down repositories if they link to PII data (or infringe copyright, or damage U.S. national security): http://web.archive.org/web/20180619172528/https://github.com... so not acting on this specific repo with SSN numbers could be seen as a poor/shoddy job on their part. Github is certainly in the dominant position to mitigate spread of PII data, so they should have their stuff in order.
> We demonstrate CryptoNets on the MNIST optical character recognition tasks. CryptoNets achieve 99% accuracy and can make more than 51000 predictions per hour on a single PC. Therefore, they allow high throughput, accurate, and private predictions.
Machine Learning. Apply neural networks to encrypted data and avoid privacy issues.
1943 First mathematical neural network model
1958 Learning neural network classifies objects in spy plane photos
1965 Deep learning with multi-layer perceptrons
2010 ImageNet error rate 28%
2011 ImageNet error rate 25%
2012 ImageNet error rate 16%
2013 ImageNet error rate 11%
2017 ImageNet error rate 3%
2019 Pre-AGIA good example of this is "move 37" from AlphaGo. This move surprised everyone, including the creators, who were not skilled enough in Go to hardcode it: https://www.youtube.com/watch?v=HT-UZkiOLv8
> Towards Biologically Plausible Deep Learning
> Neuroscientists have long criticised deep learning algorithms as incompatible with current knowledge of neurobiology. We explore more biologically plausible versions of deep representation learning, focusing here mostly on unsupervised learning but developing a learning mechanism that could account for supervised, unsupervised and reinforcement learning. The starting point is that the basic learning rule believed to govern synaptic weight updates (Spike-Timing-Dependent Plasticity) arises out of a simple update rule that makes a lot of sense from a machine learning point of view and can be interpreted as gradient descent on some objective function so long as the neuronal dynamics push firing rates towards better values of the objective function (be it supervised, unsupervised, or reward-driven). The second main idea is that this corresponds to a form of the variational EM algorithm, i.e., with approximate rather than exact posteriors, implemented by neural dynamics. Another contribution of this paper is that the gradients required for updating the hidden states in the above variational interpretation can be estimated using an approximation that only requires propagating activations forward and backward, with pairs of layers learning to form a denoising auto-encoder. Finally, we extend the theory about the probabilistic interpretation of auto-encoders to justify improved sampling schemes based on the generative interpretation of denoising auto-encoders, and we validate all these ideas on generative learning tasks.
Any AI curriculum worth its salt includes the many scientific and philosophical views on intelligence. It is not all alchemy, though the field is in a renewal phase (with horribly hyped nomenclature such as "pre-AGI", and the most impressive implementations coming from industry and government, not academia).
And eventhough the atom bomb was based on science too, there is this anecdote from Hamming:
> Shortly before the first field test (you realize that no small scale experiment can be done—either you have a critical mass or you do not), a man asked me to check some arithmetic he had done, and I agreed, thinking to fob it off on some subordinate. When I asked what it was, he said, "It is the probability that the test bomb will ignite the whole atmosphere." I decided I would check it myself! The next day when he came for the answers I remarked to him, "The arithmetic was apparently correct but I do not know about the formulas for the capture cross sections for oxygen and nitrogen—after all, there could be no experiments at the needed energy levels." He replied, like a physicist talking to a mathematician, that he wanted me to check the arithmetic not the physics, and left. I said to myself, "What have you done, Hamming, you are involved in risking all of life that is known in the Universe, and you do not know much of an essential part?" I was pacing up and down the corridor when a friend asked me what was bothering me. I told him. His reply was, "Never mind, Hamming, no one will ever blame you."
Our study of (automated) intelligence is based on science too.
> A computer ... will never be able to think by itself.
Turing wrote an entire paper about this (Computing Machinery and Intelligence), where he rephrases your statement (because he finds it to be meaningless) and devises a test to answer it. He also directly attacks your phrasing of "but it will never":
> I believe they are mostly founded on the principle of scientific induction. A man has seen thousands of machines in his lifetime. From what he sees of them he draws a number of general conclusions. They are ugly, each is designed for a very limited purpose, when required for a minutely different purpose they are useless, the variety of behaviour of any one of them is very small, etc., etc. Naturally he concludes that these are necessary properties of machines in general.
> A better variant of the objection says that a machine can never "take us by surprise." This statement is a more direct challenge and can be met directly. Machines take me by surprise with great frequency. This is largely because I do not do sufficient calculation to decide what to expect them to do, or rather because, although I do a calculation, I do it in a hurried, slipshod fashion, taking risks.
You drew the wrong conclusion about something you don't know a lot about and doubled down. Good luck with that and I forgive you.
The article cites Turner, who a few years back wrote this:
> The non-financial data was found predictive in all three outcomes examined when no other ‘traditional’ credit information was used, strongly suggesting that alternative data would be useful to lenders in underwriting the so-called ‘no-file’ or ‘no-score’ consumer who have little or no payment/credit information available.
I think his quote about "laughing out" is taken out of context. No predictive modeler will throw away informative features, because she can not distinguish noise from signal after 26 variables. That's 10 to 1% of a modern credit scoring model. It may be the perspective of a regulator though (they start drowning in noise after reviewing 100+ variables).
Yes, all data that is legal to use and predictive, will get used, if not by you, then by your competitor.
And informative variables that can not be used in the decision to give a loan, are used internally to predict if the loan will be paid back. There is more to credit scoring than the initial yes-no.
I like to point out that sources of unfairness are possible, even with 100% decorrelated non-protective variables. For instance, the data collection and labeling of the data may be biased (You label re-admission to jail with current criminals in jail, and just overfitted to the war on drugs).
Or you make a biased decision based on the output of a fair model. Things like not taking into account sample size / uncertainty.
Also, breaking the law is at odds with unethical behavior. Unethical behavior is not necessarily breaking the law, but it is still nasty. For instance, from your Facebook likes (not a protected variable) I could deduce all of race, sex, origin, religion. Not against the law, but still discrimination.
Sure it is. Even religion correlates with loan risk. (Could be a proxy for social status and people from your tribe helping you out when you can't pay back). I could probably get a predictive model better than random guessing by mining your HackerNews comments or Facebook likes.
> I'm sure it'd be used somehow by modern financial institutions if it were predictive.
It is used. All data that is even remotely informative is used. To the fullest extend made possible by jurisdiction/ anti-discrimination laws.
Two ways where these systems may give out more loans than is strictly profitable, and they are both investments:
- Fairness. If you have a variable race and a zip code, you could account for discrimination via redundant encodings, while still using the feature for the optimal trade-off between a fairness criteria and model performance.
- Exploration. Concept drift (the correlational and causal meaning of variables shifts over time) can introduce wrong predictions. If all you have is few samples from a zip code, the model will always be uncertain. You can counter this by exploration and active learning: gather samples, not because this makes the model max-profit, but gather samples, to better learn how to predict these samples in the future.
But yes, giving out too many loans to minorities, may very well lead to further crisis and defaults, and tainting the credit scores of people will low access to finances even further. A bit like how well-meaning people donate money and food to Africa during Christmas time, then a few months later when donations subside, there are increases in famines. There is such a thing as "being too good".
There is no direct inciting, and the indirect inciting claims are half-baked. There is a class action lawsuit office looking for a pay day, so they benefit by keeping Jones in the news.
All people factually saying that Alex Jones incited harassment against the parents, are parroting a non-concluded court case PR campaign ("OJ did it!"), and much like these behemoth companies, are playing judge, jury, and executioner. They most likely haven't seen a clip of Alex Jones inciting, or when they have, it is taken grossly out of context.
Remember, when CNN was targeting Jones for weeks, going through his hours of video, and noting anything they found controversial. "Alex Jones is transphobic, which is against Youtube TOS, but Youtube does nothing". Alex Jones was talking about "public library drag queen reading hour". He said these drag queens looked demonic and that it is not normal to normalize this for little kids. Now, just recently, someone ("a Trump supporter") showed up with a gun to these reading hours. Two possible conclusions: "Alex Jones is a transphobe and incited his followers to violently get rid of them". "Alex Jones practices unpopular, but legal, free speech, and they try to punish him for what any of his million followers might do that is against the law". I am leaning heavily towards the latter and it is a downright shame that the first conclusion is drawn by many, without doing any deeper research.
Imagine what you could do if you knew that many people would go for the first conclusion, regardless of the facts?
- GRU has been stepping up its attacks on the West (anti-doping, chemical weapons, MH17 investigation)
- US, Britain, and Israel attacked Iran with Stuxnet
- Iran retaliated with cyber attacks on banks and electrical grids
- Iran downed a drone
- Number stations go crazy
- US almost goes for military strike, but Trump calls it off and orders US Cyber Command to cyber-attack Iran, with focus on missile systems and spy networks.
- BGP route leak (Allegheny Technologies Inc. caused the leak, and was a target for Titan Rain around 2014).
- Dutch payment - and emergency response network goes down.
Edit: According to a KPN spokesperson: "We have an indication of the cause, but a lot remains unclear. Anyway, this was not the result of a hack."
> Want to know all the code names for America's massive intelligence gathering programs? Just browse through the "intelligence analysts" who post their resumes on the public career networking site LinkedIn.
Made me paranoid enough to stop accepting invites (~1000+ pending) and drop all usage of LinkedIn. Very discomforting to know you are a target. LinkedIn will have to visibly move on this for me to change my outlook and regain trust.