The Microsoft logo and the Google logo are very similar in style: http://i.imgur.com/s7kR6jJ.png
The blocks in the middle are the colors from both the logos.
1,012 karma · joined December 3, 2014
The Microsoft logo and the Google logo are very similar in style: http://i.imgur.com/s7kR6jJ.png
The blocks in the middle are the colors from both the logos.
I myself do not like this redesign, but I also did not like the old one. I am basically very averse to changes like this (probably a mental issue). I also do not like how Chrome hijacks the window style, or how Android automatically updates to a new Youtube app with a drastically new UI. I am probably a minority though, and still capable to use custom styles to change the logo back to the old "Backrub" logo, but I wonder if there are more who see these major redesigns as an invasion of sorts into their daily routines. Just as you created a mental map of the application, they change the location of the buttons, or show a different favicon.
If it is not too much to ask (or risky to maintain), think of keeping an "old.*.com" domain. There are users who will gladly make use of that. Being able to use your own design, or turn off "with this new logo we save 9999 bytes on mobile" distractions is also UX. And it would be good UX for me.
With this new logo, for me, the counter now starts at zero for Google. A clean slate. For all I care they may be the Facebook they so desperately desire to mimic.
As memes house in agents with an energy budget, I think that shorter simpler memes have more chance to take hold and reproduce ("Make something people want").
Words in a sentence are like models in an ensemble. Simple words are more general and have a high bias and low variance ("Make stuff users want"). Highly complex words and sentences have a lower bias, but a higher variance. You need to average a lot of them to get a clear picture. That's why the sentences in scientific papers are usually so long, they need to gradually cancel out the noise.
> There is certainly some cultural variance that can be orthogonal to efficiency.
Yes, agreed! Though same with memes, certain words or symbols without any redundancy may have cultural value. You may gain energy by speaking a certain language to a certain degree of sophistication. You may have to invest energy to gain access to the information contained in symbols (or have agents "unzip" these for themselves).
I think these old German and Japanese languages may have been hard to understand for outsiders, but were used with high sophistication (you have to invest energy to access this information) among insiders. For instance the Japanese pillow words / Makurakotoba or German words for hard to translate concepts like "Weltschmerz", "Kummerspeck" and "Torschlusspanik". All short, useful words for communicating complex rich concepts, provided the agent knows the meaning of these words.
Words like "circumvent" and "environment" are close in regards to complexity. Words like "us" and "me" are close in regards to complexity.
The counting argument tells us that most strings are not compressible. It is then a wonderful feature of sensor data, natural language, DNA and computer code that it can be compressed quite a bit. This means there is a certain order in the language that compressors can use to keep the file size smaller.
There is a cognitive economy trade-off between the energy needed to keep a system running and increased complexity. Less complex language helps us save energy. We use short words for concepts that we use often. Very complex concepts and words like "disambiguation" may be described with shorter simpler words to someone who has not stored that word and general accepted meaning yet.
In this complexity view languages evolve to use as little energy/computational complexity to convey as much information as possible. The results found in this article can also be explained using this view. Parsing a sentence like "Throw the trash out" requires you to store in working memory the word "throw" 'till you get to the word "out" for the full concept "to throw out". Until you get to the word "out", the "throw" remains in a superstate (could become "throw in", "throw on" etc.). You need both words to form a mental picture of someone throwing out the trash. This requires more computational energy to the listener, and is hence ineffective. If you want your message to be heard, you have to communicate in clear simple-energy sentences. So using simpler less computationally intensive sentences benefits both the speaker and the listener.
This would readily explain why natural languages beat the random benchmark. Randomness has far less structure to use for compression by an intelligent agent. Randomness is not optimized communication, since it is more unpredictable.
In short: Simplicity and conveying information with little energy is a fitness factor that natural selection optimizes for. This is universal to all natural language speaking agents with a limited energy budget.
Because they could not translate their sense of entitlement to actual results.
> pure statisticians often scoff at the hype surrounding the rise of data scientists in the industry
> some statisticians simply have no interest in carrying out scientific methods for business-oriented data science
Statisticians are often too careful. They let tests decide if they should continue on a certain path. Machine learning researchers run blindfolded and trust cross-validation. The latter, though reckless, gets more impressive results.
You can perfectly be a data scientist coming from a statistics or physics background. Adapt to it and use your knowledge to your advantage. You can't keep calling yourself a statistician and own data science at the same time. Start automating yourselves, like the rest of us are.
With applied machine learning it is certainly possible to quickly get a working knowledge without too much reliance on statistics or difficult theory. You can compare this a bit with using a sorting function without knowing exactly how it works (but you know how fast it is and when to use it).
If you have an engineering background, take a look at the wide array of high-quality ML code and tools. Study trendy and powerful tools like XGBoost.
http://www.hutter1.net/ai/pfastprg.htm (The Fastest and Shortest Algorithm for All Well-Defined Problems)
Even if they have a terrible track record, this code implies nothing that they are currently accused of. It is not like the code is heavily obfuscated or missing other parts. Let's not forget these leaks were "planted" too and we have no idea about its integrity.
I think it is rather cheeky (and dangerous) to accuse companies of doing such horrible stuff, when all we have is a few leaks (without context) to go on.
See for instance: https://github.com/hackedteam/rcs-common/blob/master/lib/rcs... which generates test keystrokes, test users and test programs for chat. The file names of the "planted" evidence is far too obvious, and they generate a hash for it. In short there is nothing in this code that implies this is used to plant evidence.
Furthermore, I do not understand why this is called an "insider look", when it is about someone reading a blog from 2012 and basically giving a recap.
Finally, I think it is time to move past this trope of "Netflix wasted 1 million dollars on a solution they did not even use". Probably the entire top 20 of the Netflix competition had a good enough score to build upon -- no need to focus only on the solution from the (rather lucky) winners. We are six years after this competition and the community is still talking about it. Now how is that for marketing?
There is no mention of Factorization Machines ( http://www.libfm.org/ ). A technique spawned with this competition which revolutionized recommendation engines and is definitely in use at Netflix now. Instead they keep harping on 'Collaborative Filtering' when that is a very basic technique (obviously supported by Matlab).
I miss a mention of power tools (Vowpal Wabbit has fast production-ready scalable matrix factorization which can fit on a 1 million dataset on a laptop: https://github.com/JohnLangford/vowpal_wabbit/tree/master/de... ) and of customers actually running Matlab based recommendation engines on scale.
See here for an MIT experiment in blurry text transcription: http://groups.csail.mit.edu/uid/deneme/?p=329 for the unexpected accuracy resulting from crowd-sourcing.
The cognitive benefits-schtick is also used to justify the religious doctrination: "You have passed and future lives through reincarnation. Buddha existed and we have his teachings. Our goal is to become deathless.". It's scientific right? The cognitive benefits-research focuses on hospital patients and stress/pain relieve. They do not send these patients to a 10-day retreat, where they are not allowed to talk, must surrender to a master, do extreme meditation techniques for hours on end, till they self-operated on their psyche enough it's broken and they now have to repair it.
What salvation is for Christians, is enlightenment for Vipassana.
There are frauds out there who use these techniques to ensnare students, and there are students out there who ensnare themselves through being young, naive and gullible. Since the experiences are so dramatic, you get a host of uncritical people who proselytize Vipassana. Kinda like people who buy expensive Apple products will be lauding Apple, since else their investment was bad, and no one wants to admit to that. The people for who it didn't work remain quiet, especially when the master told them it was their own fault.
Vipassana for prolonged times can do much more than reading a book or eating breakfast. What if the author had gotten a psychotic episode during his tripping balls? Would that be dismissed with another fancy foreign term or garbled psychological babble? Would a master be able to spot deteriorating mental health in their patients? People report disassociation, hallucinations and hearing voices. To a qualified mental health professional that would not be scientific evidence of cognitive benefits, that would be a manifestation of latent schizophrenia. Then there are the documented suicides and self-harm... but then again, people have probably died from eating bad breakfast too.
Meditation-induced psychosis: http://www.karger.com/Article/Abstract/108125
Panic attacks and depressed episodes: http://zensydney.com/Mental-Health-and-Intensive-Meditation-...
Psychiatric complications of meditation practice: http://www.atpweb.org/jtparchive/trps-13-81-02-137.pdf
Terrible and Traumatic Experience at Goenka Retreat: http://downthecrookedpath-meditation-gurus.blogspot.de/2012/...
The Potential Downside To Vipassana: http://livingvipassana.blogspot.de/2007/06/potential-downsid...
Edit: I am not comparing meditation to heroin. I am showing that the question 'Does it cause harm to you if other people give it a try?' is a trick question, a debating technique. Answering 'no' does not invalidate the statement 'this reads like a scientology ad'. But even if I did compare meditation to heroin, so what? The article compared meditation to psychedelics.
There is a limit on the number of submissions. Usually 5 or less a day. Also precision of scores is visibly limited, but for final ranking all decimals count.
This forum post discusses competition variance and poses a metric to quantify "leaderboard shake-up": https://www.kaggle.com/c/liberty-mutual-fire-peril/forums/t/...
The Public Leaderboards are very helpful though! When you have setup a solid local cross-validation pipeline, and the public leaderboard agrees with your local evaluation, then you can try a lot more algorithms and parameters, without using any submission. Especially when working in teams this is important as you may have only 1 submission every 2 days.
Also, the more advanced Kagglers can use leaderboard feedback to increase model accuracy: Cluster the data sets with objective measures. Apply a modifier (restaurants from this region get 0.95 x previous prediction) and look at the result. If the split between public and private is random, and your clustering is objective, then an improvement on public leaderboard should reflect in private leaderboard.
The public leaderboard gives some feedback on your model performance. But when a human is in the feedback loop, then there is a risk of overfitting. Overfitting can be explained basically as: "memorizing the data" or "learning from noise, not signal".
When you overfit, you do well in cross-validation and may do well on the public leaderboard, but your predictions do not generalize well to new data.
What this team did was to submit a lot of predictions and only take the predictions that improved public leaderboard score. Then they'd add slightly random noise and try to submit again. They repeated this until they ranked nr. 1 and left quite a few other competitors scratching their heads: How did they do this? Did they find a perfect ML algorithm? Is there data leakage? Are they cheating?
When the private leaderboard was revealed, this team dropped around 2000 spots. Their good performance on the public leaderboard was purely artificial. The contest was valid and well-organized (this could happen on any Kaggle competition with little data). They did not receive a prize.
There are other benefits to ranking well on the public leaderboard: It helps with teaming up with other high-ranking competitors and you can market yourself (recruiters are pretty interested in the top 10, eventhough the competition has not ended yet.)
- Find and read some front-end designer blogs.
- Take a good (CSS) framework like HTML5 boilerplate and dissect its code.
- Remake your last favorite designed site without looking at their code. Then compare afterwards.
- Start creating your own framework for rapid prototyping. Add layout-rules to common UI elements: breadcrumbs, pane lists, buttons, forms etc.
If I did not get downvoted for my very first flame-y, half-assed or dumb replies, then I would have continued making them.
Wickedness and cruelty is forcing low standards upon a community and sucking away all the fun.
With downvoting comes a little responsibility: to not only reward good posts, but lower the visibility of posts that do not contribute to the quality or enjoy-ability of HN. That is why you need 500 upvotes to get downvote functionality.
A plain HTML site is accessible, and will be accessible in a 1000 years. A site depending on (external) JavaScript sources will force John Titor to travel back in time to find version 1.x of jQuery. Starting with JavaScript abandons principles of progressive enhancement. Sometimes there is not even a fallback/graceful degradation, reminding me of these 2001-era: "Best viewed at 800x600 resolution in Netscape"-sites.
Google holds enormous clout among SEO's. Google says they will factor in site-speed and a large fraction of the web will become faster. Google can say more sternly that using JavaScript can have ugly consequences for user experience and accessibility, but they are swimming upstream: The web seems to be moving on to fancy new technologies regardless of what their SEO says.
Not much good comes from HTML5 JavaScript fans forcing your hand. Tor enabled JavaScript, because too much of the web would break without it, leading to a poor user experience. This led to a huge security gaffe, which I fully blame on webdevelopers eschewing basic principles, to get that slideshow running.
> On April 8, the SendGrid account of a Bitcoin-related customer was compromised
If I can gather this right: SendGrid was fully hacked for 3 months on end. At least that is what they were able to recover from forensics, it may have been longer.
This sounds illogical:
> We have not found any forensic evidence that customer lists or customer contact information was stolen. However, as a precautionary measure, we are implementing a system-wide password reset.
How would a password reset help combat information that was stolen before the reset? The password reset is because the systems accessed contained password hashes. Also it may be to upgrade the hashing mechanism to be more secure than "salted and iteratively hashed".
Sendgrid's privacy policy is cookie cutter, but it contains this:
> For example, our policy is that only those individuals who need your personally identifiable information to perform a specific job are granted access to that personally identifiable information.
Apparently the employee that was hacked needed access to the data of all his colleagues and all users.
> Upon discovery, we took immediate actions to block unauthorized access and deployed additional processes and controls to better protect our customers, our employees, and our platform.
Then the Privacy Policy again:
> We will use at least industry standard security measures on the Site to protect the loss, misuse and alteration of the information under our control. While there is no such thing as "perfect security" on the Internet, we will take all reasonable steps to insure the safety of your personal information.
So apparently there were still some reasonable steps left to take, which were forced by this hack, not by 'industry standard security measures'.
> Two-Factor Authentication: We encourage all of our customers to enable two-factor authentication, which can effectively prevent unauthorized logins.
I think you can better encourage (or force) your employees to enable this, so you can prevent unauthorized logins into superuser accounts.
From the Privacy Policy you'd expect they already did this:
> Likewise, all employees and contractors are kept up-to-date on our security and privacy practices.
Then the unspecified hashing mechanism. Should you worry about the chance of account compromise again?
> salts and iteratively hashes passwords
It would be a breath of fresh air if these companies would just say 'We use bcrypt' in their privacy policy.
> Our Ongoing Commitment to Security
Your 3 month struggle with hackers. Also, three reasonable steps follow that could have been taken before this hack, like your Privacy Policy promised us.
> NOTE: We require passwords to be a minimum of 8 alpha-numeric characters. Make sure any new passwords you set conform to this requirement.
Before or after this hack? Why should the customer make sure his password conforms to this requirement? Is it even possible to set a shorter password?
> Security update: Please reset your SendGrid account passwords today. Beginning today, and in line with standard practice, we are requesting that all of our customers reset their passwords to all of their SendGrid account access points.
Why not force this? Asking nicely? Standard practice would be to force this upon next log-in and temp disable accounts that have not changed their password yet.
2: Learn how to manipulate Numpy arrays ( http://www.engr.ucsb.edu/~shell/che210d/numpy.pdf ) and how to read and manipulate data with Pandas ( https://www.youtube.com/watch?v=p8hle-ni-DM ).
3: Do the Kaggle Titanic survival prediction challenge with Random Forests. ( https://www.kaggle.com/c/titanic/details/getting-started-wit... )
4: Study Scikit-learn documentation ( http://scikit-learn.org/stable/documentation.html ). Run a few examples. Change RandomForestClassifier into SGDClassifier and play with the results. Scale the data to make it perform better. Combine a RF model and a SGD model through averaging and try to improve the benchmark score.
5: Study the ensemble module of Scikit-learn. Try the examples on the wiki of XGBoost ( https://github.com/dmlc/xgboost/tree/master/demo/binary_clas... ) and Vowpal Wabbit ( http://zinkov.com/posts/2013-08-13-vowpal-tutorial/ ). Practically you want to get to a stage of: Getting the data transformed to be accepted by the algo, a form of evaluation, and then getting the predictions back out in a sensible form.
Then next week start competing on Kaggle and form a team to join up with people at your level. You will learn a lot that way and start to open up the black box.
I found these series very accessible: http://blog.kaggle.com/2015/04/22/scikit-learn-video-3-machi...
Kaggle also recently released a feature to run machine learning scripts in your browser. You could check those out and check out Python, R, common pipelines and even the more advanced neural nets: https://www.kaggle.com/users/9028/danb/digit-recognizer/big-... .
Furthermore they already published their search engine algorithm: http://www.google.com/technology/pigeonrank.html
If you feel too good, or too senior for a programming methodology like Agile, then you are bound to not like it. Is that a fault with Agile, or a fault with your attitude?
A cowboy coder is not micro-managed, does not seem to need atomized tasks, does not need a direct line with the customer and decides on the complexity of the problem and the solution all on his own. They are also terrible to manage and hit-and-miss when it comes to actually shipping something with business value.
Bluntly spoken, Agile is there for the project managers, not the project developers. Agile should be implemented for as long as productivity increases, customer feedback loops create desirable features, and iteration cycles become tighter and shorter. If you know a way to increase those stats without Agile, then write your own methodology (http://programming-motherfucker.com/ is taken) and join management. If you don't care about those stats, and feel too good/senior for Agile, then start your own company and proof it to yourself. You don't start to measure a project's success by first counting the number of developer complaints. You could probably keep the complaints at zero, and never ship anything of impact.