HNHacker News
TopNewBestAskShowJobs

blauwbilgorgel

1,636 karma · joined February 2, 2011

submissionscomments
blauwbilgorgel··on Ask HN: Anyone try semaglutide / Wegovy for weight loss?
Every drug has its sets of mild or severe side effects.

But few drugs, including Wegovy, have a Black Box Warning. Even fewer where the increased risks in humans are unknown at the time of FDA approval, and only backed by rodent studies.

https://www.ncbi.nlm.nih.gov/books/NBK538521/

blauwbilgorgel··on Ask HN: Anyone try semaglutide / Wegovy for weight loss?
It is hard in the beginning, but it gets much easier after sticking with it for a few weeks. Your body and hunger adapts.
blauwbilgorgel··on Ask HN: Anyone try semaglutide / Wegovy for weight loss?
> A recent review of the evidence suggests that this type of diet may help people with type 2 diabetes safely reduce or even remove their need for medication.

> However, people should seek the advice of a diabetes professional before embarking on such a diet.

https://www.medicalnewstoday.com/articles/can-intermittent-f...

blauwbilgorgel··on Ask HN: Anyone try semaglutide / Wegovy for weight loss?
Indication and Important Safety Information

What is the most important information I should know about Wegovy™?

Wegovy™ may cause serious side effects, including: Possible thyroid tumors, including cancer. Tell your healthcare provider if you get a lump or swelling in your neck, hoarseness, trouble swallowing, or shortness of breath. These may be symptoms of thyroid cancer. In studies with rodents, Wegovy™ and medicines that work like Wegovy™ caused thyroid tumors, including thyroid cancer. It is not known if Wegovy™ will cause thyroid tumors or a type of thyroid cancer called medullary thyroid carcinoma (MTC) in people.

https://www.wegovy.com/FAQs/frequently-asked-questions.html

edit: For a safer way to lose weight, including other benefits in regards to reducing cancer risks, look into intermittent fasting. Success with your weight loss, great step into becoming more healthy!

blauwbilgorgel··on Doxxing defense: Remove your personal info from data brokers
Just because a data aggregation site does not show your data on the front-end, does not mean they deleted it from the back-end. So now you can charge 50$ for people to search in the "special" data pile, where people took the effort to remove it from the front-end.

These data brokers crawl publicly available information. Telling them to remove your data, only slows down the doxxer, it does not stop them at all, since the data was already shared. It is not plugging the leak, it is mopping up some of the water. A false sense of security and a clear sign to the doxxer that you care about your anonymity (so more "lulz" to be had).

A proper doxxing is also much more than entering a name in some search engines. Especially hackers do not like to be doxxed. For internet civilians who already put this data out there (on social media) a simple data broker doxxing is a mere reminder that such data is public to everyone, not just friends.

Doxxing defense is guarding your anonymity online. Everywhere. Doxxing defense is knowing when to change persona's, and when to log off. That is: If you care about it at all. If you care about keeping your identity a secret online, see: https://www.youtube.com/watch?v=9XaYdCdwiWU (The Grugq - OPSEC: Because Jail is for wuftpd).

blauwbilgorgel··on How do you get to write so well in HN?
The guidelines say to write as if you were face to face to a person. You wouldn't probably mention those nasty things, or at least you try to hide it behind constructive criticism.

Take this one step further: For everything you write, imagine the utmost authority on that topic reading your post. A blunt example: A rant about Python syntax. Imagine Guido van Rossum reading that during his coffee break.

Do not post if you can not add anything to the discussion. Assume your debater is smarter than you and knows more on the subject. This is still HackerNews. Ask someone if they have won the Putnam prize, and you may be unpleasantly surprised. [1]

Read and practice: http://www.paulgraham.com/disagree.html

Use your spell-checker. Write shorter sentences to avoid grammatical errors. Nobody is too smart for short, simple sentences.

[1] https://news.ycombinator.com/item?id=35079.

blauwbilgorgel··on Machine-Learning Maestro Michael Jordan on the Delusions of Big Data and Others
> ... how exactly none of those are the ones with the necessary PhDs in statistics and algorithms to get anything of any value done.

I see it almost the other way around: Companies strictly demand PhD's for Big Data jobs and can't find this unicorn. Yet we live in a time where we don't need a PhD program to receive education from the likes of Ng, LeCun and Langford. We live in a time where curiosity and dedication can net you valuable results. Where CUDA-hackers can beat university teams. The entire field of big data visualization requires innate aptitude and creativity, not so much an expensive PhD program. I suspect Paul Graham, when solving his spam problem with ML, benefited more from his philosophy education than his computer science education.

Of course, having a PhD. still shows dedication and talent. But it is no guarantee for practical ML skills, it can even hamper research and results, when too much power is given to theory and reputation is at stake.

In my experience Machine Learning was locked up in academics, and even in academics it was subdivided. The idea that "you need to be an ML expert, before you can run an algo" is detrimental to the field, not helping so much in adopting a wider industry use of ML. Those ML experts set the academic benchmarks that amateurs were able to beat by trying out Random Forests and Gradient Boosting.

I predict that ML will become part of the IT-stack, as much as databases have. Nowadays, you do not need to be a certified DBA to set up a database. It is helpful and in some cases heavily advisable, but databases now see a much wider adoption by laypeople. This is starting to happen in ML. I think more hobbyists are right now toying with convolutional neural networks, than there are serious researchers in this area. These hobbyists can surely find and contribute valuable practical insights.

Tuning parameters is basically a gridsearch. You can bruteforce this. In goes some ranges of parameters, out come the best params found. Fairly easy to explain to a programmer.

Adapting existing algorithms is ML researcher territory. That is a few miles above the business people extracting valuable/actionable insight from (big or small or tedious) data. Also there is a wide range of big data engineers making it physically possible to have the "necessary" PhD's extract value from Big Data.

blauwbilgorgel··on Massive electrode array will do first large-scale recording of brain activity
This will be the first of this scale. On both those sites the most channels I could find was 64. This program is planning 10.000 channels for unprecedented resolution and scale.

Since 16 channels is enough to predict if a subject is looking at a face or not, it is exciting to research what this large-scale system is going to be capable of. Next to neuroscience, it could help with healthcare (research into dementia, epilepsy and schizophrenia).

blauwbilgorgel··on Google Apps Security Vulnerability Puts Organizations’ Data at Risk
I don't see a security vulnerability, but bad security practice.

Either they: Delete the account. All is well.

Either they: Take over the account. It is common sense then to change the phone number associated with the account. All is well.

You could solve this "bug" by reading the documentation and creating a better security protocol (which is currently putting your organizations' data at risk).

I clicked the title with just one thought it the back of my mind: "If this is an active serious vulnerability then why did OP not apply for the vulnerability program and have it fixed beforehand"?

My experience with the vulnerability team has been great (one honorable mention and one pay-out). If you did not get an honorable mention then it means the security team did not file a bug report. Your feedback could probably still be used to improve the UI.

As an aside: Hunting real security bugs on Google domains is insanely addictive (because they are so hard to find). Try to generate all their different error screens. Try to find the Google property running on aspx. To practice there is also https://google-gruyere.appspot.com/

blauwbilgorgel··on Dear Rupert
Did Google notify you of a link-based penalty? You should set up Google Webmaster Tools, if you haven't already. Then you can read notifications about suspicious links pointing to your site and disavow the ones that are spammy.

I am not so sure that you are under a link-based penalty. I know you did not ask for this, but I had a look at your website's link profile, source and index health.

1. SEO: The online web form market is hugely saturated. You probably can not compete with Wufoo on terms like "web form builder". Honestly ask yourself if you currently deserve a top 10 spot for this term. Are you a top 10 player in this field? Explore more specific and longtail keywords. Create better targeted pages and page titles. "Documentation - NicSoft Software" is a missed chance. Create more content on the blog (inbound marketing).

2. Links. You do not have enough natural links to beat competitors. They get linked from webdeveloper forums by real users of the software. You also should check out Google's stance on "Powered by"-links. If this turns in the majority of your backlinking profile, you get links from a lot of bad neighborhoods. These links may thus do more harm than good. It is not an editorial link, but probably in exchange for a free version of the product. Much safer to nofollow links created for profit or SEO/online marketing purposes.

3. Site HTML is not well-structured for information retrieval. For every page on your site, the first heading is "Home". Browse your site with styles disabled. Reorder repeating boilerplate code below the relevant page content. Specify a canonical or make sure only one version is served to visitors with redirects (both www- and non-www versions of the site return duplicate content).

4. Index health is poor. Robots.txt file is indexed. There are a few inactive subdomains, an unattended Wordpress and Drupal install in a subdirectory. Includes are indexed as separate pages "/inc/footer.php". Documentation (the content "meat" if the site) is off limits for bots. Over 90% of pages on the site are in the secondary index, which is not a good sign.

About 10 hits a day from Google is far too low for any commercial site to survive on. You could get more than 10 hits on a random wordlist.

Do not solely think about in-links. Are you even linking to reputable sources yourself? End-node sites are far less interesting for visitors than hub sites.

blauwbilgorgel··on Navdy
I am not so afraid of the loss of attention. I think this takes a little getting used to, like using a route planner.

I didn't see a comment yet about placing a loose object in front of your face. During a head-on collision you will eat that thing. Don't even put a box of matches on your dashboard. CD's become chainsaws. And this hunk of plastic?

Every year, loose objects inside cars during crashes cause hundreds of serious injuries and even deaths. In this paper, we describe findings from a study of 25 cars and drivers, examining the objects present in the car cabin, the reasons for them being there, and driver awareness of the potential dangers of these objects. With an average of 4.3 potentially dangerous loose objects in a car‟s cabin, our findings suggest that despite being generally aware of potential risks, considerations of convenience, easy access, and lack of in-the-moment awareness lead people to continue to place objects in dangerous locations in cars. Our study highlights opportunities for addressing this problem by tracking and reminding people about loose objects in cars.

http://www.star-uci.org/wp-content/uploads/2011/08/ubic489n-...

blauwbilgorgel··on Google Launches Cloud Platform for Startups
Google has another program for startups without funding or accelerator program access.

https://developers.google.com/startups/

I applied and received $500 in credit for free (also free access to online training and local events). To get 100k in credit would of course be nice, but I would have no way to spend that much in a year.

In the light of this two-tier startup program already existing a lot of comments in this thread become uninformed. Stop looking a gift horse in the mouth (unless it is a Trojan).

blauwbilgorgel··on A Revolutionary Technique That Changed Machine Vision
Do you think that computers are better at chess than humans? If yes, how does this relate to pattern recognition. If not, what makes someone or something better at chess, while still losing against a computer? Is that a beautiful move? Tactics? Irrational sacrifices to cause confusion?

Do you think that a machine's situational awareness can not achieve or surpass the level of a human? If not, what is holding the machines back?

Why do you think that instinct works better to create more rational, consistent and correct predictions? Are 100 security guards better than a single security guard at dealing with ambiguities? Do you think an algorithm to detect fights, drug dealers, and pickpockets from street cams can not exist? What if a NN could detect these cases faster and flag this to a human security guard for action/no-action.

blauwbilgorgel··on A Revolutionary Technique That Changed Machine Vision
Google trained a NN on unlabeled Youtube stills. It was able to detect/group/cluster pics of cats without ever seeing a label. This still needs supervision to teach the NN that whatever name it created for this cluster, us humans call this "cats".

If the error rate gets low enough, a NN could start labeling pics.

Finally, recent work has shown that running a dictionary through an image search engine can yield high quality labeled images automatically.

Aside: Thank you for contributing to sklearn. Really feel like I am standing on the shoulders of giants when I use that library.

blauwbilgorgel··on A Revolutionary Technique That Changed Machine Vision
It is just extrapolating the error rate reduction over the last few years. Spam filters have become better than moderators in labeling spam only in the last decade or so.

When computers first started to become faster than mathematicians this was really a breakthrough. The same is happening now with object and speech recognition.

The computer succesfully completes a task. That it is not how humans intuitively approach these same tasks is irrelevant for this accomplishment. What if the results were only half as good, but the system behaved more like humans, who does this satisfy?

The state-of-the-art is capable of detecting far more than 1000 objects, does not need labeled data, is robust to changes in light and does not care about the camera used. No preprocessing the data needed, features are automatically generated (preprocessing the target labels is a bit silly BTW).

So yes, in the very near future, algorithms will be better security guards than well... security guards.

blauwbilgorgel··on Tabnabbing: A New Type of Phishing Attack (2010)
I think you are right:

"Forbid META redirections inside <noscript> elements"

but then I immediately wondered, what about META redirections outside <noscript> elements? I tested this with a fresh install of Firefox and latest NoScript, and those still work. Also: To forbid meta redirections inside noscript elements you have to toggle an option, it's not standard for non-trusted sites.

blauwbilgorgel··on Tabnabbing: A New Type of Phishing Attack (2010)
The proof of concept is foiled. But how about something similar like:

  <noscript>
    <meta http-equiv="refresh" content="600" 
    url="phish.php"> 
  </noscript>
blauwbilgorgel··on Tabnabbing: A New Type of Phishing Attack (2010)
It is ok. Thank you for submitting an interesting article.

This attack vector is very relevant to this day. It works remarkably well on mobile phones, especially with the new trend of hiding the URLs to save screen estate. Together with throwing up a fake website, you could also try throwing up a fake address bar.

Also with mobile phones you can emulate a native application. This will show when the phone is re-activated with the browser app on.

A redirect to something like a Googleusercontent domain would be tricky, even with a visible URL. The author could also mine the referrer: visitor from HackerNews? Phish the log-in form with a bad gateway message.

blauwbilgorgel··on I disagree with Turing and Kahneman regarding the strength of statistical evidence
Turing believed in fairy tales. Gödel believed in ghosts.

This was also at the start of the cold war, where the US suspected that the USSR was funding millions to ESP research.

One of the most publicized projects was the Stargate Project: Even though a statistically significant effect has been observed in the laboratory, it remains unclear whether the existence of a paranormal phenomenon, remote viewing, has been demonstrated.

I think Turing was closer to the machine learning camp than the statistics camp. On that note I'll quote a competitor in the MLSP 2014 Schizophrenia Detection Challenge:

To the people that are going to write papers for this one...

What really strikes me is the fact that stats fail really hard in this problem. I have a couple of 2-variable combinations that score around 0.87 on training set with logistic regression and a couple of 3-variable combinations with training AUC 0.9 (ish). All results were "statistically significant" at 0.001 (not even 0.01) . I have tried the same selections with SAS, SPSS, R and scikit (with regularization) . All results are consistent (and similar) with all packages, yet again they scored around 0.5 (random) in public and private leaderboard. This makes me think about all the PhDs' thesis and medical science papers I've seen being carried out on mickey mouse sets , claiming statistical significance gives credibility to their findings ... Is machine learning more reliable than stats? I say, if you can't predict it consistently on a hold out set, then you got nothing whatever the t,F,Chi-sq distributions say. KazAnova - https://www.kaggle.com/users/111640/kazanova

blauwbilgorgel··on Failing the startup game at Unbabel
Thank you. Just a small piece of feedback: I clicked the title expecting a story on Unbabel failing the startup game, ie: I thought they'd gone bust and this would be a post-op.
blauwbilgorgel··on Failing the startup game at Unbabel
I read a lot of burning bridges. Quite unnecessary.

Author calls his former job a mind-numbingly boring Java consulting gig.

Then calls his new job at Unbabel insane and abusive.

As the author does not explain how working for 30 days at Unbabel was abusive, this only reflects poorly on the author. He seems impossible to satisfy.

That code at start-ups is messy is the norm, not the exception. Highlighting this as: "a tangled mess of mindless duplication, half-implemented features and misleading comments" again reflect poorly only on the author. What did he expect as an experienced coder? Why air this "dirty" laundry? How do the people (your former and future colleagues) writing that code feel now?

Then continues to describe the horrible experience: "The team lead was the only one who knew anything about the system". Then seems surprised at that Friday afternoon meeting with the founders.

This may negatively influence hiring practices of YC companies. Want to avoid such culture fit disasters? Do no hire anyone over 30. Do not want fire and brimstone blog posts when you fire someone? Hire someone local or remotely outsource. Taking a chance on someone works both ways.

Yes, I imagine it sucks for the author and I wish such an experience on no one. But I also wouldn't want to be the startup to read this on the frontpage of HN, see yourself be misrepresented and having to consult with legal before you can even think of replying. A bad hire is unfortunate for the employee, but very expensive for the startup too: Too many of these and the company will go down. EU labour laws are far more protective of employees than US labour laws. If this was within the law, it may not have been too nice, but remember: This is (a) serious business.

There is probably a grain of truth in this story, and perhaps Unbabel made a poor business decision, but these stories can never be taken at face value, they are closer to hit pieces. Drama-bait.

blauwbilgorgel··on Machine learning isn't Kaggle competitions
Why Machine Learning is Kaggle competitions.

  I used an out-of-the-box algorithm, messed around a bit, 
  and definitely did not make the leaderboard.
Because that is not Kaggle competitions. Nearly everyone on the leaderboard is proficient in data analysis and machine learning. They all tried that out-of-the-box algorithm for their attempt. But they did not give up so easily.

  Understand the business problem
  If you want to predict flight arrival times, what are 
  you really trying to do?
This is not different from Kaggle competitions, this is a tip for performing better in Kaggle competitions. See also the GE Flight Quest: https://www.gequest.com/c/flight Those winners used industry-standard machine learning, optimization techniques, but also creative insights and hunches, like tweaking the target labels:

"A next step is to ask, “What should I actually be predicting?”. This is an important step that is often missed by many – they just throw the raw dependent variable into their favorite algorithm and hope for the best. But sometimes you want to create a derived dependent variable. I’ll use the GE Flight Quest as an example: you don't want to predict the actual time the airplane will land; you want to predict the length of the flight; and maybe the best way to do that is to use the ratio of how long the flight actually was to how long it was originally estimated to be and then multiply that times the original estimate." - Steve Donoho - http://blog.kaggle.com/2014/08/01/learning-from-the-best/

Furthermore, it is entirely clear to everyone that data science in a business setting and in a competitive sport setting is different. To say they are equal, would be to say something like: paintball is equal to being in the military. But to say that Kaggle is not machine learning is to say: paintball requires no marksmanship.

There are some very messy, unwieldy datasets on Kaggle right now. For example the Seizure Detection challenge has many GBs of raw sensor data, from just a few patients. This would require a competitor to clean, understand problem domain, understand evaluation metrics, measure cross validation and put your model in production on your laptop in the evening hours.

The author of that blogpost is invited to team up, with me or others. Let's see if we can use machine learning to improve some pressing issues. I'd also love it if Stripe can host a contest on Kaggle.

blauwbilgorgel··on How Google Works
Perhaps: http://www.google.com/intl/en/insidesearch/howsearchworks/
blauwbilgorgel··on Lorem Ipsum: Of Good and Evil, Google and China
I remembered a discussion about this before, once in 2010 and once in 2013. Already in 2010 Lorum Ipsum was translated to random words that are very prevalent on the internet: "hello world", "learn more" and "free on": http://www.xefer.com/2010/10/lorem-ipsum

Therefor I don't think that Google Translate is used by spies to communicate plans about China. Even if all translated words were insidious, then still Occam's Razor tells us it is unlikely that a public translating service is used as a modern-day number station.

The 2013 discussion had translations like: "Cisco Security" and "Corporate Japan": http://googlesystem.blogspot.com/2013/06/lorem-ipsum-google-...

That's statistical machine translation for you.

blauwbilgorgel··on DataRobot raises $21M Series A
Data scientist is the sexy term for statistician. It's almost as overused and undefined as "big data".

You could also call an all-round data scientist a unicorn, since they do not seem to exist. Data science requires a combination of maths/stats, computer science and domain expertise. A statistician who does not know how to run a regression on her dataset, or works on clicklog data without understanding CTR and ROI in a business context, is not really a data scientist.

Here is a Venn diagram: http://www.itworld.com/sites/default/files/data-science-venn...

To DataRobot: Grats on the funding and attracting top talent! Awaiting your product when you release in the fall.

blauwbilgorgel··on Hello, this is an extortion email
I don't know if you are in on the joke, but this is how XRumer actually works. It has bots or other users provide links to their auto-posted questions, in an effort to avoid detection.
blauwbilgorgel··on Hello, this is an extortion email
Spammers noticed that their own sites got nuked from the results when they used an automated tool (XRumer) to point thousands of low-quality links to these sites.

So with a paid product that was now useless (or less useful) for linkspam, they now turn it around and use it for blackmail: Pay us or we point all these links to your site.

There used to be a time when the majority of these links were simply ignored by Google. They were worthless for both linkbuilding and negative SEO. This may or may not have changed with more recent updates.

In all of this spammers are largely sailing blindly and most of their analytic "insights" are circumstantial. Pointing a large amount of XRumer links to a site may only be a single signal to start a deeper investigation: if that turns up nothing spammy, chalk it up to ineffective negative SEO, and investigate deeper.

blauwbilgorgel··on Hello, this is an extortion email
It is important to document cases like these, so others in the future won't panic when such SEO blackmail happens.

The closest form I could find to report this to Google is: https://www.google.com/webmasters/tools/paidlinks Though not regular selling links for money, I would make note about that in the additional details and provide a link to their XRumer domain list.

In my experience the folks over at the webmaster forums: https://productforums.google.com/forum/#!forum/webmasters are good at, and interested in, handling these sorts of cases.

blauwbilgorgel··on Winnowing: Local Algorithms for Document Fingerprinting (2003) [pdf]
Large scale online ML is able to learn from terafeature datasets. The accuracy is often equal or slightly worse than in-memory techniques.

Vowpal Wabbit is made to scale. You set a fixed bitsize and words and n-grams are hashed. So if you expect 2^32 unique words you set the bitsize to around 32. More data is usually better. Linear speed-ups by adding parallel machines.

I too think that HN's user patterns could be different than other web estates. With large scale spam filters like at Yahoo mail, I believe they employ two (or more) models: One fitted on your inbox, and one fitted on everyone's inbox. That ensemble model should be able to specialize on your behavior, yet still be able to detect general spam that it has already seen in other boxes.

>Is this the basic assumption that close documents also have close information entropy?

Yes, that is the gist of it. The better the compressor, the closer NCD will approximate NID.

A simple principle: Compressors do a better job on repeating data patterns. If two files or documents share data patterns, then adding these together and compressing, will result in a smaller filesize, than if you concatenate and compress two files that don't share any data patterns.

NCD works on text, but not as good as other algo's for NLP. Sometimes PAQ (very slow, but efficient compressor) is used on genome data, or bzip on binary files like virusses. I don't think it will be practical here, since for a comparison every other file would need to be concatenated and compressed. If not using a fast compressor like Snappy or Gzip this would take a while, over a simple cosine distance between tokens.

blauwbilgorgel··on Winnowing: Local Algorithms for Document Fingerprinting (2003) [pdf]
For fast and larger scale near-duplicate detection see Google All-Pairs Similarity Search: https://code.google.com/p/google-all-pairs-similarity-search... . The SimHash algo: http://matpalm.com/resemblance/simhash/ and Gensim's Similarities Class: http://radimrehurek.com/gensim/similarities/docsim.html

For blogspam detection can't PG whip up a nice Bayesian model for that ;) ? Again if you want fast and large scale spam filters see Vowpal Wabbit: http://hunch.net/~vw/ . For something more advanced see scikit-learn: http://scikit-learn.org/ which has many algo's up for this task.

A new option you may not have considered yet, is to crowdsource this task. Just offer a labelled dataset and let the HN and ML community have a go at it. I'd love to do content and URL-based spam modelling.

Page 1 of 12Next →