298 karma · joined January 7, 2011
(Another telltale sign is the fact that the link refers to Rolf Heuer as the Director General, but in fact he is the former Director General, while the current one is Fabiola Gianotti.)
But showing them the _reason_ why a certain website is blocked can become an opportunity to teach people critical thought, something that other comment threads point out.
Let me be more precise, and use tptacek's words in https://news.ycombinator.com/item?id=12999887: "fake news = spam sites built from scripts consisting largely of a backcatalog nonsense stories [...] with a one or two carefully produced fake stories as a "payload".
I think this definition is as objective as it gets, and clearly excludes things I might disagree with politically, but are not the target of this extension.
I chose "safety" because I think I copied Chrome's message when visiting a website with an invalid certificate.
So, given some training data that produced two reasonable clusters with respect to the ground truth, I have a model that I can expect to generalize well on new data.
Now, this is not what that notebook shows, because it's missing the evaluating on testing data! The main point of the notebook is that the Jaccard Distance of the tokens of the HTML of the page, despite being very simple, appears to generate a reasonable model.
The main issue I wanted to address is the fact that such a blocker must be widely installed to be useful. Therefore, the Facebook message is intended to prod current users of the extensions to ask that their friends install it as well.
For example, I would never blacklist Breitbart, because I'm not interested in censoring political opinions I disagree with. I just want to free us from the burden of these websites that add nothing to the world and leech attention from everyone.
The Italian Senate offers a SPARQL endpoint [1], which unfortunately doesn't offer access to the texts of the amendments. So I had to roll my own and create a small spider for them using Scrapy [2].
[0]: https://github.com/jacquerie/senato.py/blob/master/analysis....
[2]: https://github.com/jacquerie/senato.py/blob/master/senato/sp...
Stack: Ruby, Rails, HTML5, CSS3, JavaScript, jQuery, D3.js, Git
Resume: https://github.com/jacquerie/cv/blob/master/cv_eng.pdf?raw=t...
Contact: jacopo.notarstefano [at] gmail.com
I am in my last year of my Master's Degree in Computer Science at the University of Pisa. I'm looking for a remote internship this summer on ANY technology, not just the ones I listed.
I think it's lacking the #include/#input feature, and I'm not sure whether it supports labels and internal references, but apparently it has some form of bibliography.
I'm currently building http://www.bandol.it, a Humble Bundle clone targeted at Italian Independent Music (and, maybe, books).
I'm experiencing the dreaded chicken and egg problem: bands won't give me their music if I don't have a big following, while people won't follow me if I don't offer them good bands.
How did Humble Bundle solve this?
require 'benchmark'
def random_string(length)
result = (1..length).map { (65+rand(26)).chr }.join
result[rand(length)] = rand(10).to_s if rand > 0.5
result
end
Benchmark.bmbm do |b|
b.report("\\d") do
(1..1000).count { random_string(1000).match(/\d/) }
end
b.report("[0-9]") do
(1..1000).count { random_string(1000).match(/[0-9]/) }
end
b.report("[0123456789]") do
(1..1000).count { random_string(1000).match(/[0123456789]/) }
end
end
gives: ~/Code/ruby% ruby regex.rb
Rehearsal ------------------------------------------------
\d 0.690000 0.000000 0.690000 ( 0.712500)
[0-9] 0.690000 0.000000 0.690000 ( 0.703990)
[0123456789] 0.680000 0.010000 0.690000 ( 0.705759)
--------------------------------------- total: 2.070000sec
user system total real
\d 0.710000 0.000000 0.710000 ( 0.791722)
[0-9] 0.700000 0.000000 0.700000 ( 0.708210)
[0123456789] 0.690000 0.010000 0.700000 ( 0.713355)