HNHacker News
TopNewBestAskShowJobs

LisaG

641 karma · joined October 20, 2010

Science geek with a strong affinity for computer nerds. Director of Common Crawl www.commoncrawl.org Former Chief of Staff at Creative Commons www.creativecommons.org
submissionscomments
LisaG··on You Can’t Do Data Science in a GUI
Also, the article does state what Hadley's take on the question is: "As Wickham defines data science as “the process by which data becomes understanding, knowledge, and insight”, he advocates using data science tools where value is gained from iteration, surprise, reproducibility, and scalability. In particular, he argues that being a data scientist and being programmer are not mutually exclusive and that using a programming language helps data scientists towards understanding the real signal within their data. "
LisaG··on You Can’t Do Data Science in a GUI
Did you watch all of Hadley's video? You might get the title more if you saw/see the whole talk :)
LisaG··on Ask HN: Would it be legal if I put my web scraping scripts (lib) on GitHub?
As long as you obey robots.txt there is nothing wrong with crawling. Your code in GitHub doesn't give any indication of what sites you collect data from so there is no indication that you are scraping instead of using it to crawl in an acceptable manner. Though it wouldn't hurt to label your work as crawler scripts instead of scraping scripts ;)

Why use your own scripts and not Nutch?

Do you know about Common Crawl? https://aws.amazon.com/public-datasets/common-crawl/ It obeys robots.txt so it may not have everything you want, but it could save you part of the effort of crawling yourself.

LisaG··on Ask HN: Who is hiring? (July 2015)
San Francisco CA Full-time / Onsite

New, somewhat stealth startup, for profit company focused on social good.

We have a very talented team so far comprised of : full stack web dev, data architect, 2 junior software engineers, CTO, CEO (me), 2 marketing people, and a business operations person.

We are looking to add a designer and devops.

We work out of the top floor of my house for now - it is a comfortable space. We have funding. The team members we have so far are wonderful to work with, everyone gets along well, and we all feel like what we are building is work that matters.

Please email me if you want to hear more about the team, the stack, and the product.

LisaG··on Glove: Global vectors for word representation
So excited so see Common Crawl data be useful for such fascinating work!

I work at Common Crawl :)

LisaG··on [dead]
Love this idea!!
LisaG··on Why Aren't We Reading Turing?
Great question and great blog post! I am looking forward to reading the Homay King stuff that uses Queer Theory and will probably reread Computing Machinery and Intelligence more thoroughly.
LisaG··on Prismatic wants to build a social network around all of your interests
I played around with Prismatic before but it just didn’t grab me and I found I didn’t use it that much.

This new version is a whole different animal. Not only is it much prettier (great design) but they seem to have seriously improved their relevance algorithms. I would be very interested to hear from their data team why the relevance is so much better now - anyone from Prismatic monitoring these comments?

LisaG··on 102TB of New Crawl Data Available
There will be news about a subset sometime next month!
LisaG··on SwiftKey’s Head Data Scientist on the Value of Common Crawl’s Open Data [video]
If you don't feel like reading the paper Sebastian wrote on the Common Crawl data, he gives a summary of his findings in this video.

Link to full paper: http://bit.ly/14dxSJq

LisaG··on A Look Inside Our 210TB 2012 Web Corpus
We do think it is worth it to avoid duplicative efforts.

Suppose you crawl 3 million pages and you pay for the compute and storage costs. Then the next person who wants crawl data goes through the same effort and pays the same costs. Doesn't it make much more sense to have a common pool of open data that everyone can use? Even if the effort and costs are low, they are not zero.

For the smaller frequent crawl, we are working with Mozilla and we are will do the top pages (top according to Alexa).

LisaG··on A Look Inside Our 210TB 2012 Web Corpus
Internet Archive (currently) doesn't want to put their data on any cloud service. We believe it is crucial that people can easily access and analyze the data so we put it on various cloud platforms. We are talking with a few organizations about getting data donations that we could put in our corpus and make available to everyone, but nothing is settled enough that I can publicly comment on those potential partnerships yet.
LisaG··on A Look Inside Our 210TB 2012 Web Corpus
Limited resources are the only reason. We are working on a subset crawl of ~3 million pages that will be published weekly starting two weeks from now. But doing the full crawl takes a lot of time, effort and money.
LisaG··on Antibiotic resistance: The last resort
That's interesting. We also could revive phage therapy which uses bacteriophages (viruses that replicate in bacteria).

http://en.wikipedia.org/wiki/Phage_therapy

LisaG··on Google Fiber coming to Austin
If you are bored in the San Francisco Bay Area the problem is likely internal rather than where you live, so moving (even to somewhere awesome like Austin) will not resolve it.
LisaG··on Share code that uses new URL Search tool and win AWS credit
I hope that some of you who use/play around with the Common Crawl data will try out using the JSON files from the URL Search and then share your code.

If you didn't see the details in the blog post, Common Crawl is giving out $100 in AWS credit to the first five people who share code that incorporates a JSON file from the URL Search.

LisaG··on Share code that uses new URL Search tool and win AWS credit
"Done" is better than "perfect" should be on a sign hanging in every startup.
LisaG··on Share code that uses new URL Search tool and win AWS credit
Thanks for the catch Djoerd! We will fix it now
LisaG··on Share code that uses new URL Search tool and win AWS credit
From @djoerd Why does @CommonCrawl URL search (http://urlsearch.commoncrawl.org/ ) need 'tld.domain' format rather than 'domain.tld'? Read Google's BigTable paper.
LisaG··on How To Date Like An Entrepreneur
Love this post! Only thing I disagree with is that you only need on to say yes. The first yes might not be the best match for you. Compatibility is not quite as important in business as it is in romantic and sexual relationships.
LisaG··on Blekko donates search data to Common Crawl
Strongly agree! blekko's Bill of Rights is a great expression of their values and of why we should all be using blekko.

blekko Bill of Rights

1. Search shall be open

2. Search results shall involve people

3. Ranking data shall not be kept secret

4. Web data shall be readily available

5. There is no one-size-fits-all for search

6. Advanced search shall be accessible

7. Search engine tools shall be open to all

8. Search & community go hand-in-hand

9. Spam does not belong in search results

10. Privacy of searchers shall not be violated

LisaG··on Blekko donates search data to Common Crawl
Graue I am from Common Crawl. We don't filter for porn. A corpus of web data needs to include porn or it wouldn't be a representative sample of the web ;) We do want to enrich our sample of the web with high-value sites and that is where the blekko data will be so incredibly valuable.

I really appreciate your mention of LGBT and sexual health sites being collateral damage - we need to draw more attention to that problem. I would love to see someone work with Common Crawl to improve methods of distinguishing. Lisa

LisaG··on Blekko donates search data to Common Crawl
I am part of Common Crawl and I just wanted to say that we are super excited about blekko's donation! This is yet another demonstration how much blekko values openness and transparency.
LisaG··on A Review of the TV Show Start-Ups: Silicon Valley
Oh and the girl who says that engineers are secondary to the success of a startup and that people with ideas are the key element. There are a lot of those people around here.
LisaG··on A Review of the TV Show Start-Ups: Silicon Valley
Another archetype is the guy who tries so hard to be a "brogrammer". Sadly those guys exist too.
LisaG··on Common Crawl announces Open Source Big Data code contest winners
So cool you are putting together a presentation of results!!
LisaG··on Ask HN: Who Is Hiring? (September 2012)
San Francisco: Data Scientist, Crawl Engineer

Do work that matters on big data! Common Crawl is an open repository of web crawl data with a corpus of over 100 TB.

We’re looking for someone enthusiastic about open source, net neutrality, open data and keeping the web truly open. Common Crawl is dedicated to building and maintaining an open repository of web crawl data in order to enable a new wave of innovation, education and researchWe’re set to do amazing things this year, and there is no better place to hone your big data skills than helping us manage and process our 100 TB corpus. Plus, you’ll be working within a passionate community and have the chance to interface with plenty of talented researchers, educators, startup folks, and an incredible advisory board.

If you’re looking to do work that matters, come join us!

http://commoncrawl.org/team/jobs/ Email lisa (at) commoncrawl.org

LisaG··on Study of ~1.3 Billion URLs: ~22% of Web Pages Reference Facebook
He did this in from idea to finished experiment in 4 days and used about 300 lines of Ruby. Strong demonstration of how low the barrier can be to working with big data.

"The key lesson I’ve learned from the exercise is that given the tools and data available today, either for free, or at very low cost, it’s possible for anyone to work with relatively Big Data without too much weeping and gnashing of teeth."

LisaG··on Bored in grad school? Learn Hadoop
Site is back up. Thanks for your patience!
LisaG··on Bored in grad school? Learn Hadoop
Hi

I am from Common Crawl. Apologies for the site being down! Too much traffic from HN :) We're working on getting it back up. The Google cache below has all the contents, so please refer to there for the moment. Here's the excerpted beginning..

Learn Hadoop and get a paper published

We’re looking for students who want to try out the Hadoop platform and get a technical report published. Hadoop’s version of MapReduce will undoubtedbly come in handy in your future research, and Hadoop is a fun platform to get to know. Common Crawl, a nonprofit organization with a mission to build and maintain an open crawl of the web that is accessible to everyone, has a huge repository of open data – about 5 billion web pages – and documentation to help you learn these too

http://webcache.googleusercontent.com/search?q=cache:http://...

Page 1 of 2Next →