HNHacker News
TopNewBestAskShowJobs

jdrock

695 karma · joined October 29, 2008

Always learning https://datafiniti.co
submissionscomments
jdrock··on How a 30K-member Facebook group filled the void left by Uber and Lyft in Austin
The consensus is that U/L's aggressive political campaigning prior to the vote backfired big time. They really screwed up on messaging by not understanding their audience. Had they made it about losing 10,000 jobs, they might have won. Hell, if they had just not called people multiple times without consent, they might have won.

There are multiple efforts going on as a result of the vote. Former U/L drivers are trying to put together a rally. A non-profit group (Ride Austin) has sprung up to replace U/L. Certain council members are trying to fight the rest of the council. The state legislature is consider a bill that would pre-empt city legislation.

jdrock··on How a 30K-member Facebook group filled the void left by Uber and Lyft in Austin
The title and content of this article couldn't be farther from the truth.

1. Uber and Lyft's absence has created a huge void that remains to this day. There are crazy long lines at the airport.

2. There are multiple ridesharing companies that have sprung up in the meantime. Arcade City is just one of them.

3. I'm pretty sure Arcade City will disappear once either (a) U/L return or (b) one of the newer companies (Fare, Fasten, etc.) get more drivers.

Transportation in Austin is terrible right now. Arcade City hasn't changed that fact.

jdrock··on How to crawl a quarter billion webpages in 40 hours (2012)
This isn't particularly difficult anymore. The most interesting challenges in web crawling around turning a diaspora of web content into usable data. E.g., how to get prices from 10 million product listings from 1,000 different e-retailers?
jdrock··on A Web Crawler with Asyncio Coroutines
Here's how many URLs we crawl every second with 80legs (http://www.80legs.com)

https://dl.dropboxusercontent.com/u/44889964/Descartes%20%20...

This translates to about 700MM/month. The bump you see this month is just us adding more crawling nodes to our cluster.

jdrock··on Spot the Ball
Trying to triangulate based on where players' eyes are looking rarely works. In many cases, the goalie is looking in a completely different direction than where the ball is. I guess that's why they're getting scored on.
jdrock··on Taking Netflix’s Vector Performance Monitoring Tool for a Spin
Any word on integration with Graphite?
jdrock··on Ask HN: Who is hiring? (April 2015)
Datafiniti - Data Engineer and Distributed Systems Engineer, Austin, TX

Data Engineer

Data engineers form the core of Datafiniti. You’ll be responsible for turning customer needs into usable data in Datafiniti. You’ll also be responsible for developing tools to help monitor and improve overall data quality. You should be familiar with data structures and basic algorithms, as well as have experience in 1-2 programming languages (Javascript a plus).

Distributed Systems Engineer

As a Distributed Systems Engineer, you’ll be responsible for developing, expanding and managing a highly-scalable architecture that handles our web crawlers, database and other applications. Familiarity with the following technologies will help in this role: Cassandra, Ruby, Erlang (or another functional language), Java, AWS, and Chef.

https://datafiniti.co

jdrock··on Show HN: Scraperjs – A versatile web scraper
Let us know if you'd like to integrate this with http://www.80legs.com!
jdrock··on Show HN: Get paid for crawling the web
Our intention is to pay the developer. Of course, there'd be nothing stopping the developer from paying out to users as well.

Happy to get more feedback!

jdrock··on Show HN: Get paid for crawling the web
We're experimenting with a new way to power our Datafiniti search engine. Essentially you get paid for running a distributed web crawler through any Chrome Extension.
jdrock··on Ask HN: Who wants to be hired? (August 2014)
Let us know if you're interested: https://www.datafiniti.net/home/careers
jdrock··on Ask HN: Who is hiring? (October 2013)
Austin, TX - Datafiniti (https://www.datafiniti.net)

Datafiniti is the world's only search engine for data. We crawl and index close to 1 million websites each month to create structured, searchable data. Right now we have over 70,000,000 records on businesses, people, and products - and growing.

Build bleeding-edge technology with a team that's all good people. We're looking for experienced or junior-level sales and back-end developers. See https://www.datafiniti.net/home/careers to learn more.

jdrock··on Ask HN: Who is hiring? (September 2013)
Datafiniti - Austin, TX.

We're looking for a Sales Engineer and Operations Engineer. Come help us scale the world's first and only search engine for data!

Full details here: https://www.datafiniti.net/home/careers

jdrock··on Ask HN: Who is hiring? (May 2013)
Datafiniti - Austin, TX

At Datafiniti, you'll get to work on an absurdly ambitious problem - building the world's first search engine for data. We solve challenges that push the boundaries of what a search engine can do. Here are a few of the awesome things we work on:

  - Managing thousands of cloud nodes
  - Processing content from billions of URLs
  - Automatically converting web page content to valuable data
We're looking to fill roles for:

  - Sales Engineer: work with our clients team to implement customer projects.
  - Data Engineer: improve the coverage and quality of our data.
  - Ops Engineer: improve the scale and reliability our search engine infrastructure.
More details at http://www.datafiniti.net and https://angel.co/datafiniti/jobs/
jdrock··on Ask HN: Who is hiring? (April 2013)
Austin, Texas http://datafiniti1.theresumator.com/

We're building the world's first search engine for data at Datafiniti. Work on fascinating problems that involve working with billions of data points, building intelligent agents, scaling out massive data collection, and more.

We have a small, close-knit team that enjoys working and hanging out together. Sending an email to careers@datafiniti.net will go straight to me, the CEO & founder.

jdrock··on Come and help save Posterous from oblivion
@jacquesm please send me a list of URLs to crawl (10M+), and I'll set up an 80legs job to do this. shion - at - 80legs - com.
jdrock··on Semantics3 (YC W13) Is A Massive Consumer Products Database To Rule Them All
We're working on this problem at Datafiniti (https://www.datafiniti.net). Since we index hundreds of sources for a similar service, we can leverage some basic string comparison techniques to normalize records from different sources and fill in attributes that are missing from any one source.
jdrock··on Show HN: Datafiniti.net - We created the Google of Data
Happy to discuss any questions the HN community has about our site!
jdrock··on Ask HN: Who Is Hiring? (November 2012)
Austin, TX

Client Project Developer

Datafiniti is the world's first search engine for data. Datafiniti's search results are complete data sets taken and aggregated from the web. You can search for information on places, people, products and pretty much any data available on the web. We are converting the entire web into a single, searchable database of knowledge.

Here's an example search: https://www.datafiniti.net/search/places?%5B%7B%22v%22:%22au...

In addition to our search engine product, we assist clients with a wide variety of custom data collection projects. We're looking to hire someone that can work on our clients team as a project developer to implement and manage these projects.

We're a small, close-knit team, so you'll have the opportunity to make a big impact on the company and its future success.

Full details are here: http://datafiniti1.theresumator.com/apply/V5pGWC/Project-Dev...

You can also email your resume to careers at datafiniti dot net.

jdrock··on How to crawl a quarter billion webpages in 40 hours
No - inbound data transfer cost to AWS is free.
jdrock··on How to crawl a quarter billion webpages in 40 hours
Being the CEO of a firm that offers web-crawling services, I found this post very interesting. On 80legs, the cost for a similar crawl would be $500, so it's nice to know we're competitive on cost.
jdrock··on Looking to Get Creative? Leave the Echo Chamber. Drink. Skip SxSW.
That's not an accurate impression of SXSW. There are many companies and folks there doing fascinating things and solving hard problems. If you just go the panels and hang out at the convention center, you're going to miss out. If you have a personal network that you can tap into and leverage to meet more folks, you get a huge benefit out of SXSW.
jdrock··on Ask HN: Who is Hiring? (March 2012)
Houston/Austin TX - Full Time - http://www.datafiniti.net

Datafiniti is a small startup building the first search engine for data. We crawl the entire web, collecting data on businesses, places, people and things and then normalize all of that data into a single, searchable database.

We're looking to add a talented and passionate User Interface Developer to our tight-knit team. The UI Developer will be responsible for building elegant interfaces to super-massive amounts of data.

To apply, email careers@datafiniti.net.

More info available at http://www.datafiniti.net/index.php/careers

jdrock··on Ask HN: Who is Hiring? (July 2011)
Houston, Texas (H1B accepted)

80legs is building a next-generation web data platform - a service that will allow anyone to run SQL-like queries on all data available from the web. We're looking for talented folks to help us :)

Positions available:

  * Data Technical Lead - handling all things data, ML experience a plus
  * UX Engineer - building a search interface for 1B+ data points
  * SysOps Engineer - improving back-end performance of a complex infrastructure
More info at http://www.80legs.com/careers.html. Feel free to email me at shion -at- 80legs -dot- com.
jdrock··on IndexTank / 80legs Crawlathon (developer contest)
Not necessarily more traffic, though that's possible. More likely it was the fact that it was coming from multiple IP addresses.
jdrock··on IndexTank / 80legs Crawlathon (developer contest)
True story: We crawled them a while back (before they expanded their engineering team) and because of our distributed system, "alarms" were going off. Rather then take time to tell their system we were not a DDOS attack, they put us in robots.txt. I imagine their small team had other stuff to work on.
jdrock··on IndexTank / 80legs Crawlathon (developer contest)
Our default crawler obeys robots.txt and it looks like the /results URLs are not allowed. However.. I think you could start from a URL like http://www.youtube.com/watch?v=sAzhOSbxMiI and then crawl to the linked videos from there...
jdrock··on IndexTank / 80legs Crawlathon (developer contest)
Er.. not true, the contest plan (which is different from the free plan) allows up to 10 MB download. Registered contestants should have access to their plans within an hour of registering.

And if you don't have it, just contact us: http://www.80legs.com/contact.html.

jdrock··on IndexTank / 80legs Crawlathon (developer contest)
Specific documentation for 80legs is available at http://wiki.80legs.com. To get fancy, you may want to check out http://wiki.80legs.com/80apps to make your own custom extractors. Note that any custom extractors you write will output files in binary in 80legs, so you'll need to convert the byte array to .txt, .csv or .xml or whatever format you want.

(We post custom extractor results as binary because you can also return stuff like images!)

jdrock··on IndexTank / 80legs Crawlathon (developer contest)
Hm.. can you try again? Seems to be working from our end!
Page 1 of 8Next →