Mozilla blogger bought 1 million Facebook entries (full name, e-mail) for $5
talkweb.eu
talkweb.eu
Kinda sensationalist at the end - "DO YOU STILL FEEL SECURE?" Uh yeah, I do. I get hundreds of spam a day that Google puts into a little folder for me to erase. Go ahead and email me, and join that exclusive folder right away.
Oh you want to send me a private message on Facebook? Facebook is kind enough to put messages from non-friends in a special folder too, and I never check that.
I feel secure.
edit: refuted by Permit's comment below: http://news.ycombinator.com/item?id=4688893
Regarding facebook I partially agree with u. Sometimes the message from non-friends goes into other folder and sometimes I can see it in the regular mail folder.
All spam systems still look at the content of the message plus the reputation of the IP/domain when determining if a message should be marked as spam or not.
Seriously though, I don't think a couple of headers will cause an email to become not spam. Google's filter is probably Bayesian, and considers the content of the message, the matching between the originating server and the reply-to address, the spam history of the originating server, any links in the message, how many times it sees the exact same message, how many people mark it as spam, etc.
So maybe it will get through to the first few hundred people, and then they'll block it.
As the saying goes, if you're not paying for the product, you're the product. The new twist here is that the product (i.e., your FaceBook info) is now being sold in the open market for only $5.00/1,000,000, or $0.000005 per person.
Prices normally go down only when supply exceeds demand, so the inescapable conclusion is that there's abundant oversupply of this product in the open market. Yikes!
[1] I'm using "pay" loosely here. I have no idea what they had to give up or produce in order to get the data, but presumably FB received something of value.
[1] Not necessarily a good indication, as this may be a last ditch effort convert _some_ value out of their app development efforts, and who knows to how many buyers the seller has sold this data.
No, it doesn't cost anything to create a Facebook app.
I don't think that rule applies for digital goods where the cost of reproduction is zero. The supply is infinite.
Takes more servers and fatter pipes to support 10,000 downloads a day rather than 500.
If you want to talk in totally abstract terms, digital goods in general tend to have marginal costs associated with them. In the context of this discussion, there is no supply and demand factor, there are no marginal costs, and there is no market force called scarcity.
2(N) does not yield W, regardless of cost to copy (N).
To get W, you will need to do something more. This will not be cost-less. That's the more general case.
Thus, the cost of hosting is a marginal cost (probably zero in this world of pastebins and digital lockers). The fee taken by the payment processor is a marginal cost. The cost of finding twice as many emails is not.
notatoad 1 day ago | link
I don't think that rule applies for digital goods where the cost of reproduction is zero. The supply is infinite.
To sum, "the cost of reproduction" is <not> the cost of "supply", unless the supply is assumed fixed. Thus the second sentence does not follow per-se.
For example, Adobe Photoshop probably costs a lot to design. It has really high fixed costs, because you need to hire good developers and implement a bunch of advanced operations. However, once Adobe pays the fixed costs, the marginal cost of Photoshop is pretty minimal: packaging, printing a DVD, maybe some marketing. It still costs a lot because the fixed costs are so high, and there's not much competition.
Conversely, a plumber has relatively low fixed costs: a truck, some tools, and some training. But plumbers also cost a lot, and this is because they have really high marginal costs: they have to spend an hour at the house of each and every customer.
So I agree with you, there may be high costs in acquiring email addresses to sell. My point is that they are in no way marginal costs.
example: reseller> pays adobe every month/quarter example: adobe> pays versioning costs every 24 months
Provided you shrink the window of analysis, you can say "already paid for inventory, just amortizing it". But in that case, you don't have unlimited supply, you just have whatever you paid for.
In the case of adobe, despite having "unlimited copies" of CS5, they would (eventually) run out of supply of salable product if they did not version into CS6. So while its trivially true they could make unlimited copies of CS5, its not a great idea to perceive this as unlimited supply. The supply that matters is the part people are willing to pay for--this is the marginal information content-- not the marginal bit content of what is delivered.
In some ways I don't think we're disagreeing, just focusing on different elements of the analysis. My larger point was exactly that -- keep in mind the broader elements that are considered as relevant by CxO.
THe CEO of adobe makes decisions, for examople, about how often to incur the marginal cost of versioning the next Creative Suite, how rapidly and how much to budget, etc. COO of facebook looks at the marginal cost of data centers for the next 200 million users, etc, in part because s/he is looking at timeframes and scales which are not the same at the level of a project team, etc.
A 2MB file does not cost twice as much to email send to someone as a 1MB file; you aren't going switch to a different internet connection or email provider because of your file is twice as big. The first 1 byte is very expensive and every subsequent byte has no observable marginal cost until 20 orders of magnitude later.
You can't create more <valuable> data per-se by making X copies of the same data (in the sense of it having value for marketing/analytics). that is just monetizing existing data. ie, The marginal cost to relicate a set, provided it was given to you for free...just assumes away a non-trivial part of the equation....getting the data.
At scales of 100m to a 1Billion...is not trivial or costless. Lastly, if you only have one set of data (say 5m users of data), that is a finite supply. You could have 2 sets (10m users). Thats not the same as having 2 copies of 1 set (of 5 million). A customer might pay per user for a lead, but wont pay twice for two copies of the same info. now, if somebody shows up with 100s million, it might impact supply/demand (depending on comparabilit/uniquenss). But those differences cannot be assumed away at zero cost, imho. Hope this make more sense, was not trying to argue just for the sake of it.
________________
[1] That scales with data entry, etc (if nothing else) at the origin (ie, this is a FB user cost == per user). Even if its non-cash its ~$0.85c per 15 minutes of time for a western eurpoean ABC1, back of the envelope. And that scales linearly.
I didn't actually intend to say that the cost here is mainly bandwidth or X for any X. My point is more than it can be true that the marginal cost of anything can be so many orders of magnitude less than the non-marginal cost for certain ranges that it isn't worth considering. Data sets based on facebook profiles almost certainly come from people approving shady apps or someone set up a crawler in a way that is able to get a lot of information before being detected as a crawler. In either of those scenarios, the person who set it up effectively paid a flat upfront cost and ends up with X number of users, and there are no linear costs (no manual verification or paid data entry at any point). They cannot spend half as much time and get X/2 users or even maybe X/100 amount of information.
In practice they could spend more time getting more users to give access to their random app, but the marginal cost function is just insanely nonlinear; its effectively 0 at some places and probably tends towards infinity an order of magnitude higher than that.
The file contains "just" full name, e-mail and URL. Thieves got the information thanks to their Facebook apps (no idea of its name), it could happen with any third-party app.
There were the obvious checks for CAPTCHAs when too much activity was detected, but other subtleties as well. If you looked at too many people's profiles, emails wouldn't be displayed as text, but as images. A person would be unlikely to notice as the pages looked identical, but dynamic changes like that make it harder to scrape some things. Introducing even rudimentary OCR requirements is enough to turn away a lot of programmers.
I'm not saying it's not possible to pull off. But Facebook has set it up so any money you might make this way will likely not be worth the development time required.
To be perfectly honest, I've kind of fallen out of love with web development in the last year and have taken more of an interest in algorithmic trading. I appreciate the interest, though. :)
Many people talk about caring about their security in an almost idealistic view; few actually care in application.
e.g. try https://graph.facebook.com/1112112584 with curl or so ... (sorry random member from the published list)
Some spammers have probably been harvesting that API for a long time ...
Also, this doesn't give you the email.
I have never noticed this with any of my other unique emails, just the facebook one.
Try 3.5 million: http://fiverr.com/palash1987/provide-3500000-facebook-emaii-...
One of my clients has a FB app so work with this stuff daily. It's so easy to build a full profile of of an active user's life, their interests, their friends, their work and education history. Their geotagged photos and check-ins tell you exactly where they like to go. Pure gold for marketers / spammers.
The information is just too accessible and valuable for people to not abuse it.
OFC there's an oversupply since they can give out an unlimited amount of copies of this same million FB entries.
Big Data is great because it's super re-usable and can be purposed for anyone's specific need.
Unfortunately, as usual your average user thinks nothing of clicking through the permissions page on FB without reading or understanding what it says.
http://graph.facebook.com/4/picture
You don't get the email but that would be really bad.
There are phone books out there that list more information that that. Do you feel secure?