Google Squared
techcrunch.com
techcrunch.com
I have to note that this sort of data scraping shown here, by Google, with no incentive left for anybody to actually travel to the original source web pages, seems to make the first groups case much, much, stronger.
I'm envisioning consultants, like SEO folks, but these are WDO folks. (Web Data Obfuscation). Web Data Obfuscation is the term I just made up for structuring your web page so Google couldn't scrape it easily. The page would still appear in the google index, as that is a benefit to you, so you would make available enough data to be indexed, but not enough to be usefully scraped.
I believe that once you publish data (particularly data, but even other kinds of information) on a public, indexed web page, you automatically relinquish control over how it will be used.
That's just reality, and it's also the most profitable way to view online published data, from a global perspective. It is better for all of us if the act of publishing data on the web grants an automatic licence to the downloader to mash it up any way he sees fit. The alternative scenario, where you have to ask for permission for every bit of data, is frightening.
Just because it's "big Google" who is doing the mash-up doesn't make it less ethical than if it was some start-up coming out with a new product (or, say, Wolfram Alpha).
Mash-ups are ethical only as long as they provide more money in the pockets of the people you are stealing content from. If they do that, fine. If they aren't, your mash-ups are not ethical, and neither are Google's..
Particularly when it's factual data, rather than, say, an article. You might recall that the copyright acts do not protect factual data.
But there is what's legal, and there is what's ethical. If I traveled around the country, measuring the heights of roller coasters for my website rollercoasterheights.com, I did it to get people to come to my site. If that information is harvested from my site and displayed elsewhere, I've done a lot of work for nothing.
This is pure speculation, and I personally doubt its true. First, much of the content of on the web amounts to echoing, filtering, and distorting high-quality source data. If you can tap that source data, are the gains you realize from what's been derived from it worth the extra cost of going out and doing the work to acquire it and then separate out the noise?
Second, there is a lot of real or perceived value in knowing where your data is coming from. Tons of companies pay tons of money to get data and statistics from sources that they can trust. Until Google's willing to vouch for these results beyond "we tried", the bespoke curatorial approach is going to keep capturing these dollars.
Then you could also assign sources a reliability rating based on how acurite the information they provide is compared with other reliable sources.
Kind of like a new pagerank but for data integrety.
Search camera on Wolfram and I expect you get lost of data on the history of camera's and other such stuff. Google seem to be offering to provide a list of camera models with some pertinent data for each... I suspect the latter will be more "useful" (especially as Wikipedia would probably be fairly reliable for the camera background...)
I think your point is valid: a curated database will be the #1 source for info on camera history and facts. Unfortunately that means Wolfram is competing with Wikipedia not Google... and that is probably even worse for them :(
- get a bunch of structured, verified, curated data - use mathematica to understand and reason about the data - use NLP to expose mathematica to the web surfer.
What google wants to do is get a bunch of structured data from unstructured data. Great, but competition for Wolfram Alpha? Doesn't look like it right now.
But I guess we'll have a better view of this in a few weeks.
it beggars belief that, after dozens of stories on the topic, each with its own universe of commentary, most of which pointing out basically what you've said, that they are unaware that their premise is flawed.
2) The interviewers are surprisingly immature and unprofessional. They are downright rude to the demonstrator and are clearly ignorant about the amazingly cool technology that they are privileged to be seeing.
If I ever do a search as generic as 'camera', it's very closely followed by a more detailed search (eg, camera olympus "flash time"). Usually, however, my first search is the most detailed, and I become more vague if I require more results.
As I'm sure most of us are aware, this helps quickly deliver the outcome / answer we seek from searching, but I'm continually surprised at how many people bang a series of vague search phrases into a Search Engine and spend time sorting through the chaff.
Google Squared seems to prompt the sort of thinking I either do before hand, or as an immediate result of seeing 8.6M results. Google-fu enhanced.
(My favourite evidence of this is from Allyn Gibson's blog, which is routinely located by people Googling the phrase "things that happen on my birthday". Think about it, or read #5 on this list http://www.allyngibson.net/?p=1686 )
Google is likely carefully studying their competitors continuously and taking prudent actions as necessary. This product doesn't seem like an answer to Alpha as much as it does an interesting 20% project growing up.
It matters not whether the "New Stuff" from Google is a competitor or reaction to Wolfram or completely unrelated. What matters is the media thinks it is competition and generates all this Google vs. Wolfram media hype :) Wolfram get some sympathy as the underdogs but ultimately it just pushes Google brand (they still just hold the public opinion of not being evil: so if you see a Google vs X discussion you imagine it is Good vs. Good battle and that both products are OK).
It's a tactic Google have always used - often to superb effect :)
You know what Alpha reminds me of most (as a tester - try it out yourself on 5.18)? 'Insert Field' in MS word, where you can insert the date & time, or a reference to a cell in an Excel spreadsheet, or some other piece of external data, and have it automatically update every time the document is loaded. It's sort of like having widgets but being able to call them up with a simplistic natural language interface.
Not all of Earth's <Alpha: world population> people will appreciate this, but it will probably be popular in <Alpha: countries with most universities per capita>. (see?)
Then today they flip and write everywhere that structured is the next big thing.
Ironic, that's all I'm pointing out. I don't know yet which technologie(s) will have an impact (note that it's not an either/or choice, both WA and Squarred can succeed).
Once it's deployed across Google's entire cloud it'll speed up nicely.
Here you have two powerful companies, Wolfram and loopt. Google has interests in technologies similar to what these two companies offer. They didn't buy/license loopt's technology, so why should they buy/license Alpha? There is no basis for what rms said.
1. Google is coming out with a new product called Google Squared. It's related to searching data sets.
2. Wolfram is coming out with a new product called Wolfram Alpha. It's related to searching data sets.
3. rms observed that #1->product and #2->product are related, and if their technology doesn't explicitly overlap, WA might be a good candidate for acquisition in order to broaden Google's hold on searching data sets.
4. You issued a non-sequitur. It's not related to the thread at hand, and I'm not even sure why Google would even want Loopt. (Please explain your rationale why Google would want Loopt, cuz I don't get it, but know that explaining it won't strengthen your argument against rms).
5. Sparknotes Version-- Your argument is this: Google didn't license Y's technology. Y's technology is from a powerful company *Therefore, Google won't license X's technology, since X is a powerful company.
6. What you're trying to say is that Google won't necessarily go around trying to license technology from every company that has intersecting interests.
(To satisfy your model:
1. Google is coming out with a new product called latitude. It is related to finding people around you.
2. loopt has a product called loopt. It is related to finding people around you.
3. rms observed that #1->product and #2->product are related, and if their technology doesn't explicitly overlap, loopt might be a good candidate for acquisition in order to broaden Google's hold on finding people around you.)
Not everything on Hacker News must be
P1
P2
P3
--
Qhttp://news.ycombinator.com/item?id=567298
I knew when I wrote that it would probably be received poorly; indeed, it was volatile. But it was creative, and there was some truth to it.
HN is too simplistic. I would venture to say it ultimately doesn't work, with evidence in responses like "Pot, socially" being #2 in the top comments under lists. HN is fun to experiment with, though.