225 karma · joined May 9, 2018
In the meantime (unless you are dealing with video) most text and image datasets out there that avg Joe needs can easily be stored/processed entirely locally thanks to cheap terabyte drives/multicore chips these days. People just haven't realized there isn't that much useable textual data OR that local computing doesn't require all the overhead of handling millions of requests a second. This is Google problem not an avg Joe problem that is being solved with cloud compute.
But...if you just want to run your own personal search engine say...
Then for Wikipedia/Stackoverflow/Quora size datasets (50GB with with 10GB worth of every kind of index(regex/geo/full text etc) ) you can run real time indexing on live updates with all the advanced search options you see under their "advanced search" page one any random Dell or HP Desktop with about 6-8GB of RAM.
Lots of people do this on Wall Street. People don't get what is possible on desktop cause so much of it has moved to the cloud. It will come back to desktop imho.
1. You assume stopping "revolution" is the goal. It is not. All you have to do is look at what outcomes "revolutions" in an info saturated/low attention span/consumption culture based society have produced in the last 20 years.
2. In 2013 the Washington post did an estimate on how much was being spent on the surveillance society in the US. Everyone reacted in disbelief. Do you know by how much that budget has increased in the last 5 years? Are these people just brain-dead to be spending this kind of cash? Ofcourse not.
The US has made its share of mistakes over valuing freedom and squandering potential. China is overvaluing control and will make its share of mistakes too. We have to learn from both sets of mistakes to arrive at the right balance of where society should go in high noise info saturation environments. These are very new environment that society hasnt been in before and the right path ahead is not as obvious as people think.
The real damaging data to people and society at large is a different set of data. It's the publicly visible counts next to every thought and utterance reinforcing misguided beliefs and behaviour up and down the food chain constantly. Any experienced shrink, psycologist or educator, marketing/PR expert knows applying the right amount of feedback at the right time is critical to how people process info.
Remove/delay/reduce the visibility of like counts/view counts/upvotes/retweet counts that are displayed and the world will be a different place overnight.
Fred Rogers was critical of it even back then. Since then, consumption culture has exploded, objectification of women, violence, escapism whether it is in games, drugs or your social media, news streams is at an all time high. For a Fred Rogers type message to get through the volume dial on all this other stuff has to be dialed down. Otherwise the very valuable message gets lost in the noise.
Kim Kardashian has a 110 million followers. After the Russians (or whoever...maybe just that "400 pound dude in the basement") get her to become the next president I am waiting to see what boywonder shows up and says to Congress.
We know what he and his exec staff thought about fake news before the election. Now who here thinks fake imagery is not a thing? Just take a look at the content produced by the top 100 accounts.
There is going to be a special place in history for the damage boywonder has done. People will have to invent new terms to describe it. Nothing like it has happened ever before.
The people/environment around the person going through the trauma makes a big difference to outcomes imho. And usually it "takes a village". Because though most people who care will want to help, they each bring different strengths and weaknesses to the table. It's a magic combo of the different strengths that make a diff. And realising that and raising group consiouness about it I agree is badly needed these days.
I suppose similar tools have been created by the internet archive folk.
But you are right it should be much easier to archive stuff in 2018 and it isn't esp thanks to all the JavaScript and XHR happening.
Edit: I just took a look at kiwix (been a while) they seem to also now archive stackexchange sites not just mediawiki...so looks like they have different archiving tools for different sites.
Hopefully the day will come where we don't have to become grad level psycologists to handle our new info environment. The professions I was thinking about were air traffic controllers, soldiers, pilots, cops etc whose training involves coping with constant high noise environments. We have lot of lessons to learn from them and their training systems.
I have a feeling psychologists("choice architects"), socialogist etc with software experience will soon start making an impact in dismantling some of the misguided pieces of the current architecture.
they will never be trained to handle
that no previous generation has ever handled
What we are watching unfold across any and every issue, are symptoms of people breaking down due to that overload.
To handle this new over saturated info environment (which up until now has been sold to everyone as a good thing) one must look to professions that are trained specifically to handle such overload. And there are many.
What I don't like is that the most hyper, overenthusiastic, over the top characters send a signal to kids that it's the only route up the ladder. It is not.
Back in 2006-7 I attended a conference where there was a big panel on what the future of news was going to look like. Lots of newspapers were shutting down. They made 3 predictions - news would be replaced with pandering to the lowest common denominator in the arms race for eyeballs, ads and marketing content looking like news(fake news) would exponentially increase, real journalists would be replaced by the talking head/oped class.
They concluded that the only way out, was to hope that revenue generated from the transition from print to digital/mobile would prevent these three things happening.
Its been a decade now and that transition is more or less complete and many have figured out how to stay afloat and yet all three predictions seem to have come true.
People underestimate this.
So pick subjects that interest you or are sure to tickle/provoke/excite the people you intend to show the work too.
And that's hard work. "Work is love made visible"
Ask Google who played the mens semifinals of Wimbeldon three years ago and Google will tell you it indexed 6 million pages to provide a link that may or may not have the 4 names I am looking for. Why is it doing all this pointless work? And why is it that dumb in 2018?
We have got so used to what it does that lot of people have stopped asking questions about how it does things and wether all the stuff it does is required.
Wolframalpha, Freebase/SemanticWeb/Wikidata/dbpedia approaches, NLP/NLU are still very underdeveloped and untapped.
Having open and distributed indexes like we see in nature with DNA is also totally unexplored because of Google type centralised index monopolies in various domains. It just takes a Gig or so to store a local offline index off all Wikipedia or Stackoverflow pages. And given the massive RAM and hard disks everyone has these days why aren't we seeing sophisticated local offline search apps?
The internet is getting exponentially more noisy day by day and in many ways its easier to find quality info going through a top notch library's index than wading through Google's. So there are lots of blindspots and areas to explore in search right now imho.
1. Targeting is much more sophisticated. Today you can round up every lunatic who hates squirrels or whatever by tomorrow morning and influence what people need to focus on.
2. Messed up social signals i.e. like counts/views/retweets/upvotes etc are just misguiding people left and right. If I don't know what to make of something/too busy/too distracted etc I fall back on these highly inaccurate proxies to make decisions. This is happening all the time and not just to the poor or semi-literate but too highly educated folks. Everyone is unconsiously nudged to support X or Y cause, org or person purely based on these numbers. Take those numbers away and it would be a very different world.