Afraid to Google a thing because I don't want the algorithm to think I like it
twitter.com
twitter.com
kinda forces me to watch something else that is new but at least somewhat wanted in an attempt to make the algorithm look at the new shiny i'm interested in and forget the previous one.
i find it funny though when amazon still make recommendations to me based on a purchase in 1998. yes, Jeff, i still want to buy wrestling VHS tapes...
For example if I consistently listen to 3 unrelated songs [A B C] on youtube together, the algorithm will regularly recommend them to me (because of my unique behaviour, not because they're similar). If I reset my history and then listen to song A, then B and C get immediately recommended even though they aren't similar to A and they don't exist in my listening history which means that the information about my listening/browsing habits is still there in the recommendation model.
Seems to work reasonably well for me.
Nobody knows, I only know I don't want to play this game.
If a real human mentions a video that piques my interest, I may watch it. If I'm interested in finding a video on a particular topic, or a specific scene from this-or-that film or TV show, I will search for it myself. If am interested enough in a certain producer's content, I may "subscribe" to them, but I will be the judge if I want to watch the latest video they uploaded.
There is one advantage: some obscure music videos. There are some rare pearls with the comments section almost exclusively thanking YT algorithm for taking them there. Happened to me so many times that I have a separate browser instance and a Google account for YT and I'm very careful what I click when I use it.
https://www.youtube.com/c/TerminalPassage/videos
I didn't even know there was a Jazz/Fusion rock scene in Japan in the 70's for instance. I never would have heard of this band without YouTube suggestions but really enjoying their work.
I did it for a couple of weeks to see what was being promoted by youtube algorithms in terms of content and ads to the more financially challenged parts of the society. Gangsta and sexually inuendo videos in neighbourhoods that although poor do not have a crime problem and young people don't dress like that. But it seems the algo is working on it. Note: This is in a European country
Last week it decided that I wanted to see rap videos and filled my feed with them. I don't listen to rap. I've never watched a rap video. I've told it I'm not interested in every single video, but it still puts them there.
I pretty much only YouTube in Incognito mode these days. Everytime I forget, YouTube manages to annoy me so much within half an hour that I switch to Incognito again.
There was a time a couple years ago I watched one episode of Arthur for nostalgia purposes and it completely obliterated my recommendations. I scrolled for hundreds of videos and didn't see a single non-arthur one. It took months for the site to become usable again.
It’s funny you mention that.
For many years I would get recommendations for Latino lgbt books. Weird as it’s not really my topic. I found a way to look at how they generate recommendations, and it turns out that I bought a book about a gay Chicano growing up for a college class in 1998 or something.
I don’t think the user meant to imply they had gained any special insight into the algorithm other then personal confirmation that purchase as far back as 1998 are still being incorporated.
I don't use online music services, I discover music on various platforms and and download mp3s and keep a local library.
I avoid Youtube at it's defaults, I use 3rd party apps and VLC to do most of my watching, other than my subscriptions I tend to skim the Home page very rarely.
I do not use Netflix or other streaming services, I try to hunt down DVDs/Blu-Rays and prefer ripping them for my personal library.
My only problem is exclusives, as a fan of The Witcher series, I do feel like I am missing out, but if I feel a really strong urge I can always borrow an account from a friend, create a temp profile, watch the series and delete it.
Their convenience features just add more inconvenience to me.
It's insane that I feel it helpful to take on "multiple personalities", but there it is.
Part of it is that these algorithms are fairly one-track. They can mix it up a bit, but it's always too much of one thing and too little of another. They can't truly comport with the reality that someone can have multiple interests and tastes.
You should be able to trust a popular open-source project, because you trust that there are folks who will go in and take a sticky beak at these kinds of things. This why a lot of security folks prefer open-source software over close-sourced solutions.
Surely, if Chromium were cross-correlating profiles, then it'd be on the front page of Hacker News in short order.
I'm unaware of any potential problems relating to fingerprinting and so on, but for your use case it doesn't sound like it'd be a problem.
https://support.mozilla.org/en-US/kb/containers
https://support.mozilla.org/en-US/kb/profile-manager-create-...
https://www.huffpost.com/entry/data-and-goliath-digital-surv...
I can recall in the mod-late 1990s when Mainstream America was starting to go online for the first time, everyone was training their kids "never give out your real name." From that simple bit of stranger-danger paranoia, we built a lot of communities as psuedonymous by default-- your AIM "screen mame" was rarely your given name, you could have different usernames on each forum, your email address probably referenced your favourite sports team or anime character.
This inherently constrains aggressive "passive" personalization. Without an obvious canonical identity, you don't want to try to cross-profile too aggressively, because user "hakfoo" on site A may well be different from user "hakfoo" at site B, and you had to assume that any ID you tracked was limited or transient: when you go off to college or apply for a job, you're probably not going to want to be slapping "PonyGirl1987" on your resume.
I wonder if it's that the algorithms are limited to being one-track or if they're overoptimized to being one-track though. I assume they have a profile somewhere that looks like "12% coin collecting, 31% travel to Paraguay, 9% 2004 Ford Focus Repair, ..." and then they offset that data with what content produces the best revenue/engagement/metric of the week. In the process, many secondary interests simply get demoted to the pont you see nothing but tire change tutorials.
The problem was (almost) no one was making money off it. Online advertising, as it was originally done, was a joke. Most of the early internet was based on ideals of quality, community, and freedom. None of these make much money.
So people started hunting for what DID make money, and they discovered data harvesting and targeted advertising. And like a cancer, that unholy pair devoured most of the web. Information really WAS power, like all those breathless articles and hacker manifestos in the 90s said. Power and MONEY. And the new masters of this realm have dedicated themselves to taking as much of YOUR information as possible. They need it, like Elizabeth Báthory would have needed the blood of virgins if she was an actual vampire and not just insane.
Realistically, on the modern net you need to: Use a VPN, kill cookies and ads aggressively, have multiple accounts (or avoid logging on at all in some cases), and never, ever use 2FA except with financial institutions. That's the bare minimum. No point in just complaining that tech companies are evil all the time. The smart organism adapts to its environment, and uses the tools at hand.
I eventually realized I could skip around and enjoy the parts I liked. Freely. Without that feeling of being watched.
Flashback.
One precedent is Nielsen Audio's Portable People Meter. https://en.wikipedia.org/wiki/Nielsen_Audio#Portable_People_...
I assume that if surveillance can be done, someones are doing it.
Knowing this, I leave it on because I want everyone to be forced to watch what I watch.
Is Roku actually capturing info about the content of the stream, or just the fact that I'm streaming from the PC? I'm fairly sure they aren't so blatant as to grab screenshots, but are they doing some sort of analysis of the video stream directly?
Just stuff I wonder about.
This is exhausting.
It would just be nice if they were configurable, understandable and actually worked for me. Some systems’ “We’re showing you this because ...” is an important first step, but it needs to go way, way, further.
Big companies probably don’t do it because they’re afraid of overwhelming a user, disclosing too much about their own algorithm and the things they know (“We’re showing you this because someone you hung out with on Instagram just before you met”).
Thus it’s probably up to open source projects again to make algorithms that actually work for people, not against them.
Are there any such projects already? Are there projects trying to dissect common proprietary recommendation systems? They’re some of the most mysterious influences on current society.
- Control: what you see is based on your explicit upvotes & downvotes, not what you happened to click on.
- Transparency: you see content from feeds that posted content you upvoted and from other users who upvoted what you upvoted, but did that before you. For example, when you upvote a link it will tell you that "You will get more content from 2 users that also liked it and from 3 feeds that posted it". And when you see a recommendation from other users you can see what likes in common do you have with those users (“We’re showing you this because ...”).
- Fairness: the amount of attention every user and feed gets from you depends on how useful their past recommendations have been. Ie, the higher their signal-to-noise ratio has been for you - the more prominently their other upvoted items will be in your "feed".
Give it a try and let me know what you think. I'd like to do a "Show HN" post for my project soon and your feedback would help me prepare for it.
If you like to read a bit more here are my announcements with discussions:
- https://tildes.net/~tech/u7f/linklonk_a_link_aggregator_with...
- https://www.reddit.com/r/RedditAlternatives/comments/mpqnpl/...
There are of course other challenges preventing people and corporations from entering this field but at this point I don't even think there's hope. They have not only destroyed everyone else in the private sector, but they partnered with governments and political entities in such a way that their continued rule is basically cemented in place. You won't get the recommendations that you know would be good for you because they're not trying to do something good for you, they're trying to extract more value from you.
If you spend a minute too long looking at the menu for a movie it will not only auto play, but for the foreseeable future Netflix will assume you’re watching it, want to finish watching it and watch similar stuff.
More than most other services Netflix doesn’t trust your ratings or your lists when making up recommendations.
The discovery is specifically what makes Youtube, Netflix or Spotify sticky beyond just content hosting, and even at suggestions, pirate websites are better at it.
And yes, they had to get rid of ratings, because some people's feelings got hurt because they were being rated so low. Can't have that pesky objective reality rearing its ugly head.
I use YouTube via RSS. My quality of life is incredible.
You could also just put the channel URL into your RSS/Atom reader and it might convert it automatically.
I wonder how many features like this exist that I could use if I knew about them.
Does anyone know if Netflix/Hulu/etc. have an accessible database of their content? Most of my search results reference a Netflix API that seems to no longer be accessible.
The book, written in 2002, was very prescient.
I am not afraid, but annoyed to not have "manual mode" internet where everything I do has no hidden consequences and is simply what I asked.
The only “algorithm” that sort of work is Amazons book recommendation, but I’m not sure that not just based on what others have bought.
I had to reset all suggestions to get away from it. Whenever I now see a video like that about the culture wars (or about covid for that matter), I open an incognito browser and search for the title. It's nuts that this is necessary.
It's come to a point that I now hate suggestion engines with a passion. I wish I could simply only see the tweets/videos/whatever from people I follow/subscribe, chronologically. I can't figure out whether the people making these services are simply incompetent or downright evil.
The AI future is now and it's awful.
I find Spotify similar. I try a suggested playlist. I check out after 2 or 3 songs. It doesn't get the hint.
Anti-woke/alt-right type content is provocative, controversial and makes people engage with. Either because they want to enthusiastically join in or because they want to dunk on it for how ridiculous it is. And the people who are into it are generally really into it, so they'll keep watching hour after hour of it.
YT doesn't care, as long as they get as many people watching videos and looking at ads as they possibly can.
Maybe that's the effect and not the cause.
Personally I find it easy to dismiss bad suggestions by looking at animal videos or vtubers, those reliably displace everything else. And that everything else was just strange Minecraft roleplays for small children, not racism lately.
And to counter all this, often the system is actually surprisingly insightful in digging up obscure crap from god knows where that's really neat, and right up your alley, that you never would have seen otherwise.
And then it goes back to filling your whole front page with Simpsons clips.
I've never opened any of these links, am not behind a VPN, am very much not Arabic and do not understand the language, and - with all due respect - am completely uninterested in Islamic holy scripture.
Why, even after all these years, Google thinks this is content I'd want to see is a complete mystery to me. Something I'm going to have to live with for the rest of my life it seems.
I have kept a log of all IPs I regularly connected from, and only a single one of those originated in another country - another European one.
Never lost a device (jinxed now) or bought or even temporarily used a second-hand one.
An ad or intrusion do seem to be the most likely culprits, but ads are blocked everywhere, and I believe my security hygiene to be pretty decent.
This has been going on for over a decade now. Due to the sporadic nature if these happenings, I'm almost starting to think there's some old (but still wrong) data stuck on some edge node somewhere - or perhaps someone at Google is actively teasing me for being so critical of the company :)
yep the hyperparameters for the search suggester model need a lot more tuning
https://www.nbcnews.com/tech/security/can-government-look-yo...
That isn‘t even necessarily bad! It‘s just problematic that it could easily be used for immoral surveillance. It‘s imo a good thing to identify likely future terrorists. It‘s not a good thing to surveil members of e.g. sexual minorities or peaceful dissidents.
I‘m a recent immigrant, so I am sorry if my English sounds weird.
They only stopped because they got caught, and enough Google employees caused a stir about it internally. The executives who supported this project and kept its existence secret within Google are still with the company. Who's to say they won't try again, in China or any other country?
https://theintercept.com/2018/11/29/google-china-censored-se...
>Locating core parts of the search system on the Chinese mainland meant that people’s search records would be easily accessible to China’s authoritarian government, which has broad surveillance powers that it routinely deploys to target activists, journalists, and political opponents.
Google search, map and translate display a pop up. Youtube search displays two (!).
Bing translator is the most dishonest - for anything else than a trivial sentence, it requires a captcha to be completed, pretending that it detected abnormal traffic (I know it's a lie, because if I don't use private mode, it doesn't do this).
The services above essentially defeat address bar (keyword-based) searches.
It's interesting (and not in a good way) times. I've changed almost all the services I use, which is something I've never considered before. I have to say that there are valid alternatives, though.
If there's any reason to believe that Google tries to defeat incognito mode and associate that history with a particular Google account, I haven't heard it, and it seems really unlikely, given the things people are expected rely on incognito mode for.
I don't think Google tracks these things in regular search either. My interests profile doesn't include my favorite porn.
I have to agree with [1] that Google is completely in the right here.
I believe that they're just trying to make life harder for those who browse in private mode, in order to push them to always stay logged in. After all, Google built a browser based on the concept of user being logged in (this was one of the primary reasons why I've moved from Chrome).
This is nothing new, as at least some news paper had explicitly rules against it, for a long time.
On an extreme and hopefully far fetched scenario, I image this becoming common practice, and Google says that you have only a certain amount of searches before you must log in. I think the current situation as a sort of middle ground (or step towards).
> I know it's a lie, because if I don't use private mode, it doesn't do this
If you don't use incognito/private browsing then you have an ordinary-looking set of cookies, so it makes sense you wouldn't be challanged with a captcha.
You'd think there'd be a plugin that gathers some common tracking cookies, and swaps them around with other users. So the adtech firms see a tangled mess of location/IP/interest data.
If the randomization is log enough, the cookies pollute the data while hopefully remaining "conventionally looking" enough that they won't be discarded outright.
Right. I find it very troubling (in a general sense; I don't doubt that there are technical grounds) though, that browsing in incognito/private mode is a flagged behavior.
So, one channel for how-to's, crafting, and tech talks. One for instrumental and foreign language focus music, etc..
https://twitter.com/jjcollinsworth/status/139066623945075507...
After 9/11 and the Patriot Act, librarians fought the government to keep the reading habits of patrons private as a core duty like doctors pledge to do no harm. Now it's difficult to get people to understand why clicking through to an attractive hammock while browsing Amazon registering in 900 databases and manifesting in ads and cold calls selling tropical vacations and being flagged on some government system as 2% more likely to be a flight risk if let out on bond (who am I kidding, AI doesn't give percentages, it just gives conclusions, looking at a hammock might be that stray pixel that can turn an OCR "O" into a "Q") - difficult to get some people to understand why that would make you nervous.
The answer to that isn't "use TOR to get to a VPN to browse amazon, and pay for that with a burner debit card loaded with bitcoin" or whatever works this week. There isn't really a hammock.
edit: also Amazon has that figured out, some researcher there figured out that they can identify you based on your click patterns and timing. Right now you can choose between 1) an app that will just click on everything silently, or 2) another app by a professor at some midwestern university that will jerk around the timing and positioning of your clicks in a way that throws your fingerprint off. The first app has been banned by every store, and triggers 80 warnings and a waiver that has to be signed with two-factor even if you manage to root your phone and sideload it. You have to compile the second app yourself, with a weird toolchain, and it draws in 640Mb of npm libraries. It's already been updated three times in response to Amazon's countermeasures, and the professor just wrote a paper about the entire method probably being ultimately doomed.
This isn't restricted to just Amazon either; as I understand it, this is the core of the current reCaptcha version (and even more so, its evil twin, reCaptcha V3, which is seemingly so effective it can rely on those mouse events exclusively in most cases)
As a dev, I google a lot to figure out how to do X.
My experience with DuckDuckGo was that it added 2-10 minutes of filtering for usable results every time.
Do you have an example of a search you've done where the results were much better from Google Search than from DuckDuckGo?
Due to the very topic this HN post is about, I don't think this is a good method - your results might be vastly different from mine.
But let's look at the number of relevant results among the top 5 results for some thing I recently wanted to learn about.
"Comprehension categories" - a rather specific term from category theory
https://www.google.com/search?q=comprehension+category https://duckduckgo.com/?q=comprehension+category
Google: 5/5 (for me) DuckDuckGo: 1/5 (for me)
"comprehension categories category theory" 5/5
I'd hardly say it's 2 to 10 minutes extra work, especially for such an obvious example
Here you go: https://imgur.com/a/CssDXS3
I use stimulating music (not relaxing) for work mainly, playing in background, and don't want to keep finding new music. Novelty is an important factor in stimulatory music. So this has been great for me.
(this is a HN joke)
(but I did start building my own music player)
Which begets an arms race of proxies, and fake proxies sponsored by Google, fuzz proxies blowing up the signal-to-noise ratio for your account, TOR proxies for the "double-hush-hush" searches. . .
Who knew that the act of finding stuff would be such a voyeuristic delight?
But to what extent would the algorithm know about you if you clear cookies/storage data though? I also imagine changing IP addresses help… does Google use IP as a source for profiling?
If you’re logged in to a service on multiple devices you have now linked multiple fingerprints to your account that can be compared against third party data.
A VPN will not protect you from this tracking.
(I’ve plugged this site before, so I feel like I should say I am not affiliated with it)
Gives you the entropy your browser is giving out in number of bits.
On desktop, set Firefox to purge everything when closing the browser. Use Chrome only when necessary to do strictly Google-related things.
Personally, disposable virtual machines on Qubes OS for browsing make me less stressful about that.
Weird, can't reproduce on mpd/ncmpc. (stop using hostile software!)
But seriously I was just explaining to some friends how my use of Audioscrobbler two decades ago has stuck with me, in that sometimes when I jump around in a track I have a quick thought of whether my listening to that song will have been counted. Even such a simple dynamic created long lasting behavioral effects, and these days we're all just swimming in such external context with every piece of remote code trashware we're goaded into using.
This is a quarter step away from voodoo and south seas cargo cult rituals.
Without the protective ideological shield, that reaction would have been deemed "hostile" and "toxic" by the twitteria.
I was very curious about how this is evidence that women specifically get berated on the internet rather than evidence that everyone is a potential target. Obviously from this interaction alone you can't tell if it happens more to women or not but... I thought it was kind of funny to say "this thing happened to me as a man and now I really understand that this thing happens to women."
I can't know whether his comment was reasonable without looking through the comments he was referring to (and neither can you), but even given that he was being unreasonable, you chose to be unreasonable in turn.
Nah, I can. It seems quite toxic regardless of what has been said to him.
> even given that he was being unreasonable, you chose to be unreasonable in turn
I don’t see what is unreasonable about shaming anti-social behavior. When you see any “unsolicited advice”, even between two same gendered persons, as a manifestation of patriarchy you are bound to come off as toxic.
If you must use gmail, use imap. The only other “sticky” service they offer (that benefits from a login) is docs. So, use some other office suite, or keep a separate browser installed just for docs.
Problem solved, I think.
That would beat browser fingerprinting as long as browser A running in VM A looks different enough from browser B running in VM B.
The average browser is unique out of like 300 thousand other browsers but with the right settings you can get that down to more like 1 in 10.