Maigret: Collect a dossier on a person by username from thousands of sites
github.com
github.com
I've worked on search engines, and depending on who is using them, that balance gets struck differently on the ROC curve. For legal matters, for example, they want every record that might match. For ad-hoc (google style) queries, nobody reads the second page so you care more about Precision @ 20 (or really, at 3)
Respectfully disagree. You're right in principle if you build a tool like that for yourself. But since this is Open Source, you have to take into account that people who don't understand that will use the tool as well and then use that as "evidence" in whatever arguments they're having with someone.
Respectfully disagree with you respectfully disagreeing - this line of thinking can be used to argue against almost any information sharing.
i.e. Should we stop governments releasing statistics that might be misinterpreted by an uninformed press? Should we stop open access to medical journals because untrained readers might use them for incorrect medical advice? Should we stop companies releasing public annual reports, because investing consumers that are untrained in reading financial documents might misinterpret them?
> But since this is Open Source,
You can add the functionality you would like it to have, and share with the others.
For more fine tuned solutions, there is always a DoD contractor available near you :-)
I find myself reading pages 2-5 quite often, because page 1 just didn't give enough results, and I doubt I'm that much in a minority ?
(I'm talking about actual generalist searches, not people that use a global search engine as a replacement to bookmarks or directly searching, for instance, Wikipedia.)
If I'm looking for a specific issue I'll sometimes try out things as deep as 10 or more pages of search results if nothing on the first 100 ish hits selects the issue, but then only if I can't think of any keyword variations to use that might get me a better result match. I don't expect the average user to go even remotely that far.
And many queries that are solved with 0 visits, thanks to the information being pulled out from the wiki, or giving up and rephrasing the query.
After those, then yes, most queries will be answered by the first page. It's what retrieval is optimized for.
TLDR: on the second results page, each result gets <1% click through rate.
So it‘s not nobody, but statistically speaking not very many.
The question ought to be "conditional on not finding the result on the first page, how likely is the user to go to the second page versus balk, or re-try a different query?"
I'm fairly confident that number is higher than 1 %, but I don't have the data.
The end result seems to be that this tool decides you're interested in "dating", "porn", "stocks" and tags you with a "ru" country code - despite not owning any of the accounts that the determination has been based off of.
The README should come with a disclaimer really.
All in all, I personally feel like it is a good thing to cycle through usernames throughout life.
—BuyMyBitcoins
Joking aside, there's definitely value to rotating usernames frequently. I've started using random strings on various sites because I really don't see an up side (for me) to being trackable from site to site and definitely across time. (I use very long random strings for my banking usernames because I don't trust them to have enough bits of entropy in their passwords.)
A Internet comment from 10 years ago might cost you a job, or it might cost a friendship. But at least I’m not living a facade about being a perfect and flawless individual, and that helps me sleep better at night.
The most hilarious/sad was the insurance provider domcura which advertised how they got some kind of award for their great processes, yet writing to 2 or 3 different service emails that they are sending me confidential documents resulted in nothing until I wrote to their data protection officer.
https://en.wikipedia.org/wiki/List_of_the_most_common_surnam...
Curiously, for a supposedly common name I have never once met a John Smith.
In Italy - for the record - it would probably be Mario Rossi (as an example on a mockup form), but conversionally it would be Pinco Pallino.
There are a few dedicated English wikipedia pages:
https://en.wikipedia.org/wiki/Placeholder_name
https://en.wikipedia.org/wiki/List_of_terms_referring_to_an_...
https://en.wikipedia.org/wiki/List_of_placeholder_names_by_l...
it is interesting to see the slighty different use in the various languages.
It was a recipt for a medical purchase, at first I thought I was getting scammed. What tipped me off was the email was sent to firstnamelastname@gmail.com and NOT firstname.lastname@gmail.com. That was the day I realized google would even do that.
I ended up using the phone number in the email to contact the person and forwarded the email. And yes, they had my first and last name :)
He got a congratulations email on a bmw purchase one time. Had a good discussion about cars, we are both gear heads.
From those emails, it's nice seeing how people support each other and apparently I'm in demand giving scripture classes.
But my e-mail address has been used by real people to subscribe to services in Sweeden, Turkey and somewhere South America. At least language helps to sort things.
Dots don't matter in Gmail, so these email addresses are the same:
bonus point for checking the activity to see if it looks human or if it's taken over by a bot.
> Maigret collect a dossier on a person by username only...
The About field:
> Collect a dossier on a person by username from thousands of sites
"on a person" seems to imply that they'd belong to someone in particular. Obviously if you have any experience of creating accounts you'd know that's unlikely to be the case, and it's not written in a promissory tone. But it does imply it.
I often use the same username(s) on multiple forums where people discuss similar topics because I want other people who visit the same forums to recognise me as being the same guy.
I've found this username taken a few times, and it has bugged me every time. While it is from Tolkien's Elvish, it is obscure and then misspelled (actually a portmanteau). I've had to add a prefix, such as "TheReal" to it.
More rations of rum, i think thats how the british managed it. I wouldnt complain, much.
Does anyone know the story behind the fork?
I wonder how hard it would be to add the functionality: go back to the email address that has registered these accounts, find any other names they've registered, and search off those. (EDIT: err, wait, I bet that's the "recursive" functionality they mention).
[1]: more accurately, you didn’t go through the list of all vendors and didn’t object to their ”legitimate interests”, if that was even an option.
I also also create one-off usernames in addition to my 'permanent' name. That way everybody sees the permanent name and assumes I'm too stupid/ignorant to use different names, so nobody ever suspects the other names are me. (Using different emails because I do the same thing with email addresses and - when I can afford it - phone numbers).
You must live in a location where you have very little fear of political oppression. If you're an Iranian or something these days the situation is very different.
My situation is a bit unusual in that I almost certainly have a globally unique name (my last name is a corruption of an uncommon name in Norway; said corruption occurred two places independently, and there are less than 500 people with that last name in Norway, and probably about the same in the US, and to date I've seen no indication of anyone combining it with my first name), my first name plus last initial is also unusual enough that I only rarely can't get it as a username.
I use the same username (which is my real name as well) pretty much everywhere.
Or don’t, I’m not your mom.
`correct horse battery staple`
just wondering how memorable that technique is.
edit: indeed
This seems like a great tool for stalking people; particularly the recursive functionality (for tying identifiers together).
I'm not saying the world is worse-off due to the release of software like this - indeed, you could argue that publishing these sorts of OSINT tools allow folks to take a proactive and protective view of their own information and e.g. make changes to their own profiles/privacy settings.
But the question of ethics seems like something all-too-often glossed over in the infosec world. Software doesn't exist in a vacuum.
The only difference here is that one is more limited and open source.
"Privacy is dead, get over it."
I realize this may be a failure to communicate because of a difference in shared values but I come from the perspective that harm reduction is worthwhile even if you can't fully eliminate the problem. That you can't prevent all crimes doesn't mean you shouldn't try to prevent crime.
Not to mention that the risk profile of having three-letter agencies come after you or having a random obsessive weirdo "collect a dossier" and share it with other obsessive weirdos who get a kick out of making your life miserable because they project their life failures on you is very different. Namely, if you're prominent enough to have an entire state agency try to destroy you, it's probably over an active decision you made (e.g. focusing your life on political activism) whereas for random weirdos to harass you and attempt to drive you into suicide you just need to show up on their radar long enough for them to start making up reasons to hate you.
The difference is that the three letter agency has a budget, process and middle management so at the end of the day someone needs to justify why they're expending resources on making your life miserable whereas the random obsessive weirdo just needs to convince a bunch of other obsessive weirdos that you're a garbage person they can turn into their "lolcow" of the day/week/month/year.
The three letter agency is more likely to inconvenience you out of apathy than to actually try to destroy you intentionally. The random nobody is more likely to try to drive you into suicide for a thrill.
I have however, encountered death threats or obsessive/hostile behavior from people online who decided they have a beef with me for one reason or another. I'm much more inclined to be wary of some unstable individual having the means to find my identity (and therefore my location)- potentially using this to act on their emotions IRL.
In reality, as mentioned by other commenters, I'm not anything they're interested in, at best I'm a datapoint that advertisers try to categorise into a demographic (male, 30 - 50, doesn't give a shit about cars, into tech).
And anyway, they're already doing it, but impersonally.
It's the person who is very personal that I'm not enthused about.
Not only are you happy to never challenge current power in any meaningful way, but you are sure you never will be. If there is a revolution tomorrow and the Trumpists are in charge, or the Marxists, or the Russians or whatever, you're still sure you will be fine with that order too and will never rock the boat.
You're also confident that if you do nothing wrong by them you have nothing to worry about.
Instead, you've been taught to fear the weak, the nobodies. The nameless savage in the night. As if they have any more reason to hate you than the powerful have.
Similar tool for phones -> https://github.com/sundowndev/phoneinfoga Collection of OSINT Tools- > https://osintframework.com/
And if you want to pay you can just use https://www.maltego.com/ and https://pimeyes.com/en
It is readily available.
You can't change the NSA, but you can resuce the amount of abuse.
It's the same with guns, USA vs Europe for example.
I am NOT getting over it.
I always think it’s strange when people make the argument that it’s good for harmful tools to be out in the open because people can/will change their behavior because of them.
99.999999% of people will never see this repo, never know things like this exist, and don’t visit sites like HN but can, and probably will be, affected by it in some way.
> 99.999999% of people will never see this repo, never know things like this exist, and don’t visit sites like HN but can, and probably will be, affected by it in some way.
Yup, I do think you're right. And honestly, the folks who'd probably most benefit from being able to run one of these tools to empower themselves with data that can help them change their behaviour/privacy settings - they're probably actually the folks least likely to be able to install and run the tool.
For an analogy, consider another tool, an app you can side-load and use to unlock any Prius made between 2010 and 2013. Is this beneficial? Certainly some affected Prius owners will be helped by its release. But affected Prius owners will be (probabilistically) harmed. Even if you're one of the group that was helped (probably arguing the exploit was already well known prior to the tool's release, perhaps?) I think it is dishonest to assert that the release of the tool is "good" from a utilitarian sense.
I wish it weren't so.
I'm not saying you're necessarily wrong in general. It's possible this is bad for the world. But I maintain that for myself and the people I care about, it's more useful than it is damaging.
It will be used for many reasons, but the stalking is the most obvious.
Not all of my usernames are as resistant as this one, unfortunately.
It's essentially like arguing that releasing open-source ransomware toolkit is beneficial. I mean, maybe it's your right, and one can make some strenuous arguments about how it helps the "defense", but really, it just makes it easier to be a terrible person on the internet.
If this wasn’t on GitHub, we wouldn’t be aware of it. However, it would certainly be circulating on underground forums. Is blissful ignorance better?
Granted there the solution was a rather simple one but I feel it's at least worthwhile for more to be conscientious of singular identities online and what info is disclosed publicly with them.
[1] https://en.wikipedia.org/wiki/Firesheep [2] https://threatpost.com/facebook-kills-firesheep-new-secure-b...
If this truly causes worry, adopt better opsec and create generative usernames for different sites. If not, assume anyone will easily be able to link your bowel issue subreddit comments to your LinkedIn profile.
The more availability this type of tool has, the less professional its users are. As a result, while it makes easier for you to see what “they” can see about you, “they” become much more personal and bitchy than credit bureaus or ads companies who already have it. E.g. a credit bureau would never sell your comments or “private preferences” to your boss to step over you in career. It increases attack vectors enormously.
Is blissful ignorance better?
It is in this case, I believe. It’s like spreading free covert time-travel-enabled surveillance devices among general public. Someone will pat it onto your back just for dark fun.
Also, if we don’t raise issues like this, the next actual fix will be “don’t reuse writing styles and vocabularies on different sites”.
Realistically, there is nothing you can do to avoid that future. The best thing you can do is choose to accept the risk or act to prevent the issue.
This tool could be used to teach people, e.g. in OSINT challenges, it could be used to gather information for a pentesting job, it can be used to teach people about best practices online etc ...
Or just... look yourself up. Like most of the people on this HN thread have done.
Anything you put online is perpetual and will be used against you, any innocent hobby will or adolescent joke will become a major transgression at the right time in your life and for the right audience. Even slight grammatical idiosyncrasies in the words you type can be used to root you out by motivated parties.
Curate you official public profile like you are preparing to run for office, one day you just might; anything else should be anonymous.
On the other hand the docs are the right place to warm about potential legal implications.
Ethical decisions are down to the individual.
Unfortunately, lots of people think ethics is what the law allows, or what government says, or even what teachers teach. Most individuals don't take the time to realise or uncover and then act upon the ethics they have innately - aka following one's heart/conscience.
Worse still, many paychecks depend upon the unethical, so don't expect any critical introspection or changes soon.
My opsec is pretty awful but I at least try to not be "simply type in the username I always use into google or some tool on GitHub" bad. I at least want somebody to have to use a few hours or pay a data broker to reveal all my secrets because my bet is that nobody cares about me enough to do that.
I mean of course it could work, but except if there is a very special interest for an account, we may need to wait some decades before a list of "linked accounts" containing it in particular could emerge, right?
can't find a comparison of them though; any idea how they differ?
Comparing the options described in their readme files, maigret adds --html, --pdf, --tags beyond what sherlock supports. And, it adds documentation beyond a README file.
So many people with empty blogs and the same user name. Unless the blog is auto-created, I don't see the point.
I do find all the .ru tags to be concerning. I can imagine out-of-touch security types at Paypal deciding to sanction me based on it without any recourse.
Interestingly, searching by e-mail gave the most garbage results, likely because most services don't expose the e-mail address directly. I guess it's an advantage that neither my usernames nor my legal name are particularly unique.
I'm seeing stuff I probably need to go clean up from like 10 years ago that's no longer relevant
Turns out that one of the websites I'm registered to uses Gravatar to display profile pictures. I've naively been using the same WP account to set a picture for all my emails, believing that you couldn't get information about my other accounts using a simple picture.
I guess I was wrong, and I'll probably split my Gravatar account this evening.
An even better trick is to invite them for a podcast recording. Then all this research is suddenly very normal and expected.
Deeply unsettling. Having had my share of encounteres with people online taking personal interest in an unfriendly context, this is worrying. I don't need to run a search of my info to know to use different usernames, but an automated feature that allows just about anybody determined enough to easily find not only linked accounts but archived data. I'm glad I no longer use social media, at least infrequently and not the major ones- but the fact someone can find archived information? Time to make some adjustments.
My operating assumption of Palantir as a company is that they have a very advanced system of a similar shape that tries to accomplish the same goal of linking accounts across services, but I have no insider knowledge.
But my idea was more along the lines of supporting API keys for many different sites, taking your time to configure all these API settings for all these different "site scraping engines" and that would give you the activity of someone you might be tracing. Their comments, their posts and so forth.
I was thinking this would be great for snooping employers who want to know what their WFH employees are up to.
Funny enough I was even imagining support for Steam, and friend request coming from these robots that in rare cases you could even see live as an employee started playing a game.
It should provide at least some numbers/graph to "measure" your presence on the web, like I do not now, number of GitHub repositories, instaram photos and so on.
A user with little instagram photos but a lot of GitHub repositories is different from a guy with a lot of Blog posts, Pinterest entries, Twich videos...and so on.
It's nice for an initial scan of a target but I expected a more comprehensive report. It's a good start but it could use a lot more.
There are various fantasy name generators, for D&D and the like.
And then give a nickname for account in password manager. And communicate your username to your friends through other channels if you need to.
https://en.wikipedia.org/wiki/Jules_Maigret
Terrific detective books
We also scraped individual reactions from social media apps to get a _very_ detailed profile on what they engaged with (like using the "Angry" reaction emoji when Trump said something stupid vs using the "Angry" reaction emoji when someone AOC said something stupid).
Never released it in the wild for obvious ethnical reasons, but was an interesting technical challenge. Also led to super interesting insights – like learning that videos and text links were watched by entirely different audiences on Facebook and Twitter [1]
[1] https://twitter.com/rishdotblog/status/1483329729302515712