Opensubtitles.org breached – Email addresses, IP addresses, Passwords, Usernames
forum.opensubtitles.org
forum.opensubtitles.org
edit: at the time I also was young and naive [0] and thought the open part of opensubtitles meant that people submitted subtitles in their free time out of their good heart and the site was simply hosting and organizing them, no problem with displaying ads on the website or at the beginning of the subs I thought, to make up for hosting costs. Subs hunting was so much fun at the time /s. But having to register to "prevent abuse" was a tad too much.
[0] still am by the way, thanks for asking :)
This was actually optional, and always was. I think the issue was that they had made some changes in the backend, but the addons didn't update the code to accommodate. At least this was the case for Kodi. The opensubtitles addon has essentially become abandonware anyway. There are now third-party addons that provide the same functionality but without a login, and across different subtitle websites.
anonymous, scrapes a handful of sites including opensubtitles + subscene
Oh, well. I'll give a try next time I launch kodi, thanks for the tip.
now on the website I would hit the limit very fast, but through VLsub it's essentially unlimited
I googled vlsub opensubtitles not working and I still find relevant threads https://forum.videolan.org/viewtopic.php?t=152485
I have been using opensubtitles since they opened the site though and I couldn't really tell when the problem appeared and disappeared.
and yes, it is still sometimes freezing or has issues, but not sure what you expect from service you don't pay anything for, not even in forms of ads if you use it through VLsub
It's from my boy Andrew from the anti-vlsub-opensubtitles cell #3948 in Wichita. Thanks for outing us, we have to rebrand our whole operation now and update banners and usernames, etc. We were so close to convincing everyone to stop using vlsub and opensubtitles :(. /s
> so not sure what's your deal but you are clearly trying to mislead other based on maybe many years old experience, while reality for many years is that nobody asks you for login/pwd in VLsub
So, you don't have to login unless you have to login.
Got it.
Also, this:
https://github.com/exebetche/vlsub#2013-09-05-version-0910
> 2013-09-05 (version 0.9.10)
> Add possibility to set opensubtitles.org username/password in config menu to avoid download limit to unlogged users
You can add this to the list of reasons you needed username/password for vlsub at some point, when you hit their download limit. Which is something that used to happen a lot with fansubs.
or you can completely delete whole string including timestamp, it's not like any hardware of software videoplayer cares about missing number
Dustin has now locked that GitHub to only previous contributors so users are once again left voiceless and powerless in the continuing war against their privacy.
We've been thinking of the business owners and the children [2] for decades now. It's time to start thinking of the users, people like you and me who are exploited constantly for everything they have to be unceremoniously discarded in a waste heap once they've been used up.
[1] https://github.com/disposable-email-domains/disposable-email...
Firefox Relay is a disposable mail service, so it makes sense to include it in a list of such services. I wonder how they handle mail aliases like GMail allow with + or . as separators.
Anyway I deleted my account.
1. The repo in question is simply a list of domains. People can choose to block whatever domains they wish. This list holds no inherent harm or ill will and is not a "war against your privacy", it's simply a list of disposable email domains to do with what you wish.
2. This isn't in any way related to security breaches or disposable email domains and appears to be an entirely unrelated topic.
This much is obvious by spending less than five seconds reading the README plastered on the front of repo. You're also welcome to peruse some of the PRs and issues in the repository if the README wasn't clear enough. In fact, go visit the last comment in the PR I linked: a contributor points to a Hacker News comment for reasoning why operators might want to make use of this blocklist.
I'm sorry, but yes it is related as the operators of repositories like this (and there are others) are deaf to the reasons why users make use of these disposable mail domains. That reason is because breaches like this happen daily and these breaches are a threat both to the public IDs we use for emails and to the private passwords we use, should they be poorly secured.
Wait, what? Definitely lesson not learned:
- sha256 is not the proper way to store passwords, it's still vulnerable to the same attack as md5, rainbow tables, because it's a FAST algorithm (sure md5 is also poor for collisions, meaning it's worse, but practical attacks for lists of hashed passwords are rainbow tables). At least with salt+pepper it limits attack surface, but instead you should be hashing your passwords with:
- `password_hash`[1] should be used instead of `hash_hmac`, with the algorithm being "PASSWORD_BCRYPT" or better[2]. This is a slow hashing method, meaning anyone trying to rainbow-table attack your passwords will have a hard time.
- A common technique is not to delete old passwords, but instead to rehash them with the new algorithm. This would be useful for moving from sha256 => bcrypt, since collisions on sha256 are not practical, but if the original hash was md5 then I think it's fine to delete the md5 passwords and require a new one. Good luck to those who changed their email in 15 years though.
[1] https://www.php.net/manual/en/function.password-hash.php
[2] I haven't followed the space too closely for 3-4 years, I'm not sure if bcrypt/blowfish is still the recommended algorithm or there's newer better ones
While bcrypt is already much, much better than using MD5 or SHA256, the best practice is to use Argon2.
In practice, it is more important to use a hashing algorithm that is designed for passwords (e.g. bcrypt, argon2, scrypt) than it is to choose the best one. As this breach shows, many sites are still using insecure hashes, like MD5, because it worked 15 years ago.
I haven't had to write any password storage code since before Argon2 won the 2015 competition so I haven't gone deep on the "side channel attack vs. memory hardness" tradeoff question when picking the mode. I am curious which mode others are using.
The specific modes are useful dependent on the threat model you are protecting against. Argon2i protects better against timing/side-channel attacks and Argon2d protects better against brute-forcing attacks.
In my opinion, a developer should (in most cases) focus on application security instead of picking the perfect hashing algorithm, as application security is far more crucial.
Adding long salts on the other hand requires the attacker to create an infeasible number of rainbow tables (one for each possible value of the salt).
And bcrypt includes stores a unique salt next to the hash for every password, making rainbow tables completely useless.
EDIT: I mean, I understand maybe why in 2006 you used md5 without salt. But a few years ago I when I tossed together my first webapp (a scheduling system for the community I'm involved in), I just googled "password hash" and immediately bcrypt came up as a recommendation; there was a package in my target programming language, so after half an hour of research and 5 minutes of programming I was done. I don't understand how the opensubtitles team, after having had their password database compromised, came up with "use sha256 without salt" instead of "use bcrypt or one of the many libraries which takes care of all of that for you".
You note that when you googled "Password hash" you got pointed to a decent password hash. But, knowing what questions to ask is half the battle. Too many people figured hey, I should use a cryptographic hash and got pointed to MD5 (or SHA1 or even SHA256) because that's what those are.
The words you should have googled weren't "password hash" but maybe "web authentication" and then the answer is clearly you shouldn't use passwords or any sort of shared secret.
You can have a copy of the authentication database for the toys system I maintain, but it wouldn't help you sign into it because all the information in it is public, a trick we've known how to do in principle for decades, and which works today, on the public Web, with readily available devices (e.g. my phone) and yet, here we are on Hacker News discussing which password hash is best like it's still 1985.
I'm pretty sure it wasn't anywhere near ready to use five years ago when I wrote v1 of my webapp. And googling again, it's still not clear whether it would work on whatever random software setup some of my zealous-for-software-freedom colleagues use. With passwords I don't need to worry: if they can render the HTML+CSS (even the JS is optional, because some of my colleagues prefer NoScript), they can log in.
So e.g. computing a rainbow table for all possible Windows 2000 passwords (up to 14 characters but you're actually only doing 56-bit inputs) in the LANMAN scheme took ages, and produced a fairly large file, but having done so that's it, you, or anyone you give the file to, can reverse LANMAN hashes into working passwords almost instantly.
(Microsoft's LANMAN hash lacked salt, and is stupid in various other ways, MD5 would actually have been a better choice than LANMAN, because you allow arbitrary input passwords and thus good passwords are stronger instead of impossible even though MD5 is much too fast for a good password hash)
Yeah, you could rebuild a rainbow table yourself u til you find the collision, but you have to search every bit of the potential hash space.
[1] https://hackaday.com/2012/12/06/25-gpus-brute-force-348-bill...
What I mean is, the stored data contains which algorithms it is. So they can in their code or configuration change which algorithms to use and how many times it should hash. Then on login they can verify the password against the hash and also check if the stored hash needs to be rehashed against the current set settings, then it can create a new hash from the password the user entered on login and store that in the database.
Then you get automatic hash upgrades to match the current settings of the hashing of the passwords on the site with basically no user interaction (other than the act of logging in to have the password in plain text).
There are lots and lots of options. Seems like they are a php shop, so even moving to any of these open source PHP server libraries[0] would have been a better move.
Disclosure: I work for a third party auth provider, FusionAuth.
I expect the parent's concern came from experience with PBKDF2, where the length is unbounded. It's good to consider possible denial of service attacks: if someone submits an enormous 1 MB password a PBKDF2 hasher can be knocked offline for 60s. Sha256 will quickly crunch that attack to a more manageable length.
https://en.m.wikipedia.org/wiki/Bcrypt
This sounds undesirable to me, so I'd support sha256 and an alternative.
I'd also recommend adding a max password length to any API. No point in allowing million character passwords.
Is that why websites sometimes have low maximum password length requirements ? Ex: must be less than 20 characters.
Websites (like some banks used to) that have less-than 20 char limits for passwords are purely bad security strategy.
There needs to be a limit, yes, but surely something like 100 characters, or even 50, would be more sensible.
Banks all have systems that will stop you from attempting to bruteforce them.
I've run across that at least once and it was a total pain to troubleshoot and figure out.
"A password one megabyte in size, for example, will require roughly one minute of computation to check when using the PBKDF2 hasher."
> If you care about the security of your account use a 20+ character random password and do not reuse that password at any other sites. There are several excellent password managers that can generate and remember such a password for you and make it easy to use. Here's a list: <link to list of password managers>.
> We allow all normal US printable characters in your password: upper and lower case A-Z, digits, and <list of punctuation>. Set your password generator to length 20+ and to use mixed case, digits, and symbols and you will be fine.
> If you think you might have to manually type this password at some point, you can use a reduced character set with a longer password. If you use just mixed case letters and digits make the password 22+ characters long. If you use just mixed case letters make it 23+ characters. Letters all of the same case? Make it 28+ characters. 32+ characters of hex is fine, too. Heck, you can make it all digits if you use at least 39 of them.
> If your password manager offers other options, such as patterns like groups of digits or pronounceable syllables separated by some symbol, that too is fine as long as you make it long enough. Make it long enough that your password manager gives it its highest strength rating.
> The exact upper limit on password length that our password entry fields allow might vary from time to time as we update the site, but will always be at least 64 characters.
NIST says there's no reason to limit password length any more[0]:
"Memorized secrets SHALL be at least 8 characters in length if chosen by the subscriber. Memorized secrets chosen randomly by the CSP or verifier SHALL be at least 6 characters in length and MAY be entirely numeric. If the CSP or verifier disallows a chosen memorized secret based on its appearance on a blacklist of compromised values, the subscriber SHALL be required to choose a different memorized secret. No other complexity requirements for memorized secrets SHOULD be imposed. A rationale for this is presented in Appendix A Strength of Memorized Secrets."
This could also be problematic.
Password Shucking https://www.youtube.com/watch?v=OQD3qDYMyYQ
First off, you'd have to assume the attacker knows the bcrypt hashes are bcrypt(md5(password)) – an attacker wouldn't always know this
Also it assumes there is password reuse, but that the password is strong enough that the md5 is uncracked.
> > I know I'm probably stupid but... how is this different from a dictionary attack? Instead of trying a list of known passwords, you try their md5s. If the md5 hasn't been cracked before, chances are that the password is strong enough to resist being cracked now. – nobody Jul 17 '20 at 14:55
> Because there are plenty of MD5s in the wild that A) just happen to not have been cracked yet because they weren't interesting enough to stand out, but B) once an attacker can figure out that that MD5 is inside a really interesting, high-value-target bcrypt, they might spend a lot more effort to crack that MD5. So it's not just a dictionary attack; it's a dictionary attack of passwords that are currently unknown but might be crackable with additional effort. And that effort is much less than trying to crack that password if it was only inside a pure bcrypt. – Royce Williams Jul 17 '20 at 15:00
https://security.stackexchange.com/questions/234794/is-bcryp...
So the assumption is: There is a breach A of an low-interest target with MD5 hashes and a breach B of a high-interest target with BCrypt(MD5) hashes. As A is not interesting enough, people don't invest the time to crack A's MD5s. But as B is super interesting they will use A as a dictionary source to then know on which MD5s they should invest a high amount of time, as it will help them crack the high-interest target B. Note that no specific user association takes place, like in the presentation about password shucking by Sam Croley (above Youtube link), where usernames/emails of A and B are correlated.
I think this is a bit more plausible than Croley's take on it. Because if I have identified a high interest individual, I would already invest a lot time to crack the MD5 password.
And yes, what you said bears repeating: All of this attack lives in the small space where the password is too strong to be cracked from a simple MD5 hash when you are mildly interested but not strong enough to prevent cracking when you are deeply interested – for varying degrees of mildly and deeply interested. Overall I would like to read about real world examples where this made the difference and how that password happened to fall into that region.
Still much better to have md5 directly in your db.
I fixed something like this just 4 years ago. :|
It's a very interesting attack, highly specific to the high-number of breaches, high password reuse environment we're in that enables at-scale password cracking.
I don't think it invalidates this advice completely. You should watch the talk and eventually add a global pepper (assuming it does not leak), and of course do the final bcrypt(md5(pass)) -> bcrypt(pass) migration upon user login.
https://security.stackexchange.com/questions/234794/is-bcryp...
bcrypt(md5(password) + salt) + salt
the problem with password shucking would be that they just do a bcrypt(md5) over the list of md5 hashes they have and check if they exist in your database.
but if each hash is salted they would need to run every their complete md5 hash list through bcrypt for each account instead of once per database.
I'm pretty sure that it did not take "that hacker" 15 years to find this vulnerability.
Pretty strange forum post.
here are his videos since his website is dead and he is also not active on twitter anymore https://www.youtube.com/user/BranoGege/videos
> In August 2021 we received message on Telegram from a hacker, who showed us proof that he could gain access to the user table of opensubtitles.org, and downloaded a SQL dump from it.
Wait, they got proof in August and release the info only now? Was this because they were trying to be in a talk with the hacker? It wasn't really clear, but it's a long time...
https://forum.opensubtitles.org/viewtopic.php?p=46845#p46845
Edit: typo
Oh yes, what could ever go wrong with keeping already compromised passwords? Plus it was clearly a white hacker that helped them secure the website, the required fee was for this service.
/s
Never saw anyone suggest paying for avoiding password leaks
If you pay them, they might keep their word. After all, how much would they get paid for a list of passwords compared to a ransom payment?
If you don't pay them, they are more likely to either sell or leak the lists.
In this case, although easy to say now, the correct answer is not to create public web applications until you know how to store passwords correctly and host databases securely, it's not like the information isn't readily available.
Risk management is important as there is no way to know what website has any known or unknown security holes in it. (Especially those built years ago)
When possible use password manager with End to End Encryption (E2EE). Maybe Independent Security Audit too.
BitWarden is open source, both server and clients. There's even a third party implementation of the server in rust that seems to be mostly compatible. Both the official and the third party server are self-hostable.
BitWarden runs their infrastructure on Azure, uses Azure managed databases and Azure managed backups. So I'm quite comfortable having my passwords there.
The only feature I miss is that I'd like BitWarden to send me an email with all my passwords, encrypted so that I own a backup of my data too in case the worst happens.
Not just mostly. Vaultwarden also offers server-side features usually only found in paid Bitwarden plans.
I wish this was a supported feature on @gmail.com domain (not the +{string} thing). I like the idea, but keep thinking that I'll be uniquely identifiable on every database, since it's almost always going to be just 1 entry for the domain I own.
Every computer I normally use will have a copy of the file in the case that I can't access my gdrive (which I don't see happening really).
Is this portable, like Authy does it? Last time I moved phone it was a real pain to move authenticators...
https://hacks.mozilla.org/2018/11/firefox-sync-privacy/
I read and understood just enough of this when it was published to know it would make the foundation an object of ridicule if it was ill-conceived, but I obviously rely on implementation details being tickety boo.
There is a separate autofill feature, but that works quite rarely, maybe 10% of the apps support that, but that makes things even easier.
(I've used it quite successfully when it was in a separate app called Lockwise)
(If like me, you find the idea of a password manager not acceptable.)
yes spam could be annoying, but it's not like people use email only to register on one site
It can be tempting to forgo the same security considerations when programming backend applications which aren't ordinarily accessible to the public. I guess this is a reminder why the temptation must be resisted.
Ouch!
I hope they learned their lesson: Security is an ongoing effort.
When I joined my first company in 2010, to my horror, they were using plain text passwords for users
To be fair to them it took till around 2008 for this to become widespread opinion but the signs were on the wall around 2004
The only reason I even set up an account was so that one of my Kodi extensions could hook into it.
I don't really mind that they got breached. I hope they recover easily and continue to offer their service.
And I will not change that password, unless I am denied access. Don't really care if people know it.
Funny, though, they got breached approximately an hour after I went VIP - it's the only way to download an entire season of subtitles, and their maint window correlates to me having spotty subtitle coverage across most of my library. So I gave in and gave money.
Then it's like "invalid password!"
(paraphrasing) "We were stupid 15 years ago and have been lazy ever since" is a slap to the face.
Maybe they should have done something about their platform instead of watching anime with subtitles.
Those people should have their internet privileges permanently removed and the whole site should be burned in a dumpster like the dumpsterfire of security they had running for fucking 15 years.
> If companies like microsoft, facebook, twitter, nintendo or zoom can get hacked, what are our chances as a tiny team to not endup getting attacked ?
It's not about them getting attacked but they weren't the target of a three letter agency Throwing weaponized 0days their way either.
Anything that is remotely considered best practice would have helped:
Like having strong passwords for the "SuperAdmin" account that was compromised. It's called SuperAdmin for a reason don't you think?
Not using unsalted hashes in the first place?
Investing some of your ad revenue in making security updates to a system that was already bad the second it was conceived?
Their whole statement is insulting
15-20 years ago it was fun to sign up for dozens of new services and we reused passwords everywhere. Now every one of these accounts is a liability and a potential vector for bad things to happen. The worst case scenario is when an account of yours is breached and abused for a long time, possibly with legal and financial ramifications, and you don't find out until much later.
A couple of months ago I happened to glance at my Gmail's spam folder by chance, and one email way down the list immediately grabbed my attention. The subject line had *my password* for an old and unused account that I'd forgotten long back and the sender was trying to extort bitcoin if I wanted the account back.
The big issue is with online shopping. The sender has your name, address, email, phone number and possibly your CC number, a prime target for identity thieves.
I'd be careful with this because a lot of services will allow you to reset your password with just SMS verification once they have your number. So while Password+SMS is stronger in practice this usually becomes Password+SMS OR Customer Support+SMS which is decidedly weaker.
With Covid-19 pandemic requests for our API went skyrocket, our servers could not handle so much traffic, so we decided to limit API usage just to authenticated requests for User Agents. In short it means, you have to provide opensubtitles.org username and password in LogIn() method.
We are avare some of user agents doesn't implement authentication for numerous reasons.
What you can do for fix, if your app is affected: - let your users know about this change - implement LogIn() into your app with user authentication - contact us for pricing, so your app can send unathenticated requests (e.g. if you have free and pro version of your app, it is possible free will work only with authenticated requests and pro works as before)
More info on [our forum](https://forum.opensubtitles.org/viewtopic.php?f=11&t=17110)
We are actively blocking IPs, which are sending unathenticated requests
The whole thing has just given a solid impression of being poorly managed, and the way this breach notification is written really reinforces that. It's unfortunate as I would be happy to donate money but... not to opensubtitles, which seems to be too deep in a hole to realistically get back out of it.
Is there some way we could help OpenSubtitles decentralize and maintain a p2p or federated database of subtitles? This could probably prevent many modes of failure that have required to restrict access to the API over time.
Depending on the injection vulnerability data can be exfiltrated, there are tools like sqlmap https://sqlmap.org/ which make it pretty easy to dump tables via injection
Only caveats:
- Github wouldn't like millions flocking to a repo to grab files.
- I'm not sure sharing of subtitles fall under 'copyright violations'. It's a bit of a grey area. (it's just text right?!)
So are books unfortunately.
I would really love if there was an open subtitle repo though. Hell I'd even help make it myself but xkcd/927 and whatnot.
Not sure how up-to-date it is kept, but for me it works reasonably well, except for obscure films.
Looks like I'll just delete my account and never come back then.
Are there any webauthn only sites, besides demos?
[0] https://github.com/bitcoin/bips/blob/master/bip-0032.mediawi...
Do you know how that supports use cases like if someone wants to change their flight from a hotel computer? I wouldn't want to expose a "master password" to a computer I didn't trust.
Given any choice in the matter I would suggest using your own equipment to change the flight (e.g. a smartphone), even if it's less convenient.
Read as: In the last fifteen years we couldn't be bothered to fix our glaring security hole.
I've seen a lot of people on HN talk about how they never store passwords because they don't want the liability, and I've seen some talk about how they do since it's simple. In my experience, hashing with Argon2 and keeping that in the DB (for toy personal projects) has been very easy, especially to how scary keeping passwords seems to have been made.
I guess my real question is: If I'm hashing with an actual password acceptable hash method (salt and all), is there any reason I shouldn't still do this? I'm legitimately interested, because for all I know, I'm experiencing the Dunning-Kruger effect.
You can store your password yourself, sure, but how certain are you that you're doing it "right" and that "right" today will be good enough tomorrow? 15 years ago MD5 wasn't perfect, but it was good enough versus the other common/stupid methods of storing the password directly or storing it encrypted vs hashed. How certain are you that Argon2 is going to stand the test of time for another couple decades?
They just had a breach like 4 months ago
That's a good excuse for it to be like this in 2006. What about 2021?
Why can't it be? Does everything require a money motive? Is it impossible that people want to do something good?
You need resources to deal with that and it's not gonna be a clean fight.
On HN, apparently. See the recent thread on Wordle.
The reality is likely to be: they make a very small amount from ads and user donations that might, if they're lucky, cover the costs of hosting. The warez scene is a subculture and community for the people who participate in it.
It is depressing that HN participants are so often mystified by the idea people might be motivated by intrinsic or social factors to engage in peer-based production (whether legal or not) when so many of the companies they run or work for exist only because of peer-based production efforts like Linux and other FOSS software.
Opensubtitles has a VIP program at $15 a year.
It's quite easy to find the person who runs the site and, according to their CV, this is basically their job. That'd make Opensubtitles a for-profit piracy site, i guess.
Guess what: Non-profit != for-free && Non-profit != for-a-loss.
If in order for opensubtitles to fulfill its envisioned role needs full-time attention, it is legit to pay yourself or hire an employee, paid by the ad and or subscription money. That's what also happens on all registered non-profit organizations.
No, the 0.1% of subs for obscure indie shows that didn't have native subs doesn't change that.
I'd guess about $1-3M/year from ads.
Not to mention having this on GitHub would probably get taken down almost instantly due to copyright infringement.
it's not just ads, but also paid API access AFAIR
nothing wrong with that, just not sure where you come with idea he would bother, if there would not be good money in it