Archiveteam are backing up SoundCloud
archiveteam.org
archiveteam.org
Like every other Unlimited service, because people like this abuse the shit out of it then when it inevitably dies they run around exclaiming "dont blame me brah they said it was unlimited...., shouldn't have called it unlimited"
For reference, I don't store pirated content, rather content that I have a license for, but cloud providers have no way of knowing that. Unfortunately the dispute processes are unreliable, so when it's time to backup a media project, I play it safe and encrypt.
I'm pretty sure that they just assert that the backup is illegal even if it came from a licensed copy. One argument (used by Nintendo IIRC) is that the official media/servers are too reliable to require backups, and thus anything purporting to be a backup is really for another purpose and thereby not exempt under the statutes authorizing backups.
My entire Nintendo DS and 3DS cart collection was stolen in a break-in at my place. You can bet that instead of repurchasing I simply bought a flash cart.
I wouldn't trust such services to respect your rights in any capacity. In fact, I would argue that the "our incredible journey" trend is a form of property damage (a storage rental place can't just burn their store to the ground with customer's posessions still inside).
> Respect copyright laws. Do not share copyrighted content without authorization or provide links to sites where your readers can obtain unauthorized downloads of copyrighted content. It is our policy to respond to clear notices of alleged copyright infringement. Repeated infringement of intellectual property rights, including copyright, will result in account termination. If you see a violation of Google's copyright policies, report copyright infringement.
Given that they cater to businesses, I don't think they could do automated scans.
Well... yeah? If you let people advertise 'unlimited' but they can't actually deliver, you end up with a kind of market for lemons situation.
Not backing up SoundCloud or the entire Internet for $5 a month
Well, it still kinda is - the physical drives have to live somewhere... ;-)
Actually Google only offers unlimited to business and education accounts, so they're absolutely not referring to personal data.
People are too f'in litteral, first
Unlimited === Store everything every created
PERSONAL in the context of the discussion would include documents created by a person in general course of business, my point is very clearly and only pedantic trolls do not understand the meaning of point I was attempting to address.
Business use would be for the business that signed up for the service, not for storing an entire copy of SoundCloud
What if that business is Soundcloud? As far as I can tell, Google places no limits on business users, numerical or otherwise.
-----------------------------
There is a strongly utilitarian argument to not allowing such false statements.
It devalues the products of people that aren't bullshitting you. Say with fake-unlimited the "real limit" is 4TB before they start terminating you, but a different provider provides 5TB of capacity.
Because the former is allowed to outright lie, there is no way for the latter to effectively communicate that they are in fact offering a better product, instead they too have to make a bullshit "fake unlimited" claim to compete. Now because nobody has to actually back their claims with anything, they are infact massively incentivised to cut the "real storage" limits, because it will cut their costs, and they can still keep making the same claims.
Its a market for lemons[1] race to the bottom, and everyone loses, producer and consumer because scamming liars cannot be reliably assessed beforehand. So consumers lose faith in the entire market segment, and providers offering actual legitimate services become unsustainable.
[1] https://en.wikipedia.org/wiki/The_Market_for_Lemons
-----------------------------
If you allow sellers to lie about information-opaque things like this, you drive the entire market to a shittier equilibrium, it should absolutely not be allowed.
It wouldn't be abuse if it was actually unlimited. They uploaded a finite (if large) amount of data to a service advertising infinite capacity.
I don't think OP is arguing that it's against ToS or anything like that. He is arguing that if people just upload stuff in these "unlimited" services all willy nilly soon they won't be unlimited anymore. The price will be the same, but they will drop the capacity to something reasonable. Which in turn might hurt some legitimate users.
I'm all for people backing up anything they feel like is necessary, but they just need to know that if they are going to be uploading Tera or even Peta bytes of data to "unlimited" service without paying much for it they are living on borrowed time.
Companies should be responsible for what they put in their marketing. If someone calling the bluff on the "unlimited" plan makes them "downsize" the plan to actual ~10TB + $x per Y additional TB, then so be it. It probably doesn't change the reality in any way, and at least the company is no longer lying about its service.
Obviously you are entitled to your opinion and again you are not wrong, false advertising is bad, but again if you are intentionally uploading stuff just to upload stuff be prepared to lose most of it when the limits come crashing down. Like the guy who had/has over Petabyte of video on Amazon if Amazon decides "OK unlimited was an bad idea, let's give everyone 10TB" where is he going to put the rest of this 1014TB of stuff? If the answer is "just let it get erased" then congratulations you are the reason why we can't have nice things.
I agree with your general point though. One thing that helps is that there's a certain throttle because of network bandwidth even if that isn't capped or deliberately throttled.
If my alternative service had a much more reasonable offering with a high limit, that 95% of both companies users data would fit in I'm going to get screwed because I refuse to lie and call it unlimited.
That's more applicable when the other party is not being wrong by calling a service "unlimited", when it is not.
(and, arguably, assholes when they inevitably take it down with because "oops we didn't mean, like, unlimited unlimited")
no youre not wrong.
am i wrong?
youre not wrong walter. youre just an asshole!
all right then.
If you want to design a cap - Goog probably has an idea of the kinds of file people store (whether it's pics or docs or videos or sheets etc...), pick a number of those that's impressive and is just lower than what your 10% power users has, and then use that to influence the package size.
2TB of files can easily be turned into "half a million photos" or like their previous ad campaigns, show how it can store every photo from birth to university for your kid or something. Or a love letter every day from first date to goodbye.
If you can't compete on the number, don't compete on the number.
This is why I so strongly prefer services that don't bullshit me. Tell me what you're exactly selling, for how much[1]. If promise to livestream setting your marketing department on fire, I'll pay double.
[1] If it is free, I already know what it costs and am not interested, thanks.
I am sure there will be people defending the car company as really offering unlimited colours, but obviously they have to restrict it to black, because some clients were unreasonable to expect the company to paint their car Neon Vermilion.
Or to put it another way, "You can have unlimited storage, except that it is limited..."
Unlimited is a marketing term used to express simply to the consumer there are not overall limits placed on your storage provided you adhere to the rest of the terms of service.
In the context of data cloud data stroage when Google, Amazon and the rest talk about "unlimited" they are referring to unlimited PERSONAL storage of data you create as a person, this would include backups of your personal computer, photos, important documents,etc
Not backing up SoundCloud or the entire Internet for $5 a month
So if you are a amateur photographer then storing 10TB of photos you took on the service is acceptable, downloading 900TB of music files you do not own, you did no create and have no permission to "as a backup" because the service is going to go under is abuse.
Well, that's doublespeak then, and if someone calls them out on it by actually testing the claim they make, so be it.
I don't understand why people seem not to mind being lied to their faces, as long as it's "just marketing".
Because most rational people use common sense and logic to come to the understanding that when a company is offering you "unlimited" storage for your PERSONAL FILES, they do not intend for you to go out and download SoundCloud as backup in case the SoundClould Service goes under
Just like when a "All you can Eat" buffet does not intend this to mean "All you can eat in your entire life" where by you fill grocery bags full of food to take home with you
How am I a "knuckle Head" in this situation, when I go to a buffet I eat a normal human portation of food inline with price I am charged for the meal
I do not eat 25 plates full of prime rib for $5.
I do not abuse business simply because "I technically can because it is in the rules"
I fucking hate people that look for these types of technicalities to exploit in society. These types of people are exactly why there are pages of Terms of service, and why we can not have nice things, because people can not be trusted to not abuse shit.
Come to think of it, I'm mostly in agreement with you. I do feel there's a difference between "all you can eat buffet" and "unlimited storage" (or "lifetime warranties"). The former is more of a reasonable and well-explained offer; the latter is more of a bogus marketing claim. I detest bogus marketing claim.
2 wrongs do not make a right and all that.
I'm not going to eat 25 plates of anything but, at a buffet, I have no issue with mostly going light on the cheaper fillers. And I'm not going to worry about it if I end up getting a "good deal" on the meal as a result.
Wait what? I mean the cost of admission is usually more like $50 but that's practically the SOP for Brazilian steakhouses. I have fasted for days just so that I could eat more.
Which specific clause of the TOS is this violating?
In the context of data cloud data stroage when Google, Amazon and the rest talk about "unlimited" they are referring to unlimited PERSONAL storage of data you create as a person, this would include backups of your personal computer, photos, important documents,etc
That's your interpretation. Nowhere do they actually claim or imply that you're only supposed to use it for data you create as a person. In fact, it'd be absurd, considering that sharing files is built into the system.
--
That companies lie to us repeatedly under the guise of "common sense", as if they followed the same standard when applying their unreadable TOSs against us, is bad enough.
Corporations are not your friends, and they won't hesitate to block you if you start becoming a liability. Assuming good faith is absurd, defending it publicly is grotesque.
In the United States for instance, making a copy of any digital media is generally illegal, period. The two major exceptions here are if you can prove "fair use", or if you are an archive or library.
"Fair use" is, of course, a very fuzzy term. Fuzzy enough to give companies enough wiggle room to terminate if, as probably most large archives would be, a person uploaded terabytes and terabytes of copyrighted media to their drive. (If said person shares copyrighted links in particular, that usually is explicitly called out in storage TOS clauses... but even if not, I think it would be difficult to claim "fair use" for a personal upload of Soundcloud to your Google drive.)
In the Archive Team's case, it looks like the Archive Team is using the Wayback Machine from archive.org. (http://archiveteam.org/index.php?title=Dev/Infrastructure) Libraries and archives have their own set of rules allowing limited copying (https://www.law.cornell.edu/uscode/text/17/108), in addition to the general "fair use case". My guess is due to questions of Soundcloud's longevity, archiving Soundcloud would qualify.
Think of it as we are all contributing to keeping this data safe and available.
> Not a home connection, just google abuse. GCE has ~40Gbit to each server and >400Gbit peering with amazon
It sounds like they paid GCE for 900TB of data transfer?
Lets assume they pay 0.005$/GB, which results in a 4500$ bill(@ 0.005/GB) or 45000$ bill(@ 0.05$) for Soundcloud.
Actually it's $0.02/GB if your traffic is in the US and you do 5PB/month: https://aws.amazon.com/cloudfront/pricing/
But yes, they likely have a non-public deal.
Even given that they could restrict themselves to songs marked okay to download, how much of that will be DJ mixes containing copyrighted songs?
I'm just wondering because Soundcloud actually has support to specify your copyright terms, which does not default to "everyone can download this", so it's an interesting case..
The website works now, it just says "selective content"
I'm sure SC serves up a lot of content per day, but how do you think they will react by suddenly having someone download all of their 900 TB or whatever it is in one day? How much will Archiveteam be contributing to SC's downfall by suddenly causing them a huge unexpected bill?
As someone who really wants the SC content backed up properly, I nonetheless see how this raises some interesting legal issues.
http://webcache.googleusercontent.com/search?q=cache:eCl5VSB...
In response to your concerns; I think ArchiveTeam simply doesn't care. They are very firm in their convictions, and they don't exactly listen to requests to not archive things.[0]
If you're curious, here[1] is the initial discussion that AT had. People bring up copyright concerns.
[0]:http://webcache.googleusercontent.com/search?q=cache:fmhU2zS...
[1]:http://archive.fart.website/bin/irclogger_log/archiveteam?da...
If you've ever seen a Jason Scott talk, he isn't the sort of guy who gives a shit if you DMCA him while he's sucking up all of your bandwidth archiving your content two days before your servers shut down.
https://en.wikipedia.org/wiki/Online_Copyright_Infringement_...
This is not to say that people don't send "DMCA notices" for anything and everything to anyone and everyone, but those notices are not following the law. (Also, a lawyer can always send a demand letter demanding that someone stop any behavior, but a properly-constructed DMCA notice gives the recipient an extra reason to follow it compared to a run-of-the-mill legal demand—"[a]n OSP who complies with the requirements for a given safe harbor is not liable for money damages", as Wikipedia puts it.)
Not to mention that dying websites or recent acquires are likely to be able to have the legal muster to start threatening archivists.
Obviously, there's more than just the bandwidth cost, but assuming they pay $0.02/GB for CDN traffic, we're talking about $18k. It's not nothing, but I doubt it'd change their outlook in any meaningful way.
I should add that the ArchiveTeam doesn't download everything in a day, but rather uses a distributed crawler (ArchiveTeam Warrior) run by volunteers. They rate-limit the crawling rate as needed in order not to overwhelm the site being archived.
So you should factor in the bandwidth cost for CDN traffic AND origin pulls, which if you're serving from AWS (and not using CloudFront), is $0.08/GB.
The size of your working set also influences your overall CDN bill. Storage isn't free.
The storage cost is a good point, and I don't know which CDN they use, but I expect the songs that weren't "hot" would be evicted from edge caches pretty fast, so that should be no significant increase.
To put things into perspective, an article from 2009 estimates that Spotify used 84TB of traffic per day[1], back when they had 5 million users. Soundcloud claims to have had 40 million registered users in 2013 and 175 million unique monthly listeners in 2014. These numbers aren't easy to compare, but either way I think it's safe to say that 900TB should be a small fraction of their monthly traffic, and shouldn't make a huge dent in their financials.
[1]: https://www.theguardian.com/technology/blog/2009/oct/08/spot...
Why stay with a host that bends you over a barrel for transit? Even Netflix doesn't use Amazon to deliver their data heavy content, as their pricing is pie in the sky high.
The only problem is that I don't know whether IPFS has any way to gauge availability, so I'm not sure if the team could tell which files were only hosted by a few people.
It would absolutely be the best move for IPFS to be used in this case - maybe something like the AkashaApp guys, albeit for audio-media.
Edit: The Akasha App for those who aren't yet familiar with it - https://akasha.world - brings together IPFS and Ethereum to make a truly distributed peer network for persistent content.
I imagine that's one of the reasons why it's not ideal for archival content.
You can pay pinning services, but what's the point if you're just going to pay someone hosting it?
I'd rather have this archived on Siacoin, Storj, Swarm or any other distributed network with actual incentives to keep things around
Which is why I find Swarm a better solution. It is literally IPFS+Ethereum with additional support for ENS lookups, deniable storage, redundant storage, etc. This allows for far better privacy and being able to compensate the loss of parts of the file, both features lacking in IPFS itself.
The current swarm testnet performs, as per my experience, better than IPFS in terms of bandwidth and latency.
The time horizon of "archival storage" probably starts at a hundred years - that will need some structure to have a likelihood of persisting for so long.
This reminds me of one of the principles of camlistore's data schema, which explicitly says that they made their schemas overly-explicit so that future digital archeologists can re-create the schema purely from examples[1].
It's a shame that camlistore feels more like a very long experiment over a polished and usable backup/archive system.
[1]: https://camlistore.org/doc/principles ; There used to be a more explicit explanation, but I couldn't find it.
You can participate in the effort of course. Have a few hundreds of GB and a good connection ? Head over to http://archiveteam.org/index.php?title=INTERNETARCHIVE.BAK/g... and follow the steps !
I don't believe this is true. IPFS doesn't have any built in way of easily distributing parts of an archive, doesn't support (as far as I know) any form of erasure coding, making overhead quite high and requires that you use its own weird block + hash scheme for integrity.
It's also very immature, we don't know if IPFS will be around in 10 years and we don't know what kinds of bugs it will have.
IPFS is a great tool and it has its uses but I don't think archiving is one of them yet.
https://github.com/cetra3/rustcloud
It would be an absolute shame if Soundcloud disappears. There has been so much music I have discovered on this service.
There are ads. Audio ads between songs.
It's a node tool built a few years ago to download the playlists of users through your command line. Might be helpful for a situation where you'd like to back up your own playlists.
You'll need to get an API key - no sure how feasible that is at this moment.
[not all, because archive.ort needs to be a bit more careful, but they have a decent symbiotic relationship]
> The website is temporarily unable to service your request as it exceeded resource limit. Please try again later.
I suppose I prefer an archive over the blog being unavailable
http://web.archive.org/web/20170717083540/http://archiveteam...
it is interesting that their blog is not good enough for archive. link above and below are different (first one is updated but second one is not, because link contains some extra info). they will be able to save a lot of space by finding duplicate.
http://web.archive.org/web/20170606104512/http://archiveteam...
Don't get me wrong, I appreciate the work they do, and without them lots of content would simply disappear. But solving this problem should be at the core of the protocol itself (Xanadu, anyone?[1]), not depending on the resources and goodwill of a single team.
Just like IPv6, I don't think the problem will be solved as long as there's a patch that somehow works.
So as long as there's a patch that somehow works we have at least a solution. If that patch wouldn't be there we'd have nothing.
There is no perfect solution in sight and this a good one until a better one comes along.
I disagree. Expecting networks and software to act as an immutable medium is a fool's errand. The internet was never meant to provide a permanent cultural archive, and it's not actually a "problem" that it doesn't, because that's not what it's for.
Backups should be a service, not a feature of the network or the protocol itself. I think that what Archive Team does represents the correct way to approach the issue.
I've written my own Soundcloud offline audio player, but didn't distribute it because it was against their TOS.
There's no such thing, copyright is automatic, it applies as soon as the author puts the work in some tangible medium (like an hard drive).
If you then gave a performance and someone recorded it, the copyright in the recording would lie with them - but they would not be able to distribute it without also having your permission.
(People forget that these are different and the songwriter royalties are quite lucrative - famously e.g. the Beatles almost all the songs are joint copyright Lennon/McCartney, not the band as a whole)
So if I record and master a random riff, I technically have a copyright on it?
To be fair, I was mostly talking about ones with a record label, where you might be hounded by a label with a lot of legal representation.
Would putting it on a private Google Drive work? What about a Drive that's searchable and anyone can have a listen, if they so desired?
Shitty drawings made by five year olds get the same legal treatment as Picasso, etc, etc.
To me that means "stuff that's most likely fine to preserve and most likely isn't found on other places".
Also according to Sound Cloud's ToS by using them you are granting all users rights to "to use, copy, listen to offline, repost, transmit or otherwise distribute" your content. So if Archive Team downloads everything they can (that does not in itself violate copyright (i.e. they are not Metallica songs)) there should be no copyright issues.
Which means you're only allowed to redistribute content through the facilities provided by Soundcloud. You're not allowed to simply download music and share it outside of the website.
That said, simply downloading music from Soundcloud isn't copyright infringement. You have to do that anyway to listen to the music. But redistributing it (outside of Soundcloud) is illegal unless the copyright holder has granted permission to do so, or the work is under Public Domain.
At least I read this only as "by default you are granting all users permission to do whatever you like", then you can restrict the access rights according to next part.
>You can limit and restrict the availability of certain of Your Content to other users of the Platform, and to users of Linked Services
I'd guess there are still plenty of stuff to backup even if some artists/performers/bands use more restrictive licensing.
I had the same issue with backing up Geocities when it went down. I figured better safe than sorry, established a very easy deletion procedure for the copyright holders and have received only a very small number of nastygrams compared to an absolutely enormous number of messages from people that were happy their content got saved.
So at a guess, yes it is copyright infringement, no, it will not lead to trouble because most people are able to recognize a good faith effort when they see it.
A takedown notice from a few large commercial soundcloud users would probably be enough, no?
Maybe because they are profit-oriented company that raised hundreds of million USD in venture capital. Asking for donations would be kinda unethical, and the founders would know that.
However, if SoundCloud somehow would be transformed into a non-profit organization...
I expect they, like any business, would take a stance primarily based on liability, cost, benefit, etc... overriding how anyone working at the company actually feels about it.
It costs them nothing to ignore/outright reject without consideration a crazy proposal like "lets give our entire database of content to a third party". Assessing the technical and legal ramifications of that proposal costs time/effort/money. Why bother?
If SoundCloud gives their content to Archive Team, they're shouldering the possibility of some kind of liability, surely. If they say nothing, let Archive Team take it themselves, from their website, they let Archive Team (who understand and are willing to) take responsibility for that.
How much does a petabyte of bandwidth go for these days anyway? It might as well be cheaper than paying engineering, legal and management to arrange for some kind of off-line data transfer.
Or if you worked with them, and both of you happened to be in Amazon, that could be a no-cost transfer.
This way they can skate the rules and save the data.
Plus startups on the way out usually want to use their data as a bargaining chip to sell whatever remains of the service off, if they give it away they lose that chip.
Although, if anyone could tell me how to do that with AES-SAMPLE HLS video, I'd be very happy.
I'm always confused when people say they've backed up some website when there are these practical and legal barriers.
This forcibly imposes a fixed recording rate, limited to the length of the track itself.
The 10 minutes I spent poking around to see what "AES HLS" was made it appear that a MITM proxy would straighten that problem right out, unless you are in an iTunes FairPlay-esque encrypted to your iPhone type deal, in which case I believe the math is against you
https://au.tv.yahoo.com/plus7/screenplay/-/watch/36206022/sc...
youtube-dl seems to find the stream URL okay, but then prints ffmpeg errors:
https://github.com/rg3/youtube-dl/issues/11636
https://github.com/selsta/hlsdl should be able to do it in theory, but it looks like WideVine DRM is involved.
Someone has a Kodi plugin to make it work but Kodi doesn't support recording of any kind (?)
Just remembered cases like Deutsche Bank vs Leo Kirch which are legal nightmares.
I think you mean libel, and I think it's a reasonable conclusion that everyone is thinking given that SoundCloud just laid off so many people.
>508 Resource Limit Is Reached
I found a lot of the content to ephemeral, things like podcasts or DJ mixes. I dunno, it just seems a bit silly to put resources on it.
Both your example require transformation, while archiving unprotected digital media is as easy as making a copy.
Soundcloud is here to stay
Or did he call the CEO, the CEO told him 'Yes we will stay'? I mean what is he supposed to say?
'No, don't bother uploading new songs, we will run out of cash soon, thanks for the call'?