Yark: Advanced and easy YouTube archiver now stable
github.com
github.com
I made a docker container to run it (https://github.com/na4ma4/docker-yark), when I get time I'll do a PR if you're interested so it isn't a separate project.
(I'll also fix it so the host is a command line argument not just changing the binding from 127.0.0.1 to 0.0.0.0)
Very nice, thank you!
Will this archive subtitles?
Will it archive comments? If is can, can comments be updated?
Also, from the readme it looks like all metadata is kept centralized instead of each video having its own metadata file.
Are containers other than mp4 supported?
Wayback machine solves for webpages, but nothing I'm aware of (short of youtube-dl-ing the video yourself and storing it somewhere openly, probably at risk of various infringements) solves this. Quite a lot of hassle for something rather simple.
It would be great to be able to immortalise them on a per-video basis, so if it's important enough, we can be sure that references made to the content will still be there in the future when needed.
Arrimus 3D recently replacing a large chunk of their 3D modeling tutorials with religious content was a pretty big lightbulb moment for me that so much of the content I rely on - not just for the initial learning of a new skill, but as a continual reference when I forget something - is so fragile.
I immediately bought a NAS and began backing up everything that I gleam even the tiniest bit of learning from using a similar project, TubeArchivist[0]. Projects like this are really important for maintaining all of the great knowledge on the web.
There are plenty of causes of delayed audio, however. Bluetooth is a big one, if your device and software aren't properly compensating for the Bluetooth transmission delay.
If you need to troubleshoot sync issues, this is a helpful tool: https://www.youtube.com/watch?v=HD4emXqHCsE
same as just "yt-dlp https://url-of-video" from a CLI
https://github.com/Owez/yark/blob/676074ee3d9e379d15e52ffe2e...
It would be nice to expose this setting via config file and/or the cli.
I'm going to cap videos to 1080p by default and have a config setting to customize this.
Importantly, pay close attention to what artifacts are uploaded to an item created, and what metadata is set as part of the upload process.
Google is massive.
We got too great a deal for so long that we many people can't see things any other way.
Youtube sells ads and is subsidized by one of the biggest ad companies in the world that happens to have a lot of cheap cloud storage available.
The scale of storage and serving is not of any real comparable scale.
The donations go to the Wikimedia Foundation which spends the money on a bunch of other social stuff that's not related to running Wikipedia at all. So unless you want to support all those other causes, you're effectively wasting your money by donating to Wikipedia's frequent donation requests. Wikipedia isn't the storage hog that YouTube is; it doesn't cost that much to keep it afloat, and the people moderating the content (the editors) are all volunteers anyway.
However, at this point, IA needs your donations much much much more than Wikipedia does.
Wikipedia is mostly text. You can download the whole thing and fit it on a SD card.
My workplace has a private cloud with some 70PB of storage, plus tape, and tons of desktops and laptops.
b) just because i HAVE 10Gb of free disk space doesn't mean i'm going to offer it up for archiving of random internet crap
c) if it was on my phone it'd cost money and now we're very actively ignoring just how expensive it is to be online in 3rd world countries if you're not a rich expat.
> My workplace has a private cloud with some 70PB of storage, plus tape, and tons of desktops and laptops.
sure. HOw much of that do you think they'd be willing to contribute, for free to backing up random crap from people on the internet that may or may not be legal and could open them up to litigation because they're not an ISP / platform and thus not protected by the Shield laws?
Really? Other than low-end Chromebooks, I'd bet way over 75% of PCs under 8 years old have 10GB free.
Why bring race into it at all? Do privileged black/asian/etc people not exist? I'm sure you didn't mean it, but it's a very shortsighted thing to say.
Also, I think you're taking OPs comment too literally. Yes, not everyone has several devices and not everyone would want to store random videos on their devices, the point being though that 10gb is a veryy conservatively small number - and it being a fun thought experiment to imagine things like this.
(From a random source on the internet, [1])
[1] https://www.petercai.com/storage-prices-have-sort-of-stopped...
It's a nice thought experiment, but we really don't need to archive all of YT. That's why I appreciate projects like this and yt-dlp that allow me to not just archive what I'm interested in, but to watch it when and how I want, without Google tracking my every move, and interrupting every few minutes with ads. Paying for YT Premium only partially solves the second issue. I don't want to see sponsored content either.
Each person gets to decide what they find important, backup and commit resources to.
We already have the sharing technology (BitTorrent, IPFS), and the backuping tech (ArchiveBox, TubeArchivist). But they are not integrated and they are not easy to configure and use for a nontechnical person. And they are unlikely to become mainstream thanks to the copyright cartel.
Many years ago there was the dream of every home having a home computer giving people ownership over their digital life. A world where everyone has their own email server, their own diaspora pod for their family, their own blog, etc. etc. and all they had to do was buy or build a box and plug it in.
Projects like freedomplug and sandstorm.io and YunoHost and others. There are even newer projects like Umbrel. And of course NASs that now have apps on them.
Arguably, NASs and Umbrel are really easy to configure and use. OMV is somewhat harder but not unreasonable.
Alas, a self sovereign world is not what people desire. People are content with letting themselves be controlled by a few megacorps. Everything is in the cloud (someone else's computer).
Also, with the rise of botnets and DDOSs, it became unfeasible to self host something public without something like Cloudflare in front.
I miss the old internet.
So even if all these technical solutions to problems only technical users care about become as easy to use for laypeople as modern web browsers are, the general public just won't care about it.
I've long believed that the blame for this lies mostly on early WWW architects. If the focus from the very start had been on sharing content as much as it was on consuming it, and user-friendly tools analogous to the web browser had been built, then the general public would be educated that the web works by being in control of your data, and sharing it selectively with specific people, companies, or the world. ISPs would be forced to deliver symmetrical connections to enable this, centralized services would be much less influential, and the web landscape would look much different today.
This was actually planned as a second phase in the original HyperText proposal[1], but was never completed for some reason. I'd be very interested to know what happened to this effort. If someone has insider knowledge, or can contact TBL, I'd be very grateful.
Alas, it's too late for this now. The centralized web is how most people experience the "internet", and that train has no chance of stopping.
There are many ideas that fit under decentralization: torrents, the fediverse, crypto, even the old idea of the semantic web (because it was about standardized formats for metadata and carrying that metadata with the data instead of having it siloed in a central entity).
All of the hype around web3 really is only about crypto, because web3 is a marketable term for speculators.
Currently I am very cautiously hopeful about what the hype surrounding Mastodon (caused by Twitters self-immolation) will lead to.
But this is all hypothetical. We don't need a global YouTube archive. We need to stop using it altogether, and replace it with decentralized services. In the meantime, the existing personal archiving solutions work well.
I prefer the file prefixed with a number that indicates "air date". 01 being the first uploaded video. The default is by index and the top of the channel or playlist is number 01 which is the most recent.
Pulls from my watch later playlist which is quite handy
It's very cool to be able to search through and remind myself of something I heard once. Not exactly life changing, but still, nice to be able to quickly drill down and find audio for something when a curiosity strikes me.
Now it's down to 1 or 2 machines depending on what's going on, so it'll take much longer to finish up, but I'm in no rush.
I call BS.
This includes data from 1995 on. The early data is backfill of radio shows that transitioned to podcasts and dumped old episodes in their feed at some point. My reader itself started in 2012, I downloaded around 7000 hours of new podcasts, which works out to 1.7 hours per day. So, around 2 hours per day, since I don't listen every day, and to be fair, I haven't listened to every podcast I've downloaded, some don't interest me. But 1-2 hours of listening a day is the sweet spot for me.
My biggest annoyance at the time was importing my existing videos (and converting them to a streamable format, generating thumbnails and hover-previews etc). Do you have any plans of allowing to import existing yt-dlp folder (in the standard layout with a bunch of mkv files, the info.json, the subtitles etc). Because my current archive contains a lot of already-deleted videos :(
BitTorrent magnet uri's or hashes are suppose to be made from torrents but they really just point at a torrent.
One could make torrents for each video, take the youtube url v param and make a hash from that and point it ("erroneously") at the torrent.
That way, provided it was downloaded before, anyone who has the url can obtain the video.
The idea needs one more trick to validate the download. I suppose one could compare a chunk of yt to the same piece downloaded over BitTorrent but perhaps there are better ideas to be had.
Eventually, with tit for tad one could swap one chunk of one video with a chunk from a different video on the same channel.
- the Distributed YouTube Archive Discord: https://discord.com/invite/PQqks7eSKc
- ArchiveTeam also do a significant amount of YT archiving: https://wiki.archiveteam.org/index.php/YouTube
- a similar, but private effort: https://reddit.com/r/Archivists/comments/5uvfpw/youtube_arch...
https://github.com/chapmanjacobd/library/blob/f778e22bf80c58...
My focus is on error handling and trying to differentiate between unrecoverable errors and recoverable ones (try different proxy) but there's still a lot of work to be done.
Also look into https://github.com/swolegoal/squid-dl
"The use-case of tabs are websites that you know are going to change: subreddits, games, or tools that you want to use for a few minutes daily, weekly, monthly, quarterly, or yearly."
i mean, it's not really important to a user which ever library it uses to scrape youtube - i suppose it's important if you want to contribute/develop it.
Is it a library for things you've watched and want to store outside of youtube? Or is this for storing content you've created / managing your own portfolio of content?
> Yark lets you continuously archive all videos and metadata for YouTube channels. You can also view your archive as a seamless offline website
This applies to ripping, too. Funimation removed Drifters years ago, but I’ll always have a copy of it because I ripped it. Of course, I need to store it so it still costs money. But I can be content that I have the content.
https://github.com/VeemsHQ/yt-channels-archive
This just provides the latest HQ of the videos + thumbs & metadata, no historic information such as changes in video titles.
You clearly didn't even skim the readme to see what this does
> Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
It seems Apple is fine with adblocking by default for web content, but inconsistently as with youtube. Hard to predict what's risky to invest dev time into
I get confused about their stance on this because they allow products like AdGuard or browsers that block ads by default
I've been taking my first steps at having a home server, and one of the things I'd love to do with it is having an archive of the videos that I have saved in my private playlists on YouTube. In my mind, the service would periodically check all my playlists, compare with what exists locally, and download any missing video. Maybe even with a nice web UI so it's easier to visually configure and use.
Does such a service already exist so I can self-host it?
You can subscribe to playlists, as well as automatically update and download videos.
I mean, there's a time and a place for 4K, but watching the zits on the face of a guy who tells you how to play C Am F G on a guitar isn't it.
Not to mention, most of those cheap 4K cameras won't have optics to utilize those pixels; no quality will be lost in 720p.
If you want gui, check out TubeSync. Web UI for yt-dlp , ffmpeg and nicely packaged in docker
Issue for downloading playlists: https://github.com/Owez/yark/issues/49
Welcome to the nightmarish world of authentication to Google products, which all have 4 different versions of documentation and not a single one up to date.
For auth it seems the preferred way to login with Google is OAuth2, that's I believe what third-party apps use, e.g. Thunderbird uses it when setting up a new GMail account.
However, for apps that don't support OAuth2, there is also the possibility of using "App Passwords" [1], I've used one in the past and it worked well. (Update: I'm just reading it only works if 2FA is enabled, which I use)
[1]: https://support.google.com/accounts/answer/185833?hl=en
It requires a local file with a list of channels/playlists. I use Jellyfin (https://jellyfin.org/) as the frontend/video player.
gasp can i run it as a github action???