The surreal joy of having an overprovisioned homelab
xeiaso.net
xeiaso.net
edit: silly typo!
I, after all those years, can finally understand my dad coming into my room, yelling at me due to the electrical bill.
Couldn't understand, only was running Seti@Home 24/7 on my Pentium 4 :D
Proof of work is not a viable defense -- it's basically impossible to tune the parameters such that the cost is prohibitive or even meaningful to the scrapers but doesn't become an obstacle to users.
It's pretty much just a check for whether the client can run JavaScript. But that's table stakes to a scraper. Trying to discriminate between a real browser, a real browser running in headless mode, or something trying to fake being a real browser requires far more invasive probing of the browser properties (pretty much indistinguishable from browser fingerprinting) and obfuscating what properties are being collected and checked.
That's already what any commercial bot protection product would be doing. Replicating that kind of product as an on-prem open source project would be challenging.
First, this is an adversarial abuse problem. There is actual value in keeping things hidden, which an open source project can't do. Doing bot detection is already hard enough when you can keep your signals and heuristics secret, doing it in the open would be really hard mode. (And no, "security by obscurity is no security at all" doesn't apply here. If you think it does, it just means you haven't actually worked on adversarial engineering problems.)
Second, it's an endless cat and mouse game. There's no product that's done. There's only a product that's good enough right now, but as the attackers adapt it'll very quickly become worthless. For a product like this to be useful it needs constant work. It's one thing to do that work when you're being paid for it, it's totally another for it to be uncompensated open source work. It'd just chew through volunteers like nobody's business.
Third, you'll very quickly find yourself working only in the gray area of bots that are almost but not quite indistinguishable from humans. When working in that gray area, you need access to fresh data about both bot and real user activities, and you need the ability to run and evaluate a lot of experiments. Not a good fit for on-prem open source.
Like they said in the presentation, git(lab/tea) instances have insane amounts of links on every page and the AI crawlers just blindly click everything in nanoseconds, causing massive loads for servers where normally there might be a maybe a few thousand git pulls/pushes a day and a few hundred people clicking on the links at a human pace.
Plus the bots are made to be cheap, fast and uncaring. They'll happily re-fetch 10 year old repositories with zero changes multiple times a week, just to see if they might've changed.
Even a if the bad proof of work requires the bots to slow down their click rate, it's enough. If they somehow manage to bypass it completely, then that's a problem.
So if you don't understand these things best not to copy them because you could unintentionally be sending some pretty strong signals.
It's also somewhat in the style of the Manga Guides, which are fantastic and an entirely different way to make some of these things accessible.
Which doesn't strictly have anything to do with developers either other than the author being part of all three.
I learned a ton so far, it's been really rewarding having to manage the entire stack yourself.
Also Talos is golden!
But wait, you also run your plex instance on there? And a git server? And wait, it stores your tax documents too?
Is it a playground, or a functional home server?
I run plex and a NAS and a VPN server on a machine in my home, too. But it's a home server. It's definitely not a "homelab" or a "playground" of any kind, any more than my freezer is a playground for storing food.
The surest way to take all the joy out of a hobby is to do it as a job. You sound thoroughly employed.
Install it, configure it, make sure to give it enough storage so you have the desired retention duration - here about 2 days - and restart it regularly to keep memory leaks in check. Once you've got it set up don't mess with it anymore. I've run it this way for close to 5 years now without too many problems. I did a thorough survey of the available free and cloud-free software before I settled on ZM and regularly check whether something 'better' has shown up but have not seen anything which matches the functionality yet. There's lots of other surveillance camera software but most developers seem to consider a Javascript-heavy glitzy UI the prime directive while the backend is more of an afterthought. I hardly ever use the UI, most interaction goes through apps (ZMNinja etc.) or other systems (OpenHAB etc.). ZM, once configured, does very well this way as long as you keep the memory leaks in the zmc processes in check.
In short ZM is a flawed but functional piece of software with the best balance between functionality and flaws for our application - stable and farm monitoring services. It may not be what your want if you want to have a web interface to that camera on the front porch.
Good way to explore ideas without having to do concensus building with those around you. Its your own little pet project!
I certainly don't consider my home server that packs a 30 W TDP CPU from 10 years ago to be a homelab. Someone with half a rack that peaks at 2 kW? That's clearly a homelab.
I think of it as the computational equivalent to a home machine shop. If you don't have some heavy duty equipment can you really claim you have your own shop setup?
It’s both for me! I’ve had a homelab in some form since high school (~12 yrs ago).
Having a home lab has gradually taught me Linux, Shells/Terminals, Ansible, Docker, and most recently Kubernetes.
I have to be careful with backing up data when I do something new. I use this setup for Home Assistant, Plex, small self-hosted apps, a personal Jenkins instance for my open source stuff, and game servers for friends.
It has taught me a lot and has had significant (unintentional) benefit to my career by exposing myself to new concepts/tech and thinking about how operations works solo vs in a team setting.
If you wanna see it: https://github.com/shepherdjerred/homelab