PhotoPrism: AI-powered photos app for the decentralized web
github.com
github.com
> User Management
> Account Roles: Super Admin, Admin
So unless I switch to the >=EUR6.00/month Plus plan I cannot add a non admin user? So my grandma gets admin privileges? Is that not a blocker for a lot of families?
Edit: From https://www.photoprism.app/plus/kb/multi-user
> PhotoPrism® Plus includes advanced multi-user functionality and additional account roles. These roles are intended for situations where you want other people to have access to your library, such as giving family members access to your pictures without granting write permissions or exposing private content.
> It is recommended to set up additional instances if you have multiple users in a family, so that everyone can manage their own files independently. This way you can avoid problems with conflicting library settings, file permissions, and dealing with duplicates.
So you're actually supposed to run an instance per person it seems. But still, then my grandma would still be an admin. I think I'd like to about that.
Maybe, but at least it's not per-user, per-month. I'm much more likely to set up something like this if I can add / remove family members as needed without thinking about the cost beyond the annual fee I pay for to run my own instance.
The blocker for me is typically ease of use. They recommend PhotoSync for transferring files to the instance. I currently use that for my family, but sync to a Nextcloud install (which I desperately want to replace). The part where it fails for me is that my family members don't understand how to verify everything is running as intended. For example, I checked one of their phones earlier this year and they had 5k un-synced photos.
What this really needs to be useful is a native app that does sync from the phone in a way that makes it relatively foolproof.
Beyond the ease of use, every link in the chain of the workflow is a point of risk IMO. For example, what if PhotoSync rug pulls everyone (I have no reason to believe they would) and starts charging a subscription? The app that runs on your phone and does the syncing is the more valuable half of the workflow IMO.
And I'll pay close attention to the reliability of the sync. I've been using Syncthing for a while now and am still not trusting it completely. Quite regularly I have to restart a mobile app for the files to get synched and I just can't figure out what's wrong. I don't want that to happen for photos. Maybe I'll monitor the date of the latest photo per user and alert if it's too old.
Unfortunately, if you've used Google Takeout or other systems that can both downsample your photos and videos, as well as actually deleting _or changing_ metadata, deduplicating becomes a big wad of heuristics.
My first approach was to build a UID based on a series of metadata fields (like captured-at time, the shutter count of the camera body, and the geohash), but that breaks when metadata is missing or edited from their sibling variants.
Just finding what assets to compare against by time proved to be problematic in some edge cases. By storing in UTC (which is how I've built systems for the last thirty years!) made any file that didn't encode a time zones be wrong by whatever the correct offset was--almost all videos, and many RAW file formats don't have timezones. The solution I came up with was to encoding captured-at time _in the local time_, rather than UTC. I also infer or extract a milliseconds of precision for the date stamp, so files with more or less precision will still overlap during the initial captured-at query.
If you're interested, more details are here: https://photostructure.com/faq/what-do-you-mean-by-deduplica...
My main point here is to expose how deep this rabbit hole goes—I'd suggest you think about these different scenarios, and how aggressive or conservative you want to be in duplicate aggregation, and what you want to do with inexact duplicates.
Most importantly: have full offline backups before you use any deduping or asset management system. When (not if) the tool you're using behaves differently from how you expected, you won't have lost any data.
I also found that there really isn't one "best" image hash algorithm. Using _several different_ image hash algos turns out to be only fractionally more expensive during both compute and query times, and substantially improves both precision and recall. I'm using a mean hash, gradient diff, and a DCT, all rendered from all three CIELAB-based layers, so they're sensitive to both brightness and color differences.
So much time later, I would hope that we expect more sophisticated and flexible tools, not less, for that task.
I used libraw to read the actual raw data from my images, ignoring possible metadata that can get changed/updated by Capture One for example. The raw data is just fed into a hashing function, to get an exact content hash. Does not work if your image is down-sampled of course, but that was actually my goal - I want to know if the raw content has changed/bit flipped, but don't care if the metadata is altered or missing.
So if you, say, change the rotation of your RAW image file (which normally will just add or change the EXIF `Orientation` tag), the image content hash will stay the same.
See https://exiftool.org/ExifTool.html#ImageHashType
(of course, your current approach should prove substantively faster, but will only work with the formats that libraw supports)
Heck, we could even brute force this by resizing images via different methods and comparing checksums. But _surely_ AI can help us here.
Simple image hashes based on small thumbnails can do similar image detection quite well (see sibling comments for links to several different approaches).
Do this anyway. Always.
If you want to go deeper, I suggest grouping everything by image content - that means, at the very least, comparing images via a resolution independent hash. (E.g. average down to e.g. an 8x8 picture in greyscale for the simplest approach). (undouble does that nicely, and has a number of different approaches: https://erdogant.github.io/undouble/pages/html/index.html)
Why not calculate hashes just for the image data and then do metadata conflict resolution separately?
I thought takeout exported the original files? (alongside metadata entered into the application)
But I switched over to czkawka, which has a better interface for comparing files, and seems to be a bit faster: https://github.com/qarmin/czkawka
Unfortunately, neither of these are integrated into Photoprism, so you still have to do some file management outside the database before importing.
I also haven't used Photoprism extensively yet (I think it's running on one of my boxes, but I haven't gotten around to setting it up), but I did find that it wasn't really built for file-based libraries. It's a little more heavyweight, but my research shows that Nextcloud Memories might be a better choice for me (it's not the first-party Nextcloud photos app, but another one put together by the community): https://apps.nextcloud.com/apps/memories
General question: does this have import from google photos? Please don't make me go through the google takeout pain.
Docker is much less manageable locally than, say, systemd or supervisor.
The only thing (which is something) that docker has going for it is that it's cross platform. But the same argument can be made for writing apps using electron. If you're going to do that, fine, but acknowledge the compromise. It's not better.
Now I just stick to debian packages.
[0] https://github.com/photoprism/photoprism/blob/develop/docker...
I'm not sure if you perused the docker files for this repo, but imho there is nothing simple about them
Edit:
I was curious so I dug into them a bit and found that the dockerfile references a develop docker image. There is a second docker file to see how that base image is built. In the steps to build that base image, we grab some .zip files from the web for the AI models. We install go, nodejs, mariadb, etc for quite a few deps, and then there is also a substantial list of packages installed. One step also does:
apt-get update && apt-get -qq dist-upgrade
Which seems a bit iffy to me. Each step calls a script which in turn has its own steps. Overall, I'd say the unfurled install script would be quite long and difficult to grok. Also, I'm not saying any of this is "bad," but it is complex.The nonsensical one would be do-release-upgrade.
The only correct way is to either rebuild the base image from scratch or just fetch a new base image. My suggestion would be the latter, just run docker pull again for the baseimage and use that, without running update.
I dunno, I enjoy not having to install 500 libraries on my system just to test an app. Also upgrading those libraries without butchering my system is also nice. Not to mention rebuilding a system is really fast. Too many pros outweigh the cons
A developer must have had a working config at some point to create the Dockerfile. Providing literally only the Dockerfile is usually just a sign of throwing hands up and saying "It's too hard!".. you should be able to package for at least one platform that isn't Docker. That's just app development, or at least it was until recently.
But yeah, it's even fully integrated into NixOS options now. You can set up a default install with one line: https://search.nixos.org/options?query=photoprism
Docker is the one thing that works on MANY flavors of Linux.
If I want to provide a tool I want to spend my time on building the tool, not building an rpm, a snap, a deb, ...
The Docker build process is significantly easier. For example, I can just pull in NodeJS 20. I can't do that on Ubuntu. It's not available on packages.ubuntu.com.
Building a deb/snap/rpm is a whole other language to understand how dependencies are set up.
And then I need to test those. I've never even ran CentOS.
Requires: openjdk=20, opencv=4.5
I’m assuming openjdk is versioned correctly and the name of the opencv package is opencv. I’m also assuming openjdk, replace openjdk with your jdk package name.You would put this line in your rpm .spec file. Do you really think the above line is hard? Maybe the difficulty you have is in never have touched rpm. Start here: http://ftp.rpm.org/max-rpm/index.html
None of that has to do with the difficulty of rpm spec, but entirely with organizational planning. You do plan while building software.. right?
Also github does the exact same thing but with deb. It works in the real world. Quite well, too!
Yes but I am trying to see if there is a better solution to the red hat freeloader problem. To spell it out, What is to stop Oracle (doesn't make sense but couldn't come up with a more evil name) from purchasing a license and releasing IronPrism or whatever by doing s/PhotoPrism/IronPrism?
That doesn't mean the code can't get distributed more widely. I guess it's a bet based on the honesty of users.
Overall, I was underwhelmed
1. Select a photo with a recognizable face
2. Click the Pencil / Edit icon in the upper right menu
3. Select the People tab
4. Enter the name/names
Well, there is a good reason. To make the install more simple. IMHO containerization makes selfhosting way easier.
When I see a software application that recommends Docker for deployment, I always assume that I'm going to need to do extra work to make it behave correctly. Whether that is figuring out how to make it auto-start, or how to forward network ports, or how to share a data partition with the host OS.
Non-docker installs are simpler. At least, for my skill set.
I see one dockerfile and 7 docker compose files (.yml)
The dockerfile does not apparently do anything useful. I'd be amazed if running that dockerfile by itself produced anything useful
Now, I don't know very much about docker compose, but I learned a bit of it in order to get this software running on my server. If I worked at it, I could almost certainly get a working install of PhotoPrism without using Docker, but it would be annoying work, and I wouldn't have any certainty. I wouldn't know that it was correct, and any time something didn't work the way I expect, I would worry that I screwed something up during the installation
Not to mention the added operational complexity involved in managing a dockerized application compared to managing e.g. an equivalent webapp deployed without containerization (systemd service file, configuration file, etc)
In any case, I found this in the docker/ subdirectory (if it was me I would have put it in the main README):
https://github.com/photoprism/photoprism/tree/develop/docker...
Building From Source
You can build and install PhotoPrism from the publicly available source code:*
git clone https://github.com/photoprism/photoprism.git
cd photoprism
make all install DESTDIR=/opt/photoprism Missing build dependencies must be installed manually as shown in our human-readable and versioned Dockerfile.
I haven't been able to find any dockerfile that lists dependencies. My guess is that the documentation is referring to some prior architecture or something similarNoticing the broken link points to photoprism/docker/develop/Dockerfile, I supposed they had moved it and indeed, just by going one parent up, the directoy photoprism/docker/develop/ contains subdirectories for lots of base systems, each one containing a Dockerfile that lists dependencies needed for each of them.
For example, for Ubuntu 20.04:
https://github.com/photoprism/photoprism/blob/develop/docker...
All the portability of Docker, plus (something close to) the ease of use of installing a distro package like an .rpm or .deb.
https://github.com/photoprism/photoprism/blob/develop/LICENS...
Selling open source or free software is fine.
As such you could fork it, tweak the code around the license verification to always return TRUE or whatever, and you would be running a "pro version" of the software.
The developer simply trusts that people who like the software will purchase a license, or the fact that the majority of people out there are not programmers and would not be able to rebuild the software for themselves. Besides, the product is also offered as a SaaS.
would be cool if also had dynamic shared albums as google photos has with facial recognition
Did I conceivably mess up the setup or is this other's experience as well?
Google photos AI is so far ahead of Photoprism it isn't even a contest unfortunately. Which sucks because I am more than willing to pay for a viable alternative (still looking...).
But after ten minutes of browsing around, I still have some open questions regarding privacy. What would I have to do to have an absolutely, unequivocally local-only install? Can I even do that? Which features would I need to disable? Especially anything related to AI and classification raises red flags in that regard.
While it is apparent that they've given a lot of thought to privacy and while their privacy policy is definitely one of the better ones I have seen, it still conflates things such as website access with usage of the tool itself.
It would be nice to have one clear, guiding document that outlines how private a private install really is.
https://stackoverflow.com/questions/39913757/restrict-intern...
Their recognition model is loose. It works and it works well enough for what this is to be useful but it's just a quick classifier. It will absolutely have misses.
You can optionally aadd in a coral dev board to save yourself some cpu load but I've never found it necessary.
I'm absolutely fine with a good enough classifier, as long as it is local. I would even be fine without one.
I use FolderSync on my android phone to copy any pictures I take to my photoprism WebDAV
Genuinely curious, if you use Android can you truly de google your life?
But, if you use grapheneos, they replace the google services with their own alternatives and sometimes let you disable the feature entirely. Also, on any rooted android. You can use AFwall+ to firewall any google connection or app. This way, even system services can't make network requests.
As far as I can tell, the only place you really need to be logged in to a Google account to get a "normal" Android experience is the Play Store, which makes sense.
I use Google Maps for example (Osmand when I can, but Google Maps is the only free app that has good traffic information) while not logged in, but my photos are hosted by Photoprism, my emails are not at Gmail, my backups are on my home server, etc.
I find myself largely independent from Google and if that's your goal, as opposed to trying to hide yourself from Google (which would be better, but I find it unrealistic for my needs), it's easy to achieve on an Android phone and certainly easier than independence from Apple on an iPhone.
I've got an older Synology NAS that currently stores my 250GB photo library with a lot of RAWs. While it can run docker I'm not going to attempt to run it there. I'm wondering if a VM on my 10700k is going to be sufficient for great performance, or if handing over the complete machine is going to be necessary. I'm happy adding more NVMe 3.0 drives for storage.
I'm also curious how version upgrades have been for you. I want to get out of the ops business but I feel like I'm going to be dragged in.
Biggest thing was to put all the photos on SSD, at least the thumbnail cache if not all.
I also run it, on my backup server, pi 4, with all ai disabled, it's slower, but probably because it's all on HDD still usable.
I would say my requirements are nowhere near yours and that you would really need to test it yourself for your usecase. Some people on the photoprism forums claim it can handle large libraries like yours but I would always recommend testing yourself.
I host it via docker on a pretty tiny VPS with 3 vCPU and 4 GB RAM and it's fine.
The deal breaker for me was that there was no way to share the entire inventory of my photos without providing admin access. Does anyone know if that feature has been provided?
(*) I used to be able to find some pictures using simple keywords like "dog" but Google finds far fewer pictures with those keywords any more. OTOH, it is amusing to see the interpretation of some of my photos in Photoprism, but I suspect that will get better with time.
> Please don't upload photos containing offensive content. Uploads that may contain such images will be rejected automatically.
> Non-photographic and low-quality images require a review before they appear in search results.
How does it know what is "offensive"? Is this configurable? I don't want to upload photos of people (for obvious reasons) but other than that, does it even know the contents of a photo to tell if a photo is offensive?
Is a photo of a shelf at a grocery story selling beer offensive?
Is a photo of ice on the sidewalk offensive?
Or is it just about human nudity and porn?
In cases like this the "burden" should be on the party choosing the content they want to see.
Everyone should be able to upload a photo of ice on the sidewalk, but the other users should define that ice or sidewalks are offensive to them.
Why do private photos need to be reviewed?
What am I missing? Is everyone just okay with giving up their privacy?
I get a bit bored of the knee-jerk "Free Speech" stuff.
2. Is is configurable and I believe mostly off by default, depending on install method
3. It is lit up in their demo instance
4. It is not a bespoke content filter, this is an `AI POWERED` app that classifies and labels photos so that they can be indexed for text search. Everything is processed through a NASNet and labeled. The upload filter just does some pretty rudimentary heuristics on the labels and decides if it will allow or quarantine the upload. To be clear, It is all server side-- the image is uploaded, its just quarantined.
The review section is not for NSFW things. It is for low resolution or low information images. Again, it runs locally. As far as I know, it is to keep your library from being clogged up with crappy thumbnails or screenshots or other junk. If you don’t care, you can just bulk approve them.
It also does facial recognition locally. You aren’t giving up any privacy, which is why I use this after resisting Google photos and just doing the folders method for years.
Edit: Okay, found it. I think it's likely that those warnings are only for uploading on their demo server linked on github and not on the locally hosted product.
https://github.com/photoprism/photoprism/blob/ff5f3ddeb1a198...
Conversely, Synology Photos (which I use instead now) has a fantastic mobile app. However, if you really want reliable and granular object/face recognition, the Syno app is a little bare bones. It does some face recog but that's really it. However, its backup feature is reliable and I don't have to worry about if it's working in the background or not. I tried Nextcloud's photo plugin for a while too, but its mobile app had issues and I just couldn't rely on it.
Another project I've been keeping my eye on is immich: https://github.com/immich-app/immich
I'm happy for now with Syno Photos but it would be nice to have my photo app in my container environment with everything else I run and just use the NAS for media storage like I intended to.
The downside is that Synology Photo's database schema and API are not officially documented, but you can find people who have documented them and since it's on a machine you own you have unlimited access to them, so it's still a big step up from Google Photos and Apple Photos which are chock full of restrictions - e.g. Apple Photos doesn't even have a web API.
I have a nas to which I automatically upload all my phone and DSLR photos/videos. I'd love to see automatic montages every not and then with some animations/music etc. without having to do them myself...
The 'memories' feature of of Apple Photos is frankly amazing.
Nextcloud was basic, but serviceable. I liked that it handled docs, contact sync, etc. The sync was clumsy and I found updates frequently broke the system.
I just tried PhotoPrism with my ~75GB photo directory. The classification is.. decent and all run locally. It took around a full 24 hours to index and classify my photos, which it did reasonably well. Google Photos' classifications is miles better to be sure, but PhotoPrism was easy to set up and works pretty well.
I no longer expose any of my services outside my local network (besides Wireguard) so I'm hoping to find a Photosync system that will only upload when on my home wifi and charging. Any suggestions there?
https://help.resilio.com/hc/en-us/articles/204754349-Can-I-f...
I don't use docs/contacts or even the official photo sync though, I just sync files through Foldersync on Android. That can easily be set to only upload on certain wifi.
Syncthing
Edit: for example, just yesterday I uploaded ~30 photos of my girlfriend. It recognized her in one photo and didn’t even identify her face as a face in the rest.
I also don’t understand what’s doing the processing for the facial recognition. Is my cpu doing it or is an external service doing it?
I have a large collection of google photos "takeout" zip files. I dont need to unzip them but I would like a photo app to view them.
They are stored on a server where i can run docker not "locally".
Does anyone have a recommendation?
"Thanks"
Anyway, you could try using something like https://bitbucket.org/agalanin/fuse-zip/src/master/
To mount the zip file and then use any tool you want to operate on the files?
I will look at the mount option. Thank you for you suggestion!
It's a... passable alternative to Google Photos. I downloaded it beforee they started advertising AI and decentralized features - I'm kind of surprised they took this marketing direction.
Plus: all thumbs of non-square images are cropped, and it makes all this per source image:
-rwxr-xr-x 1 photoprism photoprism 1689 Sep 21 2022 json/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_exiftool.json -rw-r--r-- 1 photoprism photoprism 3944 Sep 21 2022 thumbnails/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_100x100_center.jpg -rw-r--r-- 1 photoprism photoprism 7856 Sep 21 2022 thumbnails/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_160x160_crop_2a51210960c8.jpg -rw-r--r-- 1 photoprism photoprism 12881 Sep 21 2022 thumbnails/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_224x224_center.jpg -rw-r--r-- 1 photoprism photoprism 12151 Sep 21 2022 thumbnails/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_224x224_left.jpg -rw-r--r-- 1 photoprism photoprism 13601 Sep 21 2022 thumbnails/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_224x224_right.jpg -rw-r--r-- 1 photoprism photoprism 100 Sep 21 2022 thumbnails/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_3x3_resize.png -rw-r--r-- 1 photoprism photoprism 49090 Sep 21 2022 thumbnails/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_500x500_center.jpg -rw-r--r-- 1 photoprism photoprism 1824 Sep 21 2022 thumbnails/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_50x50_center.jpg -rw-r--r-- 1 photoprism photoprism 74535 Sep 21 2022 thumbnails/0/0/0/00000f1790c93fdaaca662b56326fb5d130ab45d_720x720_fit.jpg
The UI doesn't change the url to the photo you're looking at, so you can't share urls.
Images aren't links, so you can't do browser stuff like open one in a new tab. The back button doesn't work right at all.
The choice of metadata to display under each image in the listing is downright silly ('place', city, year, DoW, month, year again, time, zome, camera model, resolution, file size, city again, state, country).
The information scheme features 'albums', 'favorites', 'moments', 'labels', and 'folders', as if they were being forced to match a bunch of legacy systems. I want a powerful tag/group mechanism that I can use for many possible workflows.
The list under 'search' says 'Review 40193', clearly telling me I have that many pics to review. Clicking on it says:
No pictures found Try again using other filters or keywords. In case pictures you expect are missing, please rescan your library and wait until indexing has been completed. Non-photographic and low-quality images require a review before they appear in search results
(And I had to disable some user-select:none styles to copy paste that)
But yes, never set media hosting to open upload unless you want to become an image host for content people can't host elsewhere (or just for images used in spam campaigns).
Providing face recognition sans a cloud seems like a big selling point to me.