Edit to add: oh yeah, I just read another comment that reminded me that it also ate up lots of disk space as well.
Edit: I should expound. Jenkins seems like it has a lot of clunky moving parts. It all works, and I’d rather use it than anything else, but it’s kind of like IKEA furniture: you use it because you have to, not necessarily because you want to.
It’s also incredibly difficult to automate. I can configure Postgres with a config file or two and easily use Ansible to get the exact same instance every time. Jenkins has to be dragged into Automation Alley kicking and screaming. I partially blame this on the fact that Jenkins has nontrivial amounts of configuration that’s done via GUI. I approach a long-running Jenkins instance with the same fear and dread I approach a Windows box that hasn’t been restarted in six months. I.e. the box is now a snowflake and trying to make it reproducible and automated is going to be a bad time.
I could go on, but as a devops critter, Postgres wins every time.
It's like comparing a missile with an airplane. One gets where it needs to go faster and more efficiently, and the other one transports people.
Sure it works but from an admin perspective it’s horrid.
Only you don't. I'm not sure why you seem afraid of Postgres, but millions of people use it, across tons of companies, and even for personal projects, and it's trivial to setup and run. Oh, and those lots of plugins and stuff you mention? You don't have to use them if you don't need them. They don't even enter the picture at all.
Because I've managed database applications before?
> and it's trivial to setup and run
Jenkins is more trivial. It's one process. It doesn't depend on an external high-availability networked data service. Backup is 'cp -r SRC DEST', or a plugin if you're fancy. And, again, Postgres does not replace Jenkins, it's just the storage and querying. It adds a lot of complexity and service availability points of failure that Jenkins does not have.
Copying the jenkins directory can be done but restoring it on another fresh install will not work. There are many files that needs to be deleted manually until the instance can start without error.
pg_dump on the same data volume takes about 14 minutes to dump, compress and move to another node.
The filesystem is a shitty database. Thought we’d all learned that by now.
How dare you :) "The" filesystem is a great database.. for certain applications. Big, binary blobs of video, store, and even index, particularly well!
I think what the collective "we" haven't learned is to avoid trying to think about scale intuitively (rather than "doing the math") and to avoid extrapolating from the trivial scenario.
It's why "latency numbers every programmer should know" is still a thing.
Heavily locked stuff, lots of small things, huge number of locatable data entries, not so much. Which is Jenkins.
As for scale, Incidentally I worked on a very old filesystem based store back when we had spindles to contend with. We had four racks of Sun disk array trays. The only way it performed well was keeping it to 2Gb spindles. I’m well aware of scale issues on file systems, perhaps moreso than the Jenkins developers. We had to scale that to 4000 concurrent users.
My comment was partly in jest, and mostly hoping to spur conversation, not as a true disagreement or criticism with your comment.
Still, saying "It's OK" for those narrow use cases may be too dismissive, even if "great" is an exaggeration. There are plenty of examples where DBMSes (especially relational ones) have fared poorly in comparison.
> Outside of that it needs something that has some intelligence
I fear I'm missing your point here. Certainly "the filesystem" as in the Unix syscall interface to a hiearchical arrangement of files, lacks intelligence, but that doesn't mean the specific underlying implementation must.
We've even come a long way from every being the Berekeley Fast Filesystem. Besides the many choices of underlying filesystems (including CoW ones like ZFS and BtrFS)
> and ability to read optimise it.
I assume by read optimization you don't just mean something like the buffer cache, but a user-specified index?
> Heavily locked stuff, lots of small things, huge number of locatable data entries, not so much. Which is Jenkins.
Does Jenkins use external locking? It would be odd in light of some of the comments elsewhere in the thread touting its advantage of being a single process. Of course, even if it's using only locking internal to itself, there's a good argument that its authors needlessly re-invented a DBMS (which we've seen happen when other niche-use databases get used for broader purposes).
I'm not sure a large number of tiny files is inherently problematic for a filesystem, merely problematic for existing implementations, and some are better at it than others. What about something like libferris (assuming perfectly spherical cows and ignoring the performance implications of FUSE for a moment), which can back a filesystem with an arbitrary database?
IOW, is Jenkins-using-the-filesystem an Ops optimization/tuning problem, or is it a more fundamental problem that can only be addressed with modifying its code?
> The only way it performed well was keeping it to 2Gb spindles.
I worked with the aforemention hardware extensively, early in my career, but at a company that wrote a data warehousing (aka OLAP) RDBMS.
I'm reasonably confident in saying that your performance observations have almost nothing to do with the filesystem itself and everything to do with I/O performance in general. Large numbers of smaller spindles were absolutely required good database performance and scalability.
> I’m well aware of scale issues on file systems
I didn't mean to suggest you didn't, rather the opposite, as "the collective we" was a euphemism meant to imply "everyone but us".
Anyway, my overall point is that you and I may be acutely aware of the real, practical problems with scaling filesystems, but we're rare. Since there's nothing fundamental/theoretical that makes the filesystem an obviously poor choice at modest scale (the definition of which increases as computing power increases), the lesson does not get learned by everyone.
Instead, because truly large scale becomes rarer and rarer as computers become more powerful (CPU more than I/O, of course, but then.. SSDs), every time the lesson is re-learned by an individual/company, they think it's a new, or at least unique problem, and we end up with a re-invention of Portable Batch System (a fairer characterization than re-invention of cron, IMO).
https://github.com/geerlingguy/ansible-role-jenkins
https://github.com/geerlingguy/ansible-role-postgresql
Two ansible roles by the same author, supporting both Centos and Ubuntu. Not hugely different in complexity IMO. Installing Postgres on FreeBSD, though, is little more than
pkg install postgresql10-server
sysrc postgresql_enable=“YES”
service postgresql initdb
service postgresql startHow does Jenkins HA compare?
I've only ever used it in the internal/build scenario, never in production.