HNHacker News
TopNewBestAskShowJobs

ImGajeed76

110 karma · joined July 21, 2025

submissionscomments
ImGajeed76··on I imported the full Linux kernel git history into pgit
you should! "go install" it and you're up in a minute.
ImGajeed76··on I imported the full Linux kernel git history into pgit
i did look into this before writing the post. there's a fossil-users mailing list post by Isaac Jurado where he reported that importing Django took ~20 minutes and importing glibc on a 16GB machine had to be interrupted after a couple of hours. he explicitly warned against trying the linux kernel. the largest documented import on the fossil site itself was NetBSD pkgsrc (~550MB) which already showed scaling issues. so "never did" is fair - not because anyone tried and failed, but because it was known to be impractical and explicitly discouraged.
ImGajeed76··on I imported the full Linux kernel git history into pgit
Thanks! LWN's development cycle reports are incredible and were actually an inspiration. The goal here wasn't to replace that kind of expert analysis but to show what becomes possible when you can just write SQL against the raw history. Your reports add the context and understanding that no database query can provide.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
yeah i get that. sorry if it comes across as too salesy. but keep in mind that pgit was only meant to be a demo of pg-xpatch and wasn't built with beating git in mind. the fact that it's SQL queryable and comes close to git's compression was a nice side-effect. so the whole thing was really just built for showcasing xpatch's compression and evolved into what it is now. but yes, in theory you could also just store the git history uncompressed, which would actually solve quite a lot of issues i had :)
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
Sounds great! Yeah i have been working on a 3 layer cache in pg-xpatch so its not only in-memory cache but a little more sufisticated and hopefully uses less ram... haha. but its still not quite what i want.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
They are talking about this fossil: https://fossil-scm.org/home/doc/trunk/www/index.wiki
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
fire
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
haha yeah pretty much. but postgres already solves most of that complexity for you, so you get SQL queryability almost for free.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
yeah totally get that. the main blocker was delta compression. sqlite's extension api made it really slow for custom storage. i either had to do all the compression on the pgit side (and lose native SQL queryability) or just use postgres which handles it natively. but an sqlite version isn't off the table for smaller repos where that tradeoff makes more sense.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
haha, great that you tried! i also imported it multiple times now and it does work. but it's huge. the times actually match quite well, i also had around 3 hours, i'm surprised you managed to do it that fast actually. so yeah, i'm currently working on multiple things to improve the speed for importing and then also for analysing the kernel. but that will be something for the next post. stay tuned! as a quick teaser: it imported the 123GB uncompressed master branch into 2.98 GB pgit actual data while git aggressive puts it into 1.95 GB. but keep in mind, pgit was never meant to beat git in any terms. it really started as a demo XD
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
thanks! FUSE is actually a really cool idea, hadn't thought about that. would basically let you mount a repo as a filesystem backed by postgres. server side branches and change sets are interesting too, postgres already handles concurrent access well so that could work nicely. definitely adding these to the ideas list!
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
good question! the "pgit actual" column tries to compare just the compression algorithms, similar to how the git side only counts the .pack file and not .idx/.rev/.bitmap or filesystem overhead. so both sides strip their "container" overhead to make it a fair comparison. but you're totally right that in practice the on-disk size is what you actually pay. that's why both numbers are in the table. and yes, pgit on-disk is usually larger than git aggressive. the tradeoff is that you get SQL queryability over your entire history, which git just can't do natively.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
sounds interesting
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
hahaha i feel that
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
Yeah, I get that, and I'm fully on your side. SQLite would have been a nice fit. The only downside is the delta compression problem. Creating an extension for SQLite works, but it's slow. I had two options:

1) Do the delta compression and caching and so on on the pgit side and lose SQL queryability (or I need to do my own), or

2) Use postgres

ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
1) In the case of pgit, the "remote" database is a local docker container

2) You can do more complex analyses faster and easier (you don't need to pipe the git outputs) since it's just SQL

but pgit is not meant to replace git.

ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
The problem i faced is mostly importing large repos. But normal use should be fine.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
Accessing specific files is very fast. For sure sub second and most of the times its just a few milliseconds
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
Yes, you also got analysis commands the AI can use. I just did the prompt example before they existed.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
thanks! but it might still need some releases until it's really good. just don't rely on it ;)
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
but the difference between you and an agent is that you naturally know the history of the project if you have worked on it. the AI doesnt.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
I did actually look into writing the extension for duckdb. But similar to SQLite the extension possibilities are not great for what I needed. Though duckdb is a great database.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
in theory yes. you just need to do the full text search across the databases. pgit doesnt support it but at the end its just postgres under the hood.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
so most analyses already have a CLI function you can just call with parameters. for those that don't, in my case, the agent just looked at the --help of the commands and was able to perform the queries.
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
sounds great yes. maybe an SQLite version will come in the future
ImGajeed76··on Show HN: Pgit – A Git-like CLI backed by PostgreSQL
yeah fossil is great, but can fossil import the linux kernel (already working on the next post)