Version Control for Everything
tyoverby.com
tyoverby.com
Lore looked interesting for this purpose. https://github.com/EpicGames/lore
Its private for now, but here is printout of help: http://borg.uu3.net/~borg/?dot
Docs that often change, its completly different story. I love VCS like features where I can slam 'dot status' to see whats going on, commit when I feel I need snapshot, sync to whatever device I want. And because I run DVFS I dont even bother with backups because I have multiple copies of repo with entire history.
Disk space is far from cheap when the value of what's being stored is basically zero.
You can version everything today with something like ZFS snapshoting - and very quickly realize hundreds of gigabytes of noise aren't worth keeping.
I think you can create git plugins to show diffs in different formats for certain file-types (or even open a 3rd party tool to let you visualize it), but it is a lot of work. Most productivity tools don't have anything of the sort.
If you just care about storing the file git LFS + manual text-changelog per binary file works well, but it is annoying to get everyone on your team to work in this workflow (heck, getting everyone and all automation scripts to install LFS is already a pain).
Meaningful deltas and meaningful merges are actually a problem set that is totally and utterly disjointed from keeping the underlying change history. Deltas and merges can always be computed independently of the underlying storage format for versioned data. Git could have indicriminately used binary deltas for file storage and would still work the same on every higher layer.
I mentioned it before, see the following comment https://news.ycombinator.com/item?id=47749916
You can also use a custom difftool or mergetool that understand your formats. For example Beyond Compare https://www.scootersoftware.com/kb/moreformats
LFS is kinda close, but you don't store the files "somewhere convenient": you need support on every git remote to handle the files.
Imagine your git history filled with commits from project managers where every little change in requirements and docs has its own commit. We've tried that and ended up having to rebase our branches multiple times a day while we lost the overview of code changes completely.
The main issue that code can live in multiple branches while development but you can't apply that to docs and requirements where you need to have one centralized source of truth for everyone.
I think we will start to see more adaptation from collaboration tools to provide/accept the info in text format making it suitable for agents as a working solution. Meanwhile, I've seen that having two repos (One for Code and one for docs) is the best solution for the current tools available
I think this would be a great thing to have and know who changed what and when. It seems it’s more of a problem of the tools being able to help the user see only what he needs to see in a certain context.
The main role for projects management is handle the chaotic way of humans communication across different organization without developer being involved in that messy cycle. But having everything in the same repo will remove this decoupling and will let developers get involved in that chaos as it's part of the repo history.
I've worked in several gamedev companies where everything was in a Perforce repository. In fact in one company EVERYTHING was in it, not just code and game data, but also documents, concept art, plans, everything. If it wasn't in P4, it didn't exist.
Yet nobody thought it was a pain (after learning to use P4V at least :-P but that's a different topic) because P4 a) can handle big files just fine and b) you don't really care about the history of the entire repository, but only about the part you're working on. For example personally i had my history tab display just the engine's code directory (and even then i rarely checked it since i don't really care what Joe working on the animation code did, my main interest was on the stuff i worked - and P4V showed me both if someone else is working on the same files i am and if my files have been modified, with easy GUI tools to see both diffs and the file's development history visually - the timeline feature is honestly great for that).
Perforce has some fundamental issues in how it works that creates problems, probably because of some deeply ingrained assumptions stemming from the 90s when it was first made that were never addressed, but it also has a lot of features (especially on the P4V GUI client) that makes working on a huge monolithic repository where everything goes in pretty much seamless.
However, I still think it would be much simpler to extend the tooling for current project management platform to give AI agents the text-based user interface that they needed. This would be much simpler task than developing the tooling for huge repo to include all the chaos of humans communications alongside the code itself.
Actually I would suggest that to one of the projects that I'm currently involved in
This worked pretty well for the team I'm in. Design docs in Google Docs were really hard to keep in sync with decisions and not properly agent accessible. Initially, we were concerned lack of comments would be an issue, but this didn't really came true - a Slack channel does a good enough job. With this flow, the design author makes a decision and review results, agent propagates it to individual areas, and from those to tasks (those are in MD too), and then copies descriptions to Jira. No more "we changed a thing but missed one place that depends on it", or at least not as bad as before.
from those to tasks (those are in MD too)
please?We're doing the design docs in MD thing, but are you saying you've gone a step further and have the backlog represented in git too?
Well, yeah. But it's hard. Also, git is not the solution to this, it's just the screwdriver you have in your hand right now.
It allows you to make new history, and the old history might get garbage collected eventually.
Basically, almost everything is immutable in git, apart from tags and branches: these are mutable pointers to immutable commits.
<proceeds to hammer down the screws with the back of the screwdriver>
The crux of what OP wants seems to be transparent interfaces to the data store backing each of these services. Version control would be about tracking changes when this data mutates, which seems somewhat orthogonal to this stated need.
I wonder whether a nice solution to this would be to have a distributed architecture where each node can publish / subscribe to updates and maintain a local copy it operates on, with conflict resolution for eventual consistency.
PS: Checkout Perkeep for a take on aggregating all your data in one place.
That's the idea of Epiq: https://ljtn.github.io/epiq/
> In my view, the main obstacle here is that without serious dedication, the user experience for humans would be a major downgrade. This isn’t insurmountable, but it would be a lot of work.
I think Epiq does a decent job of addressing this.
Take a look at the second animation on the page above. It even lets you "scrub" your way through the timeline to very quickly glance at what an agent (or junior developer) did to your code recently.
The most you can do is to accelerste subworkflows and review as hell.
Noone wants middle term maintenance nightmares or unreliable code. Can look good the first month, that's it.
I think they have been overselling it so much but you need to know exactly what you are doing and I think for generating code they are pretty bad. And they will improve as they sniff some more data but that's it. Something new, something LLMs do pretty bad.
It's built on top of a content-hashable XML-ish data modeling language called CSTML: https://docs.bablr.org/guides/cstml
Binary document support isn't done yet, but it is planned.
As agents take over more of our workload, it becomes increasingly important to store the data in a dumb and simple format that’s easy to diff and verify.
I talked about some of this at LocalFirstConf: https://youtu.be/2ZEBH5Clrnc
I agree that LLMs introduce a lot of new issues. We’re going to be developing a whole new paradigm, for working with them, and we’re just at the very start.
Would also be nice if GitHub issues/reviews were in sync so reviews are accessible during a GitHub outage.
There is a port for plain git repos named Bee. That one is earlier-stage, but works https://github.com/gritzko/bee