GQL – Git Query Language
github.com
github.com
Download clickhouse: curl https://clickhouse.com/ | sh
Check out documentation for git-import: ./clickhouse git-import --help
Then the tool can be run directly inside the git repository. It will collect data like commits, file changes and changes of every line in every file for further analysis. It works well even on largest repositories like Linux or Chromium.
Example of a trivial query:
SELECT author AS k, count() AS c FROM line_changes WHERE file_extension IN ('h', 'cpp') GROUP BY k ORDER BY c DESC LIMIT 20
Example of some non-trivial query - a matrix of authors, how much code of one author is removed by another:
SELECT k, written_code.c, removed_code.c, round(removed_code.c * 100 / written_code.c) AS remove_ratio FROM ( SELECT author AS k, count() AS c FROM line_changes WHERE sign = 1 AND file_extension IN ('h', 'cpp') AND line_type NOT IN ('Punct', 'Empty') GROUP BY k ) AS written_code INNER JOIN ( SELECT prev_author AS k, count() AS c FROM line_changes WHERE sign = -1 AND file_extension IN ('h', 'cpp') AND line_type NOT IN ('Punct', 'Empty') AND author != prev_author GROUP BY k ) AS removed_code USING (k) WHERE written_code.c > 1000 ORDER BY c DESC LIMIT 500
Does this check the useragent to change the response? Clicking that link shows their home page.
;)
If I run curl on Windows, do I get this script? A PowerShell version?
Why not make it https://clickhouse.com/linux-installer?
(Pull request is in ... it should be deployed on Monday and you can use https://clickhouse.com/install.sh ). Love the feedbacks, please keep them coming!
Because someone may not have curl and use another tool your server doesn't know.
Also it seems up arrow doesn't work to recall the previous command?
Lmk what you think :)
The commit example:
select name, count(name), from commits group by name
is actually: git shortlog -sn
The tag example: select * from tags
is actually: git tag
The branch example: select * from branches
is actually: git branchThough if I regularly needed the information that this tool retrieves I would probably have memorized the relevant CLI commands by now.
2. SQL is more capable
3. SQL is more well-known
Yep. And to list worktrees you use:
$ git worktree
error: need a subcommand
Oh woops. But to show all refs instead of just the branches you surely just: $ git show-ref
0003692409f153dd725b3455dfc2e128276cfbe2 refs/branchless/0003692409f153dd725b3455dfc2e128276cfbe2
No. But I can just cut(1) out that SHA1...These SQL-like commands are probably meant to be familiar and guessable. Not terse. And then you can change the query instead of using all sorts of utility commands in some ad hoc pipeline (or look up whatever switch they threw in to turn on and off things like SHA1 as first column). If you don’t like the latter.
In addition SQL is a powerful functional programming language with joins! So that is handy too.
SELECT name, COUNT(name) AS commit_num FROM commits GROUP BY name ORDER BY commit_num DESC LIMIT 10
something like that is more complex and less easily gleaned by the git gui. or something else that's dead simple to remember how to do in sql vs probably having to open up the git manpages or cobble some script together to do: SELECT name, commit_count FROM branches WHERE commit_count BETWEEN 0 .. 10 git shortlog -sn | sort -r | head -n 10
If I run the 2nd command example against the chromium repo, it takes several minutes where a shell script using git for-each-ref takes about 8 seconds. I would also argue that this might be better expressed as commit_count <= 10 but I realize you're also using an example. This is the script I ran: git for-each-ref --format='%(refname:short)' refs/heads/ | while read branch; do
commit_count=$(git rev-list --count $branch)
[ "$commit_count" -le 10 ] && echo "Branch: $branch, Commit Count: $commit_count"
doneI was using the more complex examples from their README as well (except added a name to the second one to make it more useful).
> I would still argue I can get the same results with less typing.
so what? we type words and sentences much quicker than having to type control characters like - and |. also, you still have to know the invocations to sort and head off the top of your head without looking them up whereas a simple sql select statement is something most programmers would know how to do in their sleep.
git for-each-ref --format='%(refname:short)' refs/heads/ | while read branch; do
commit_count=$(git rev-list --count $branch)
[ "$commit_count" -le 10 ] && echo "Branch: $branch, Commit Count: $commit_count"
done
I think we can agree this is much more complicated to write than: SELECT name, commit_count FROM branches WHERE commit_count BETWEEN 0 .. 10
> If I run the 2nd command example against the chromium repo, it takes several minutes where a shell script using git for-each-ref takes about 8 seconds. I would also argue that this might be better expressed as commit_count <= 10 but I realize you're also using an example. This is the script I ran:if speed matters then your script is fine to bang out if you need it but for one off queries that you know to be correct (look at how simple the second query is) it is probably fine and most people aren't working on chromium so speed is probably never going to be an issue in the first place.
Imagine instead if there was a standard xQL that could run on any tool.
Nonetheless nice work
I don't see it in the code, but have you considered creating secondary indexes to speed up some operations? Things like looking up the path of a deleted file in a large repo?
https://repo.mercurial-scm.org/hg/help/template
https://repo.mercurial-scm.org/hg/help/revset
(but arguably enables few things like aggregations that would require the churn extension)
I would like to request a feature --> Can you allow the output produced by GQL to be in different serializable formats, like JSON?
I see a lot of potential to this if the GQL output could be piped to C# LINQ queries.
Those are two entirely separate things.
EDIT: And neither of these should be confused with "Graph Query Language (GSQL)" by TigerGraph: https://www.tigergraph.com/gsql/