131 karma · joined December 7, 2020
For fun, I had put together a GitHub bot for this purpose a while ago. It indexes all existing issues in a repo, creates embeddings and stores them in a vector DB. When a new issue is created, the bot comments with the 3 most similar, existing issues, from vector similarity.
In theory, that empowers users to close their own issues without maintainer intervention, provided an existing & solved issue covers their use case. In practice, the project never made it past PoC.
The mechanism works okay, but I've found available (cheap) embedding models to not be powerful enough. For GitHub, technology-wise, it should be easy to implement though.
For a tool I’m writing, the tree-sitter query language is a core piece of the puzzle as well. Once you only have JSON of the concrete syntax trees, you’re back to implementing a query language yourself. Not that OP needs it, but ast-grep might?
cat file.py | srgn --py def --py identifiers 'database' 'db'
will replace all mentions of `database` inside identifiers inside (only!) function definitions (`def`) with `db`.An input like
import database
import pytest
@pytest.fixture()
def test_a(database):
return database
def test_b(database):
return database
database = "database"
class database:
pass
is turned into import database
import pytest
@pytest.fixture()
def test_a(db):
return db
def test_b(db):
return db
database = "database"
class database:
pass
which seems roughly like what the author is after. Mentions of "database" outside function definitions are not modified. That sort of logic I always found hard to replicate in basic GNU-like tools. If run without stdin, the above command runs recursively, in-place (careful with that one!).Note: I just wrote this, and version 0.13.2 is required for the above to work.
It allows you to grep inside source code, but limit the search to e.g. “only docstrings inside class definitions”, among other things. That is, it allows nesting and is syntax aware. That example is for Python, but the tool speaks more languages (thanks to treesitter).
https://github.com/alexpovel/srgn/blob/main/README.md#multip...
Caught me.
> And how is it less tedious to have to select previously typed text all the time?
This is tedious, but I have that automated (AutoHotKey). So a single, AHK-managed hotkey does the equivalent of:
CTRL + SHIFT + HOME
CTRL + C
feed into tool, paste back
CTRL + V
So once done writing, I press that single button and it's done (CTRL+SHIFT+HOME select all text from cursor to beginning). To me, that's a better tradeoff than fiddling with compose keys, which I find to break flow. For very short text, compose key is possibly better; but again, once in AHK, it's a single shortcut. So once more than 1 compose key combination is needed, it's "worth it". But you're right: this is a custom setup and might not work for everyone.[0]: https://github.com/alexpovel/srgn/blob/0008cce1c71f0d83f6a31...
It grew out of a niche, almost historical need: using a QWERTY keyboard, but needing access to German Umlauts (ä, ö, ü, as well as ß). Switching keyboard layouts is possible but exhausting (it's much more pleasant sticking to one); using modifier keys is similarly tedious, and custom setups break and aren't portable.
So this tool can do:
$ echo 'Gruess Gott, Poeten und Abenteuergruetze!' | srgn --german
Grüß Gott, Poeten und Abenteuergrütze!
meaning it not only replaces Umlauts and Eszett, it also knows when not to (Poeten), and handles arbitrary compound words. Write your text, slap it all into the tool, it spits out results instantly. The original text can use alternative spellings (ou, ae, ue, ss), which is ergonomic. Combined with tools like AutohotKey, GUI integration through a single keyboard shortcut is possible. See [0] for a similar example.A niche need I haven't yet come across someone else having as well! (just the amount of text explaining what it's all about is saying a lot in terms of specificity...)
The tool now grew into a tree-sitter based (== language grammar-aware) text manipulation thing, mostly for fun. The bizarre German core is still there however.
[0]: https://github.com/alexpovel/betterletter/blob/c19245bf90589...
Thank you for the feedback! That sounds good, I'll add that.
I've spent a lot of time trying to find similar tools, and even list them in the README, but `AST-grep` did not come up! I was a bit confused, as I was sure such a thing must exist already. AST-grep looks much more capable and dynamic, great work, especially around the variable syntax.
Idea: renders your resume as pretty terminal output. Others can view it in their own terminals:
curl -L ancv.io/heyho
Pipe to a pager for best viewing. Yes, it's just a nerdy gimmick with almost no real use!I provide a GCP-hosted server that works off GitHub gists (where your resume can live in JSONResume form). However, self-hosting is a first-class citizen and easy to use as well.
Yeah, I had looked into these but for some reason that didn't work. Don't remember why.
> Looking through the repo I wondered why you would commit the complete German dictionary weighing in at over 30 MB, whereas you only need a small fraction, the words containing the umlauts (or their false matches). Surely this would be a huge performance boost?
Yes! It would be performance boost. In fact, I had a "caching" sort of functionality in the tool before. The whole dictionary is shipped (because that makes it much easier and there's almost no risk of wrong-doing just copy-pasting a word list, plus it compresses well enough), but then a list containing only special characters will be generated on first use if it doesn't exist yet.
As you noted, a lot of words do contain special letters, so the "complexity" wasn't worth it to me and I removed that. Could be brought back anytime, but it's fine for now.
` [ ] \ / { }
very easily available is a blessing. The German QWERTZ keyboard has triple occupation on some keys, which is not ergonomic and harder to type fast with.Anyway, both Linux and Windows offer fast switching between installed keyboard layouts/languages using SUPER+SPACE. This is needed in e.g. emails, where I still need Umlauts. It's just much easier to read that way. However, switching back and forth constantly is completely overwhelming and not viable. However, in German, there are perfectly and officially (?) acceptable alternative spellings for our special "Unicode"-characters. They can be typed using plain ASCII, aka a QWERTY keyboard.
So, I wrote a script to read in any text, combined it with AutoHotkey on Windows and now have a tool that, at the touch of a button, replaces selected text using alternative spellings (gruen, Duebel, Faehre) with their correct versions (grün, Dübel, Fähre). The tool could be extended for other languages rather easily. I've been using it for over a year now and recently got to release it properly on the cheese shop:
pip install betterletter
(https://pypi.org/project/betterletter/)Before putting this together, I had looked around for an existing tool. To my surprise (there's always something!), I found nothing. I guess this scratches a too specific itch: using QWERTY but wanting proper spelling quickly, while remaining on QWERTY as to not have a mental breakdown and stay at full typing speed.
After writing, select everything (CTRL+SHIFT+HOME works well), hit shortcut, text will be replaced. This takes about 2 seconds, much faster than switching keyboard layouts back and forth. If this ran as a daemon with the dictionary loaded into RAM already, the script could run almost instantaneously (most of the 2 seconds is IO, reading from disk), in linear time according to the text input size.
That is only because their templates are years behind the curve and they are slow to update. It is not an argument for the advantages of pdftex, aside from its stability, gained over many decades.
LuaTeX has been nothing but stable for me, so from a technical standpoint, there is no reason not to switch.
As far as scientific papers go, the publishers and editors probably value stability and backward-compatibility (I would).
This is simply an artefact of times past and has no technical relevance nowadays. LuaTeX allows dynamic allocation, with the available system RAM as the upper limit (so effectively, no limitations in everyday usage).
Now, I could not find a mention of memory handling in the XeTeX reference manual [4]. People are using tricks like `tikzexternalize` with xelatex [5, 6]. Especially the first point makes me think XeLaTeX inherits base TeX memory handling/limits, but I cannot confirm this.
I just know that all my problems disappeared when switching from XeLaTeX to LuaLaTeX.
Lastly, see here [7] for a comprehensive (albeit somewhat anecdotal) list of advantages of LuaTeX over XeTeX. Of that list, `microtype` is another significant functionality I rely on.
[0]: http://www.tug.org/texlive//devsrc/Master/texmf-dist/doc/con...
[1]: https://tex.stackexchange.com/q/7953/
[2]: https://tex.stackexchange.com/search?q=tex+capacity+exceeded
[3]: https://tex.stackexchange.com/a/482560/
[4]: http://mirrors.ctan.org/info/xetexref/xetex-reference.pdf
[5]: https://tex.stackexchange.com/q/438131/
Sadly, I think basing off XeTeX and not LuaTeX is a mistake. Certainly renders it unusable for me. Having Lua integration is just great.
Also, `lualatex` does not have some of the limitations of `xelatex` (memory limitations, `contours` package, ...), but I guess this XeTeX reimplementation can work on removing those implementations, so that only lack of Lua integration remains.
Also, like another person said, not having biber breaks my workflow as well, which specifically tries to leverage the "latest and greatest" of what LaTeX has to offer [0]: `pdflatex` is obsolete, so `lualatex` it is. `nomencl`, `makeindex` etc. are obsolete, so `glossaries-extra` it is. `bibtex` is obsolete, so `biber` it is. Throw in `latexmk` for automatic compilation (which the tool presented here does too, which is a biggie! [1]) and CI/CD and you have a 1970s tool in 2020s attire. Lua rounds off the picture.
Among other things, this given Unicode-native (gasp) code/documents, and great automation capabilities (`latexmk`, CI/CD, Lua).
I think a modern TeX engine reimplementation should support all of the above, which are arguably the best modern options there are.
[0]: https://collaborating.tuhh.de/alex/latex-git-cookbook [1]: I wonder if the logs are available though? aux, blg etc. are important for debugging and shouldn't be dropped outright.