Atuin replaces your existing shell history with a SQLite database
github.com
github.com
Have been putting off pushing it to Github, think I'm gonna do that today.
Though the idea of mistakenly putting credentials in the cli and it ending up in GPT is bothering me a bit.
One reason for this is that different programs have different rules for what should be escaped when, when you write a regex. For example, I think grep is a bit different from vim in this regard
I'd probably suck at it. No matter how frequently I use regex in :ex commands, I always screw up the escapes, substitutions, etc.
If you have big atuin file though, creating indexes for session and cwd are good ideas so we can get the request out quickly..
edit: but much faster than searching or asking LLM for command line parameters instead
explain this please
https://github.com/TIAcode/LLMShellAutoComplete
Forgot to add needs tiktoken, openai and fzf. If someone knows how to do that command line query and replace on other than fish/nushell, please let me know.
I would highly recommend anyone who spend a lot of time in terminal to improve their shell history by using atuin or similar tools, cannot tell how many times it actually helped me to find some important information there, about how I did one thing or another.
-
- [1] https://www.outcoldman.com/en/archive/2017/07/19/dbhist/
For comparison, my (non-work) history since 2012 (plain text) is 181k entries, and takes 25MB. I store the command along with when and where you ran it. (https://www.jefftk.com/p/logging-shell-history-in-zsh)
I'm probably being wasteful of space because I store each session in a separate file. I used to do a lot of data analysis at the shell back in the day, and found it useful to audit sequences of commands afterwards for mistakes, or to turn them into scripts.
Do you rebase your git repos regularly to delete commits older than 6 months?
Didn't they already? eg stemming
"We could learn advanced regexes... or we could just use FTS5".
Hard call. :)
Additionally, everything you describe can be phrased as a regular expression.
If there's something I do repeatedly, I make an alias or a function for it.
I guess you could set up an entropy scanner and flag history lines that have high entropy, but that might not be enough (low-entropy secrets) and might be bothersome (lots of false positives / things that are technically secrets but that you don't care if they're in your shell history).
I suppose this could be solved with either:
- Some kind of modified ssh that sends back the commands to my host
- Some kind of smart terminal that can analyze commands to build up the history
Any ideas on how to practically solve this problem?
https://github.com/cantino/mcfly
I've only used McFly and found it to be pretty great. My only complaint is the default search mode is SQL strings, so you have to use `%` for wildcards. I wish it was a more forgiving, less exact search.
Has anyone used both and could compare them?
In SQL there are sessions and transactions. In shell history - we don't have such entities and this sucks. One could configure their bash/zsh to save history into separate files, but you can't teach them later to source them properly (retaining session awareness).
Sequential context is something we're building very soon. The idea: you search for "cd /project/dir", press TAB and it opens up a new pane in the tui. This will show the command +/- 10 commands. You can then navigate back in time.
This could indeed be useful for managing that one setup command you always have to run in this project dir but never remember the name of
Good to hear, but the point stands: so you deduplicate only for the view, not in the source, and thus the source remains contaminated with duplicates (at very least they cost some disk space and increase seek time).
As for the view: so, since you deduplicate the commands - you can't lookup the context (commands executed before & after)! Because each time the now deduplicated command was executed in the past - it had its own context!
CREATE TABLE history ( id text primary key, timestamp integer not null, duration integer not null, exit integer not null, command text not null, cwd text not null, session text not null, hostname text not null, deleted_at integer,
unique(timestamp, cwd, command) ); CREATE INDEX idx_history_timestamp on history(timestamp); CREATE INDEX idx_history_command on history(command); CREATE INDEX idx_history_command_timestamp on history( command, timestamp );
ls I usually don't care about, but there are directories I regularly cd to, so it would be nice to have those in history.
I can think of a neat heuristic, which is that I often cd to an absolute or home directory, so if the path starts with / or ~ I'll possibly want to cd there again in the future. Changing to a relative path on the other hand, I tend to do more rarely and while doing more ephemeral work.
* prefix with space the commands you don't want to keep
* edit your existing .bash_history by prefixing all commands with space (then reload with history -r)
* after session exit or on next login, edit the new commands at the end (vi ~/.bash_history && history -r)
* use comments on the commands to make it easy to search (use keywords)
* group command lines by category (e.g. package manager, git, ssh, backup, find, dev)
2. I just don't want to be conscious about that every time I write a command. I'd rather edit history after I've finished some work. But that's just too tedious to do manually, I'd like to have some pre-configured heuristics applied automatically, like "never save cd/ls to history", but provide a way to overrule that rule in rare situations.
3. Absolute/partial/symlinked paths - are another separate problem :'(
Coming back several months later to be gifted with several, very similar commands of which only one is right can be frustrating. The history records the failed tries as well as the successes.
Mind, the errors give a place to start, but if it’s far enough removed from the original event it may well be ambiguous enough to send you to a search engine anyway, especially if you have the memory of a goldfish like I do.
The correct command is almost always the last one, so as long as your search results are chronological this shouldn't be an issue?
It's not just commands that are actual for the current session (like ls/cd/cat), it's also incorrect commands or just commands not worth retaining still being saved among with the useful commands.
The most precious is user's time. And when you fzf part of the command with the regularly saved history and would like to re-execute something important - you'll first get a long list that you'll first have to filter to find the command you were seeking.
So to counter your question with another: why store garbage?
Also, when I do want to reference my deep history I often find that seeing the full list of what I was doing is helpful at getting myself back into the frame of mind I was in when I ran the commands originally, which can be more valuable than seeing exactly which commands I ran.
Then why write history at all? Just discard it.
Maybe 3 in every 1k of my command lines is noise. Finding anything useful in there takes effort, so I rarely use anything older than a few days.
"log exit code, cwd, hostname, session, command duration, etc"
One day, you could search for that AWS cli command run specifically on AWS_ACCOUNT_ID=foo
export PROMPT_COMMAND='if [ "$(id -u)" -ne 0 ]; then echo "$(date "+%Y-%m-%d.%H:%M:%S") $(pwd) $(history 1)" >> ~/stuff/logs/bash-history-$(date "+%Y-%m-%d").log; fi'
makes files like stuff/logs/bash-history-2023-05-06.log
that contain 2023-05-06.11:42:38 /home/dv 3737 2023-05-06 11:42:38 cat .bashrc
then you can make some commands to grep through this ls <dir> | grep <pat> | less
vs ls ~ | grep .bak | less
ls .config | grep *rc | less
you get the ideaYou can optionally filter by the current window session, current machine, or commands across all machines you have atuin installed on
At least it is a setting for zsh
# append to .bash_history and reread after each command
export PROMPT_COMMAND="history -a;$PROMPT_COMMAND"
# append to .bash_history instead of overwriting
shopt -s histappendSometimes the rare, long-running commands are the most valuable.
If I set up history software, it should preserve all history, and as early as possible.
It looks interesting, so if I'm overlooking a way to do this with fish, I'd experiment with it.
Presumably you already have a long history of commands from past years in the original format. And if your command line usage is similar to mine, then most of the commands you will use in the future will be covered by your existing history.
So if a few weeks of experimenting with atuin ends with you deciding not to use atuin, then probably you will be fine going back to the old history files that do not include those weeks of activity.
[0] https://sqlite.org/cli.html#export_to_csv [1] https://docs.datasette.io/en/stable/csv_export.html
So you could uninstall it, and you won’t have lost a thing
In fzf, there is no noticeable lag when typing.
Compared to Windows & MacOS – lack of cli and gui full text indexing is a real setback for linux/unix
apropos , locatedb and this are solid domain-specific attempts. Hopefully this expands to indexing a lot more content.
Random side point, it frustrates me so much I cant get the pre-substitution command taht I actually type, at least in zsh. I really want to see the history of what I type! Not just what the outcome was, but learn & adjust with what I type.
A while back I hacked around writing vim commands that would run my current line (or selection) in a terminal, and then copy the command, paste it after the current line, & go to the next line (and another command to do the same but not re-copy the output). That let me better see the work I had done, and I really liked that, but it was a pretty drastic hack born of desperation.
Shell histories! Yay! Anyone have other good tools they use?
Seems more useful in the GPT age, when it could provide an opportunity for conversations about the whole task you worked on. (As well as finding commands based on their output.)
Is there a way to say, “I don’t care about which terminal a command is written to, just log them all chronologically in the same file”?
Would this SQLite approach work that way?
Edit: apparently, they made it possible to disable that behavior and the documentation is much better
> // we can skip past things like invalid utf8
No, you can't! Thanks to some bizarre escaping that happens when ZSH (and BASH, I think) dumps commands to the history file, any command with non-latin1 characters will break here and won't be read, moreover - silently! The other possibility is that you'll import wrong characters.
Grep the ZSH source for "unmetafy", you'll see; I have it extracted here: https://github.com/piotrklibert/zsh-merge-hist/blob/master/u...
- Add both files in this gist to $BASH_HOME.
https://gist.github.com/chmaynard/dbcaed11534dc54bdc90856d18...
- Install .bash-preexec.sh in $HOME (see https://github.com/rcaloras/bash-preexec).
- Append this command to your shell startup file (e.g. $BASH_HOME/.bashrc):
source $BASH_HOME/bash_history.sh
- Open a new bash window. You should see the message "bash-preexec is loaded."Enjoy! Please share your feedback in the gist comments section.
(Memo to self: Package this up in a repo with a proper README.)
Is there a way to import all my existing history from `.zsh_history`? It's a real pain to have to start from scratch.
I've also seen a couple of PRs such as [1] that consider Windows support a dead-end as of 2021, so I assumed it to be a dead-end as well. Maybe this user is on WSL? That'd seemingly work as it can get bash/zsh or anything running there.
I installed it and simple commands take noticeably longer. Is there any way to make it faster?
- how active the development is[1]
- having a distributed option (McFly does not have one)
1. Seeking co-maintainers: I don't have much time to maintain this project these days. If someone would like to jump in and become a co-maintainer, it would be appreciated!
Second, I am in awe of how good your documentation is and how well you communicate about atuin to the world at large.
Does Atuin offer any features to toggle the capture of commands into its DB? Being able to opt-in or opt-out of Atuin history on a per-command basis would be pretty useful, especially because there is also the atuin sync feature.
I usually work with sensitive information inside a tmux session because, in the default bash configuration, most commands run in tmux never make it into bash history (I believe the last pane to exit is the only one that does make it). It seems I would have to manually go in and drop rows from the DB if I set up Atuin.
One of my products, bugout, has a command called "bugout trap". Not trying to push bugout here, but thought Atuin might benefit from some of the lessons we learned:
1. Because bugout trap is opt-in (you have to explicitly prefix your command with "bugout trap --", it also allows users to specify tags to make classifying commands easy. This is really useful for search - e.g. you can use queries like "!#exit:0 #db #migration #prod" to find all unsuccessful database migrations you attempted in your production environment.
2. bugout trap has a --env flag which gives users the option of pushing their environment variables into their history. This is really useful for programs that use a lot of environment variables. The safest way to use this is to first trap commands into your personal knowledge base with --env, then remove or redact any sensitive information, and only then share (in case you want to share with a team).
3. We thought that sharing would be useful for teams to build documentation on top of. Even we ourselves have very little adoption of that use case internally. We use it to keep a record of programs we run in our production environment (especially database migrations).
4. bugout trap also stores data that a program returns from stdout and stderr - this has been INCREDIBLY useful. I do want to add a mode that makes the capture of output optional, though, as currently bugout trap is unusable with things like server start commands which run continuously.
5. In general, I have found that command line history is very personal and private for developers so collaborative features are going to rightly be seen with skepticism.
Hope that helps anyone building similar tools.
$ bugout trap --help
Wraps a command, waits for it to complete, and then adds the result to a Bugout journal.
Specify the wrapped command using "--" followed by the command:
bugout trap [flags] -- <command>
Usage:
bugout trap [flags]
Flags:
-e, --env Set this flag to dump the values of your current environment variables
-h, --help help for trap
-j, --journal string ID of journal
--tags strings Tags to apply to the new entry (as a comma-separated list of strings)
-T, --title string Title of new entry
-t, --token string Bugout access token to use for the requestThe docs need updating as we support far more data sources now!
one thing i haven't seen yet (correct me if i'm wrong...) is an easy way to get all this stuff to magically appear on a new machine you've ssh'd into for the first time. i've hacked up my own in the past but that's got issues with tunneling and multi-hops. anyone know a solution to this? maybe a feature request?
TBH I was thinking about doing this for a while now. History of my shell, which now is 3.8MB in size, is one of my competitive advantages as a developer (and a very nice thing to have as a power user). It accumulated steadily since ~2005, and with fzf searching the history is much faster than having to google, as long as I did what I want to do now even once in the distant past. I even wrote a utility to merge history files across multiple hosts[1], so I don't have to think "where did I last use that command", as I have everything on every host. The problem with this, however, is shell startup time. It started being noticeable a few years ago and is slowly getting irritating. The idea of converting the histfile into sqlite db crossed my mind more than once due to this.
For comparison: I use extended ZSH history format, which records a timestamp and duration of the call (and nothing else), and I have ~65k entries there, with history file size, as mentioned, 3.8MB. It could be an order of magnitude larger and I still wouldn't care, as long as it loads faster than it takes ZSH to parse its history file.
Currently, the tool reads ~/.mergerc, which is a JSON file with a list of SSH hosts to SCP history to and from. As long as the history file is in the same place (it tends to be on hosts that I setup, and otherwise I check in default locations) and the host has an entry in ~/.ssh/config, the tool will work. It's really just a wrapper for a few SCP invocations plus a history file (extended) format parser.
Changing servers is just a change in the config file, but it's also helpful for changing jobs, because I can quickly add a bit of filtering before the merging happens. I had to erase some API keys and such a few times, adding `filter` call here: https://github.com/piotrklibert/zsh-merge-hist/blob/master/s... took care of it.
> what is your workflow around maintaining these one off ad hoc "developer boost" type tools?
Good question. I don't have such workflow, at all. When I commit to write something like this, I try to make sure that it has a scope limited enough so that it can be "completed" or "done". In this case, the tool builds on SSH/SCP and a file format that hasn't changed in the last 20 years (at least). So, once I had it working, there was nothing much to do with it after that. The only change I had to do recently was changing `+` to `*` in the parser, because somehow (not sure how, actually) an empty command made it into the file. But that's all I had to do in 5 years time.
I'm not as extreme, but suckless.org philosophy appears to work well here. Here's another example: https://github.com/piotrklibert/nimlock - it's a port, done because I wanted to do something in Nim, but it worked for me for years and I suspect it still works now (after going full remote I stopped needing it). There's nothing much that could break (well, Wayland would break it, but I don't use it), and so there's not much you need to do in terms of maintenance.
As for language choices - these are basically random. I made the zsh-merge-hist in Scala simply because I was interested in Scala back then. I have little tools written in Nim, OCaml, Racket, Elisp, Raku - and even AWK (pretty nice language actually) and shell. That's another reason why making the tools any more complex than what's absolutely necessary would be a problem: the churn in the ecosystems tends to be too high for me to keep track of, especially since I'd need to track 10 of them.
EDIT: I forgot, but obviously the most important "trick" is not giving a shit if these things work for anyone else but me :D
> I'm checking it out
If you have Java installed, `./gradlew installDist` should give you `./build/install/bin/zsh-merge-hist` executable to run. The ~/.mergerc (on the host the tool runs) should look like this:
{
"hosts": ["host1"],
"tempDir": "/tmp/zsh-merge-hist",
"sourcePath": "mgmnt/zsh_history"
}
where `sourcePath` is a path to history file relative to the home directory.Let's say I run a command where I've pasted in a credential from my password manager: ` some-cli login username my-secret-password` (note space at beginning)
Normally this would prevent the command from getting saved in any meaningful way in my bash history, so that if I later run a malicious script, it can't collect secrets from my bash history.
With the bug here, it sounds like atuin would prevent that entry from being stored in the sqlite store, but it would still be in my shell history?
If so, this is really significant, and would stop me from using Atuin. Not letting users know about this behaviour is incredibly negligent, and honestly erodes my trust in Atuin to consider user security in general.
You can also host your own backend
Currently we rely on Github sponsors as well as our own additional funding
eg receives incoming shell history to store in the backend, and maybe do some searches/retrievals of shell history to pass back? eg for shell completion, etc
If that's the case, then I'm wondering if it could work in with online data stores (eg https://api.dbhub.io <-- my project) that do remote SQLite storage, remote querying, etc.
There's a PoC that allows it to work with SQLite too for single user setups - and we are thinking of switching to a distributed object store for our public server since we don't need any relational behaviour.
One of our developer members mentioned they're learning Rust. I'll point them at your project and see if they want to have a go at trying to integrate stuff.
At the very least, it might result in a Rust based library for our API getting created. :D
As a data point, we're using Minio (https://github.com/minio/minio) for the object store on our backend. It's been very reliable, though we're not at a point where we're pushing things very hard. :)
> You may use either the server I host
Right. And isn't this what home dirs are for?
Why is our localisation relevant/quoted?
So even if we were being controlled, you still can be confident that we can't do anything with your data - all we can see is how active you are, that is until someone finds a way to quickly break xsalsa20poly1305.
At the end of the day, Ellie and I work on this because these features actually improve our workflows. The directory search feature is probably my favourite, and the sync feature is the key feature Ellie wanted to begin with.
As an aside, the part I like the most about our field is the ability for a single or two devs to build themselves the tools they exactly need, and potentially share it to the community with low friction.
i am not a new kid on the block either but I was looking for such a tool for a very long time. I think distributed shell history across all of my servers is a big win.
How can that work. Nobody is going to pay any ransom for just a shell history, and there are ways to get it out of a SQLite database. Wouldn't it be simpler just to encrypt the original .bash_history?
That stored context can then be used to query the database (e.g. filter the history to only show commands that were executed in the cwd).
These queries are the point of using sqlite, not anything security as far as I can tell.
I am not sure how many writes per second you have on your shell history, but sqlite is not only up to the task, everything else is overkill.
Additionally not having to run a database, but having your history in a file has advantages for that usecase as well.