Show HN: How did I live without Pipe Watch?
kylheku.com
kylheku.com
Why do I need this? What problem does it solve for me? I know it samples pipes, I don't know why I would want something that does.
A bit like if ‘pv’ (https://linux.die.net/man/1/pv) and a pager (eg ‘less’/‘more’/etc) had a child.
I’ve typically solved this kind of problem with ‘tee’ but this looks a much cleaner approach.
It could be a very useful tool for any Sysadmins who still find they need to occasionally SSH into a server for real time log parsing (which is a dying habit these days). Or for DevOps and Developers needing to debug.
When pw isn't updating the display, it's still streaming and discarding the data, which is a big semantic difference. It means that the source of the data will not block, for one thing. But if you need to see all the data or search through it, then pw is the wrong tool, at least by itself. You need to make other arrangements to capture the complete data.
Hmmm, I don't know what that says about me. Checking logs is 2nd nature to me. I guess that makes me old if the "kids these days" are doing something else, but what is the something else? Just not looking at the data and spinning wheels debugging? New tools available? I'll admit I've never looked for anything other than the location of a log of a service I'm investigating.
As a matter of fact, we do not have SSH access to our servers. Everything happens over a web interface that has been secured.
So it's not that checking logs is a dying habit, it's process of SSHing onto the server and running shell commands to query them.
Though I'm not suggesting the latter is bad practice. In some domains it's suboptimal but in other domains it is a pretty reasonable response to fault diagnosis. IT is a broad field :)
Yes it is. Today's assumption seems to be that everything is being done "web scale" with hundreds of servers in multiple datacenters across the globe.
I'm not so fortunate/unfortunate (depending on one's viewpoint) of having to deal with that kind of scale. I have less than 10 servers I maintain so a lot of these new fancy viewers is overkill for me.
Plus if you have any kind of auditing and compliance where you need to retain tamper proof logs for any period of time then you've already done the other half the work by streaming those logs off the servers.
There are several layers to the complexity of "web scale orchestration" however, and you definitely can get the same tooling/features on a much smaller scale with less complexity involved.
I.e. you can rent a few baremetal servers from hetzner, give them an internal lan interface and setup k3s on that internal network. Now you can add capabilities as you want and need them: I.e.
* deploying longhorn for a distributed storage backend with easily configured backup.
* Certmanager for cluster wide https certificate provisioning
* Promstack to get healthchecks etc
* Loki for the just now mentioned logging aggregation
Each of these capabilities are usually just one deployment command (kubetcl apply/helm install) and will enable this capability for all existing and future deployments on the same cluster.If you have a weekend to experiment and want to explore what it actually feels to use that kind of tooling I can only encourage you to maybe try it out. The whole experience has gotten pretty nice over the years, and while it is pretty much the same you already know... Just differently, it does solve it in a very reusable way, which makes it ultimately (at least now after about 1 year of usage) much more straightforward to figure out then just a bunch of shell scripts
After I got over that it's done wrong, (as in: not what I've come to know/expect) I kinda got to like it, especially for small scale stuff (a 4 node HA cluster with 64gb ram each costs slightly below 200€ per month on hetzner, and they're using nvme storage, so really fast IO)
I never suggested that was the case. I just don’t consider then in scope for this topic.
I have managed numerous such devices over my career. In fact most of my home networking gear would fall into that category of “produces logs and SSH but not a server” but none of which runs anything UNIX-like (so you couldn’t install this tool even if you wanted to). However they do support a bunch of different tooling for pushing logs and metrics over the network. Which then brings me back to my earlier point about log collection.
Also if we take your example of embedded devices with low memory and storage, then installing additional utilities like this one when you can already do the same thing (albeit less conveniently) with busybox probably isn’t a good idea. So they’d be out of scope for this thread too.
I’m not going to say that there’s no such thing as a non-server platform that produces logs that this tool would be useful for (I don’t believe in absolutes!) however it is reasonable to assume that servers and desktops, ie more classical computing devices, are a vastly larger demographic for this tool. Hence why I focused my examples around that.
I regularly install debugging tools onto target devices (strace, tcpdump, ...) even though those don't go into a production image.
Currently, I'm using pw on an embedded system, for capturing detailed traces from an application. It's very useful because it decouples the rate of the traces from the serial port, while still letting me get the info I need out of it.
The stripped executable is tiny: 20 kilobytes. It's so small I'm stashing it in a configuration flash partition which survives re-flashes of the kernel and filesystem, so I have the utility there no matter which image I load onto the target.
This particular man page is from 2012 or earlier.
Or other things!
tcdump -i <if> -l | pw # -l for line buffering
Now you have an interactive network monitor that, in some ways, has better ergonomics than Wireshark.
Want IPV6 only? :g IP6
done. Okay, only packets from bob to alice: :g bob.*>.*alice
Only UDP: :g UDP
Forget UDP, how about http: :r
:g http
Kind of thing.Use it only for debugging builds and not for production builds (obviously).
NAME
man - format and display the on-line manual pages
SYNOPSIS
man [-acdfFhkKtwW] [--path] [-m system] [-p string] [-C config_file] [-M pathlist]
[-P pager] [-B browser] [-H htmlpager] [-S section_list] [section] name ... pw stands for Pipe Watch. This is a utility which continuously reads textual input from a
pipe or pipe-like source, and maintains a dynamic display of the most recently read N
lines.
If pw is invoked such that its standard input is a TTY, it simply reads lines and prints
them in its characteristic way, with control characters replaced by caret codes, until
end-of-file is encountered. Long lines aren't clipped, and there is no interactive mode.
The intended use of pw is that its standard input is a pipe, such as the output of another
command, or a pipe-like device such as a socket or whatever. In this situation, pw ex-
pects to be executed in a TTY session in which a /dev/tty device can be opened, for the
purposes of obtaining interactive input. The remaining description pertains to this in-
teractive mode.
In interactive mode, pw simultaneously monitors its standard input for the arrival of new
data, as well as the TTY for interactive commands. Lines from standard input are placed
into a FIFO buffer. While the FIFO buffer is not yet full, lines are displayed immedi-
ately. After the FIFO buffer fills up with the specified number of lines (controlled by
the -n option) then pw transitions into a mode in which, old lines are bumped from the
tail of the FIFO as new ones are added to the head, and refresh operations are required in
order to display the current FIFO contents.
The display only refreshes with the latest FIFO data when
1. there is some keyboard activity from the terminal; or
2. when the interval period has expired without new input having been seen; or else
3. whenever the long period elapses.
In other words, while the pipe is spewing, and there is no keyboard input, the display is
updated infrequently, only according to the long interval.
The display is also updated when standard input indicates end-of-data. In this situation,
pw terminates, unless the -d (do not quit) option has been specified, in which case it pw
stays in interactive mode. The the end-of-data status is indicated by the string EOF being
displayed in the status line after the data.Then they show a fancier example of a specific pattern of recvmsg(11,...) happening a specific amount of lines below poll(). That ends up showing the flow of data to a specific socket.
Or without triggers, it only shows you the latest when you tap the space bar or some other key.
tail -f logfile | pw
the main behavior change that pw will introduce is that during periods when the logfile is rapidly growing, it will prevent spewage to the terminal. It will only update when there is a lull in in the output exceeding the short timeout (default: 1 s), or else once every long timeout (default: 10 s). (These values were hard coded at first, then they became command line options, then dynamically adjustable too).The development of pw was inspired by a discussion in the GNU Coreutils mailing list by a user wanting behavior along these lines from tail itself.
Then the other feature ideas came.
For instance pausing: we can suspend the display indefinitely. With tail -f, that's easy: you can just kill it. You can't kill a program that's taking input from a pipe without breaking the pipe, though. You can pause the terminal output, but not indefinitely without causing a backlog that could block the producing application, and when you resume, all the backlogged spewage will show.
The reasons I like it are that it's "an interactive streaming text viewer", along the lines of a readonly text editor designed for presenting inhuman amounts of data in a way that's useful for us humans. Think of the times you want to use your favorite text editor to interactively search and inspect streaming data, but can't because the stream is continuous.
- doesn't stream characters to the screen so fast that humans can't read it
- allows for interactively searching the stream and showing context
- operates on a large ring buffer, so all of this can be done without ever attempting to read a fixed subset or the entire stream
It reminds me of what Casey Muratori did for refterm: https://youtu.be/hNZF81VYfQo
This can relieve a pain point if you ever need to work on a serial port, and as you are typing in a command, the other side sends text, which overwrites the characters you were typing (depending on how the local echo is configured). It reminds me of a line based serial terminal emulator, where the input is a separate text box, and the output is a larger box above it.
Oh, and it also seems to have methods of sampling the data as it comes in, so that it can be configured to only show the messages you are looking for, and also counting them?
The source file is also tiny: https://www.kylheku.com/cgit/pw/tree/pw.c
Very Unix
- pw currently misbehaves if you Ctrl-Z suspend it and then try "bg". It doesn't handle the SIGTTOU signal telling it that it's not allowed to send to the TTY due to being in the background. (It won't die, but you may have to suspend it again and do "fg").
- the ideal behavior would be: if we Ctrl-Z, bg, the program runs and just reads the pipe in the background.
With that functionality, we can take any program that spews logs onto standard output and just throw it into the background indefinitely without having it block.
$ spewing-program | pw &
[1] 123
$ bg
[1]+ spewing-program | pw &
To see the latest bit of logging from the program, we foreground it.A program which produces detailed logging that cannot be turned off or adjusted dynamically will benefit from this.
I did that to a different program and misremembered.
No, scratch that; I did that in the man page source code, but forgot the C source file.
[Issue fixed now.]
Before setting setting the trigger, I unpaused the display. Every input event from the TTY also refreshes the display in free-running mode. So as I'm typing /poll[Enter], there is display activity with each key.
Then when the trigger is set, things start going due to the frequent triggering.
Things might have been less confusing at the start of the video if I had shortened the long timeout.
Said he while using find, dd, rsync, tar and others every day. But in all seriousness new cmd tools should be given a chance, it sometimes feels like people believe posix tools were perfected in the 80s and further development makes no sense.
For example, last week I was grepping through some security logs looking for anomalous logins. There were a bunch of other security events in the log file. Eventually my command looked like `cat /var/log/auth.log | grep -v Disconnected | grep -v cron:session | grep -v 'Invalid user'` and so forth until I only had the "interesting" lines left.
Maybe the right approach to this tool would be to put the lines into a tree based on the longest substrings.
grep -v -e Disconnected -e cron:session -e "invalid user" /var/log/auth.log
would also suffice (or you can use -E which would enable the usual "Disconnected|invalid|cron").https://www.electronics-notes.com/articles/test-methods/osci...
This was used as an example on monitoring a MySQL restore. The command was:
pv sqlfile.sql | mysql -uxxx -pxxxx dbname
It seems that pv copies the file to standard output, MySQL then consumes this. The normal command is mysql -u user -p pass < sqlfile.sql
But this shows no progress output.The difference being, pv is more like a fancy version of cat (and more, but that’s the primary use case). This samples data that is being piped around, so you can see the actual data, not just the progress.
alias lzf="fzf +s --tac --bind 'enter:select-all+accept' -m"
With the added bonus that you can press return to release the filtered lines to stdout. Really useful for interactively filtering logs.If there is some dynamic pattern in the logs that repeats (or that you can cause to repeat by executing some test steps on the software which produces the logs), you may be able to use pw to latch on to that pattern.
E.g. suppose there is detailed function enter/leave tracing in the log. You can discover that certain functions are traced when you do something, and then filter for that material to remove unrelated spewage, and perhaps set up a trigger to capture the scenario you are reproducing.
Once you know that, you can use other tools to find the same thing. You know what to look for in off-line analysis of a complete log.
This is not just for logs. If you pipe
tcpdump -i <eth> -l | pw
you basically get an interactive network monitoring utility.E.g. add:
:g thishost.*>.*thathost
and now you're just seeing packets from thishost to thathost.Great this tool looks cool. The UI turns me off. Just like almost every Unix UI. It's not about how good the UI is either. It's about the fact that every new tool requires reading MAN pages.
More tools need to be like nano because I simply don't have time to read the MAN for every single thing.
From the man page. "watch - execute a program periodically, showing output fullscreen".
I use it as "watch -5 -d <command>". "-5" is to run the "command" every 5 seconds and "-d" is to highlight the difference from the previous output.