HNHacker News
TopNewBestAskShowJobs

pixelbeat

1,141 karma · joined September 12, 2008

http://www.pixelbeat.org/ http://twitter.com/pixelbeat_
submissionscomments
pixelbeat··on ASCII Art Weather
Oh cool, this is using my ansi -> html conversion script!

http://www.pixelbeat.org/scripts/ansi2html.sh

pixelbeat··on Vitter's reservoir sampling algorithm D: randomly selecting unique items
GNU coreutils' shuf since v8.22 (Dec 2013) implements this to minimize memory usage when sampling large files or unknown numbers of items read from a pipe.

For example this command never finishes, but consumes a constant (small) amount of mem.

  seq inf | shuf -n1
pixelbeat··on UTF-8 Encoding Debugging Chart
I previously wrote about this common double encoding issue at http://www.pixelbeat.org/docs/unicode_utils/ which references tools and techniques to fix up such garbled data
pixelbeat··on Docker Official Images Are Moving to Alpine Linux
Note the coreutils-single package in Fedora rawhide which can be used to get coreutils down to about 1MB. Details on that change at http://pkgs.fedoraproject.org/cgit/rpms/coreutils.git/commit...
pixelbeat··on dd built-in progress introduced in coreutils 8.24
That's more awkward and outputs a separate line per interrupt. Also to do that programmatically (to provide progress to a program using dd underneath) is tricky to do robustly due to the default disposition of SIGUSR1. SIGUSR1 is really only a hack provided on systems that don't support SIGINFO (which is easier to generate and whose default disposition is easier to handle robustly)
pixelbeat··on dd built-in progress introduced in coreutils 8.24
As mentioned in the online manual at http://gnu.org/software/coreutils/dd that command line syntax is inspired by the DD (data definition) statement of OS/360 JCL
pixelbeat··on RemixOS – Android for the desktop
Yes something like termux might be a sufficient addition to "android". Note a standard GNU coreutils build is about 15M, though one can configure to use a multi-call binary like busybox, reducing the install size to about 1M. See the coreutils-single subpackage in Fedora rawhide for example: http://pkgs.fedoraproject.org/cgit/rpms/coreutils.git/commit...
pixelbeat··on Reviewing 20 years of email clients
s/email/mozilla email/

Thanks for the reference ;)

pixelbeat··on Linux Performance Analysis
Note that's virtual memory. For determining real RAM usage for programs, I find https://github.com/pixelb/ps_mem extremely useful
pixelbeat··on GNU Coreutils Gotchas
Yes GNU spends a lot of time improving performance. A couple of examples from the most recent release, which you might think were too simple to optimize significantly:

The yes command (which is generally useful for generating repetitive text):

    $ yes-old | pv > /dev/null ^C
    ... 55.8MiB/s ...
    $ yes-new | pv > /dev/null ^C
    ... 3.44GiB/s ...
Details on that fairly simple change are at http://git.sv.gnu.org/gitweb/?p=coreutils.git;a=commitdiff;h...

Also we more than doubled the speed of wc -l (by avoiding function call overhead):

    $ yes | pv | wc-old -l ^C
    ... 230MiB/s ...
    $ yes | pv | wc-new -l ^C
    ... 558MiB/s ...
Also we now generate an infinite stream of integers more efficiently too:

    $ seq-old inf | pv > /dev/null ^C
    ... 13.3MiB/s ...
    $ seq-new inf | pv > /dev/null ^C
    ... 497MiB/s ...
pixelbeat··on GNU Coreutils Gotchas

    xargs basename -a < a.txt
pixelbeat··on GNU Coreutils Gotchas
Discussed at http://lists.gnu.org/archive/html/coreutils/2011-01/msg00080...

Summary of that is you can filter using xargs like:

    get_file_paths | xargs basename -a
No point having two ways to do something, especially when there are caveats about enabling stdin processing
pixelbeat··on OS X 10.11 buffer overflow with deep filesystem hierarchy
There have been a couple of recent changes to the latest GNU version of that related to flexible array members:

https://github.com/coreutils/gnulib/blame/master/lib/fts.c#L...

https://github.com/coreutils/gnulib/commit/49078a78

pixelbeat··on Requestdiff – Send two HTTP requests and visualize any differences
Personally I use these from the command line:

http://www.pixelbeat.org/scripts/urldiff http://www.pixelbeat.org/scripts/idiff

See also mergely which supports diffing URLs: http://pixelbeat/programming/diffs/#mergely

pixelbeat··on How to Shuffle and Sample on the Command-Line
This uses a few utils and techniques to deal 5 random cards:

https://twitter.com/pixelbeat_/status/587703133717057537

    paste -d '' <(printf '%s\n' $(seq 2 9) T J Q K A | sed 'p;p;p') \
    <(yes $'H\nD\nS\nC' | head -n52) |
    shuf -n5
pixelbeat··on Using htop to generate a live website background
Neat thanks!

You can generate HTML (and from that png or whatever) with http://www.pixelbeat.org/scripts/ansi2html.sh like

    $ COLUMNS=80 timeout 1 htop > t.ansi
    $ ansi2html.sh --bg=dark < t.ansi > t.html
I've requested a -b, --batch option for htop, to simplify this mode of operation. https://github.com/hishamhm/htop/issues/282 Then you could just:

    $ COLUMNS=80 htop -b | ansi2html.sh ...
pixelbeat··on “I have [bash] history back to ~2003”
It's a joke
pixelbeat··on “I have [bash] history back to ~2003”
You owe the GNU coreutils project $392

https://github.com/diafygi/gnu-pricing

pixelbeat··on The Fastest Blog in the World
It's great to see focus on performance, especially after all the recent stories on web bloat.

I've noted a few techniques I've used on my blog to get pages served in a single request to users, including using SSIs and avoid cloudflare's "rocket loader"

http://www.pixelbeat.org/docs/web/about/#performance

pixelbeat··on Should I Use Signed or Unsigned Ints?
Some practical notes for handling/avoid signed integer overflow

http://www.pixelbeat.org/programming/gcc/integer_overflow.ht...

pixelbeat··on GNU coreutils 8.24 released
No, just that sparsification will work better.

1. cp will detect holes < 128KiB

2. cp --sparse=always it will avoid speculative preallocation used on file systems like XFS, which can have a significant impact on sparsification.

pixelbeat··on GNU coreutils 8.24 released
One thing not mentioned in the link are some performance improvement details.

For example the yes command (generally useful for generating repetitive text):

    $ yes-old | pv > /dev/null ^C
    ... 55.8MiB/s ...
    $ yes-new | pv > /dev/null ^C
    ... 3.44GiB/s ...
Details on that fairly simple change are at http://git.sv.gnu.org/gitweb/?p=coreutils.git;a=commitdiff;h...

It's interesting there are so many potential improvements in such widely used tools. For example we also more than doubled the speed of wc -l (by avoiding function call overhead):

    $ yes | pv | wc-old -l ^C
    ... 230MiB/s ...
    $ yes | pv | wc-new -l ^C
    ... 558MiB/s ...
For completeness, we now generate an infinite stream of integers more efficiently too:

    $ seq-old inf | pv > /dev/null ^C
    ... 13.3MiB/s ...
    $ seq-new inf | pv > /dev/null ^C
    ... 497MiB/s ...
p.s. I would have used the new `dd status=progress` feature rather than pv above, though pv gives more accurate values due to its use of splice(2). Something to consider for a future version of coreutils
pixelbeat··on Vim's 400 line function to wait for keyboard input
Note Emacs uses http://git.sv.gnu.org/gitweb/?p=gnulib.git to abstract away platforms differences
pixelbeat··on iTerm2 Shell Integration
The popup on command completion is a neat feature actually which I've appreciated in gnome-terminal on Fedora 22, nicely integrated into the notification system
pixelbeat··on When ‘int’ is the new ‘short’
In my experience we can't rely on manual handling of these integer overflow issues, especially with changing compiler behavior over time.

I've noted some compile time and run time checking options at:

http://www.pixelbeat.org/programming/gcc/integer_overflow.ht...

pixelbeat··on Asciinema
Another option for the common case of static output is to render command or script(1) outout directly to static html with something like http://www.pixelbeat.org/scripts/ansi2html.sh
pixelbeat··on The Unix way: A note on composability
GNU cp already has the --one-file-system option.

Possibly mv might also benefit from this option to error out on EXDEV?

Feel free to discuss these things on coreutils@gnu.org

We try to be accomodating, and your suggestions here definitely have merit and are worth further discussion.

pixelbeat··on Unix filesystems: How mv can be dangerous
True, symlinks complicate things.

We have to be careful about doing too much (adding edge cases and complexity), though in this case how about:

    canon_src = canonicalize(src);
    canon_dst = canonicalize(dst);
    int ret = rename(src, dst);
    if (ret == EXDEV) {
      /* Note if there are symlinks being recreated
         between the canonicalize() and rename() above,
         then it's better to not fall back to a
         cross device copy anyway.  */
      if (common_prefix(canon_src, canon_dst))
          ret = EINVAL; /* Treat like "copy into self" */
    }
pixelbeat··on Unix filesystems: How mv can be dangerous
http://git.sv.gnu.org/gitweb/?p=coreutils.git;a=blob;f=src/c...
pixelbeat··on Shelling Out Sucks (2012)
There are large advantages to shelling out though.

1. Not reinventing the wheel

2. Implicit use of multiple cores

Things could be improved I agree.

For example posix_spawn() could be efficiently implemented on glibc to avoid kernel overhead for fork()+exec() from large processes. Then language runtimes could use that to implement their "shell out" routines.

Also pipefail doesn't cater for SIGPIPE as detailed at http://www.pixelbeat.org/programming/sigpipe_handling.html which can be awkward.

← PreviousPage 2 of 4Next →