HNHacker News
TopNewBestAskShowJobs

pixelbeat

1,141 karma · joined September 12, 2008

http://www.pixelbeat.org/ http://twitter.com/pixelbeat_
submissionscomments
pixelbeat··on Gonix – Unix tools written in Go
It's also worth pointing out, since busybox was mentioned a few times, that the latest release of coreutils has the ./configure --enable-single-binary option to build as a multi-call binary linke busybox etc.
pixelbeat··on Gonix – Unix tools written in Go
A few data points on GNU coreutils.

1. It has a fairly good test suite, that rewrites should leverage. That can be easily done by setting $PATH to prepend the dir of the new tools, and running `make check`

2. To give an indication of the size of coreutils:

$ for r in gnulib coreutils; do (cd $r && git ls-files | tr '\n' '\0' | wc -l --files0-from=- | tail -n1); done

985050 total 243154 total

pixelbeat··on Bug 1202858 – Restarting squid results in deleting all files in hard-drive
That's a bug IMHO which I reported at http://lists.gnu.org/archive/html/bug-bash/2015-02/msg00052....

I've collated other mishandling of closed pipes at: http://www.pixelbeat.org/programming/sigpipe_handling.html

pixelbeat··on Moved ~/.local/share/steam. Ran steam. It deleted everything owned by user
The trailing /* is to delete the directory _contents_ rather than the dir itself. It would be safer to receate the dir as then you can possibly hit some of the inbuilt rm protections
pixelbeat··on Ask HN: Open-source projects that could use documentation help?
Note newer versions of coreutils will have links in the man pages directly to online manuals through http://www.gnu.org/s/coreutils/ls etc.

BTW the coreutils man pages are a subset of the full manual, and I think that adding more information to man pages can make things harder to find.

The thinking at present is that linking directly to a web page for the full manual is what most users would prefer.

pixelbeat··on My experience with using cp to copy 432 million files (39 TB)
I found an issue in cp that caused 350% extra mem usage for the original bug reporter, which fixing would have kept his working set at least within RAM.

http://lists.gnu.org/archive/html/coreutils/2014-09/msg00014...

pixelbeat··on Cross-platform Rust Rewrite of the GNU Coreutils
I see what you're saying, but then you'd have people complaining there was too much text in the man page. Note the info pages are available on the web so you could have a wrapper like:

    cinfo() { xdg-open "http://www.gnu.org/software/coreutils/manual/html_node/$1-invocation.html#$1-invocation"; }
Which you could use like:

    cinfo dd
pixelbeat··on Cross-platform Rust Rewrite of the GNU Coreutils
Note GNU coreutils has a good test suite which just calls out to the various tools from shell and perl scripts. It should be easy enough to run this implementation through it. More effort goes into the coreutils tests than the code.

One of the hardest parts of coreutils is keeping it working everywhere, including handling that buggy version of function X in libc Y on distro Z. That's handled for GNU coreutils by gnulib, which currently has nearly 10K files, and so is a significant project in itself.

Some stats:

coreutils files, lines, commits: 1072, 239474, 27924

gnulib files, lines, commits: 9274, 302513, 17476

pixelbeat··on Linux memory manager and your big data
Well with coreutils like wc etc. you get low level control over such things, but with postgres there would be less flexibility. For illustration one could use GNU dd to take advantage of the page cache only for readahead purposes like:

    dd if=clickstream.csv.1 iflag=nocache bs=1M | wc -l
A more common technique is to bypass the page cache altogether and is often use to avoid the many unfortunate characteristics of the current Linux VM. This is done usually with directIO:

    dd if=clickstream.csv.1 iflag=direct bs=1M | wc -l
Now postgres might be able to use directIO as an option?

Another related problem with too much caching when writing to slow device can be seen in this thread: http://thread.gmane.org/gmane.linux.kernel.mm/108708 That thread actually describes two problems. 1. That Linux waits too long before writing 2. When it does write large amounts to a slow device it locks out everything else

pixelbeat··on Using ‘screen’ - The Absolute Essentials
Here's an actual quick reference for screen:

http://www.pixelbeat.org/lkdb/screen.html

pixelbeat··on Twitter RSS
Which would be fine. RSS usually has a limited number of entries
pixelbeat··on Twitter RSS
This gives RSS for a particular users' public tweets.

I'd love a service to provide an RSS feed of my timeline (would require giving auth of course)

pixelbeat··on Cloudflare down again
More on the story http://blog.cloudflare.com/the-ddos-that-almost-broke-the-in... https://news.ycombinator.com/item?id=5450410
pixelbeat··on Cloudflare down again
Ok again now. Probably caused by http://www.nytimes.com/2013/03/27/technology/internet/online...
pixelbeat··on Cloudflare down again
Slowly coming back...
pixelbeat··on Anonymous Gets Into The Pycon Incident
Seems to have had the desired effect? https://www.facebook.com/SendGrid/posts/10151502570463967
pixelbeat··on Creating a simple blog system with a 500-line bash script (2011)
It's cool to see static web sites becoming popular. As for comments (dynamic) you could integrate with discuss, or as I do wrote my own simple comments app for google app engine.

Details here: http://www.pixelbeat.org/docs/web/feed.html

pixelbeat··on Python command line oneliners
Related: http://www.pixelbeat.org/programming/evanescent_python.html
pixelbeat··on Smem memory reporting tool
See also the http://www.pixelbeat.org/scripts/ps_mem.py tool for reporting real mem usage for programs
pixelbeat··on DOS is long dead, long live FreeDOS
Also dosbox is really cool:

http://www.pixelbeat.org/misc/dosbox/

pixelbeat··on ASCII bit trick to convert lowercase to uppercase and back
Yes ASCII (and UTF-8) were designed well.

http://www.pixelbeat.org/docs/utf8_programming.html

pixelbeat··on Use ack instead of grep to parse text files
I find ack slow and overcomplicated (in implementation at least). As an alternative consider:

http://www.pixelbeat.org/scripts/findrepo

pixelbeat··on ps_mem.py
I found this in my bookmarks:

http://www.lshift.net/blog/2008/11/14/tracing-python-memory-...

pixelbeat··on ps_mem.py
Wow github is slow at present. The latest script is also at:

http://www.pixelbeat.org/scripts/ps_mem.py

pixelbeat··on Measuring the shared RAM usage of a process
Here's a python script to do much the same thing:

http://www.pixelbeat.org/scripts/ps_mem.py

pixelbeat··on GNU grep is 10x faster than Mac grep
I notice these Mac tools becoming a bit stale. sort is derived GNU sort, but from some ancient version. I guess this might be due in part to these tools now being GPLv3 ?
pixelbeat··on GNU grep is 10x faster than Mac grep
For a similar but much faster tool than ack, which simply wraps `find` and `grep` in the UNIX tradition, see:

http://www.pixelbeat.org/scripts/findrepo

pixelbeat··on Charm shutting down because of unfixable kernel panics with Rails on Ubuntu
This seems like a bit of a cop out? Why not try a new kernel or separate stack like Fedora for example?
pixelbeat··on Learn a Programming Language Faster by Copying Unix
I had this idea to demo basic python concepts. Here's an implementation of ls with links to further info:

http://www.pixelbeat.org/talks/python/ls.py.html

pixelbeat··on Hacking ls -l
I enable this for GNU ls like:

    alias ls="BLOCK_SIZE=\'1 ls --color=auto"
The above is a bit hacky and not very UNIXy as it's lumping more logic into ls, rather than splitting out into functional units.

Number formatting being a very common requirement, I've proposed a design for a new numfmt GNU coreutil

http://lists.gnu.org/archive/html/coreutils/2012-02/msg00085...

which would be used like:

    ls -l | numfmt --field=5 --format=%'d
← PreviousPage 3 of 4Next →